Listening platforms give you two ways to find themes, and they are good at different jobs. Using one
for both is the most common way a listening programme becomes untrustworthy.
A curated Boolean taxonomy owns the recurring reported numbers. It runs on
everything, it is auditable, and it stays stable week to week. Unsupervised discovery
runs only on the residual the taxonomy could not place, and its output is next month's taxonomy
entries rather than this month's report.
A model re-fit each week silently renumbers and reshapes its own topics, so "price mentions doubled"
can mean the topic changed rather than the world did. For a figure in a recurring report, that is
disqualifying. Equally, running discovery across the whole corpus just rediscovers the themes the
rules already cover and buries the new signal underneath them.
Coverage, not accuracy, is the health metric. Taxonomy accuracy on placed mentions
is close to meaningless, because the rules were written against that vocabulary, so it measures
consistency. Coverage falling is the real signal. It means the conversation moved and the rules did
not.
Every recurring report passes this before it leaves my hands.
These are the working repositories, here for anyone who wants to check the arithmetic. The case
studies are written to stand on their own without them.