Insights

Intelligence

Why market monitoring fails without evidence preservation

Victory Harbour Operating Team · · 7 min read

A question that ends most market intelligence programmes: “where did that come from?”

Someone presents a finding. A competitor has repositioned, a price has moved, a supplier has entered a new category. The room is interested. Then someone asks how confident we are, and what exactly was seen, and when.

And the answer is that a tool flagged it three weeks ago, the summary is in a document, and the original page has since changed. Nobody can reproduce the observation. The finding is probably true. But it cannot carry a decision that costs money, so the decision is deferred, and the monitoring programme quietly loses its purpose.

This is not a tooling failure. It is a design omission, and it has become more common — not less — since summarisation got cheap.

AI made summarisation cheap and provenance expensive

Before language models, market monitoring was expensive and slow, which forced discipline. An analyst read a limited number of sources carefully and could tell you what they had read.

Now a system can process hundreds of sources and return a fluent paragraph. The cost of producing a conclusion has collapsed. But the conclusion arrives detached from what produced it — a summary of a summary, with the original wording, the exact date, and the source context stripped out along the way.

Two properties of the current situation make this worse than it sounds:

Sources change and disappear. A competitor's pricing page today is not the one you looked at last month, and there is no version history. If you did not capture it at the time, the evidence is simply gone.

Fluent output invites over-trust. A well-written paragraph reads as more certain than the observation behind it. When a model says a competitor “has shifted towards enterprise customers”, that might rest on a rewritten homepage headline, or on three job postings and a pricing change. Those are very different levels of evidence, and the summary does not distinguish them.

The result is a system that produces confident-sounding conclusions nobody can check. That is worse than no monitoring at all, because it feels like information.

“Confident-sounding conclusions nobody can check are worse than no monitoring at all, because they feel like information.”

What evidence preservation actually means

Not “keep a link”. A link is a promise that the page will still exist and still say the same thing, and it will not.

At the moment an observation is made, capture and store:

  • The raw content — the actual text or page as it was, not a summary of it
  • A visual record — a screenshot for anything where layout or presentation carries meaning
  • The exact timestamp of observation
  • The source identity and how it was accessed
  • The prior state — what this is a change from, which is what makes it a signal rather than a fact
  • The processing trail — which model, which prompt, which version produced the interpretation

The last item matters more than it appears. If you upgrade the model six months from now and the same input produces a different reading, you need to be able to tell that the world changed rather than your instrument.

Storage cost for all of this is negligible. The reason it is usually absent is that nobody specified it at the start, and it cannot be added retrospectively — you cannot preserve evidence for an observation you already made.

Change detection needs a baseline

A related failure: systems that report what a source says rather than what changed.

“Competitor A's pricing page lists three tiers starting at X” is not intelligence. It is a description. It becomes intelligence when it is “Competitor A introduced a third tier on 14 March; the entry price moved from X to Y; here is the page before and after.”

That requires a stored prior state. Which means the first weeks of any monitoring programme produce almost nothing useful, and this has to be said out loud at the start — otherwise the programme is judged a failure before it can work.

The false positive problem

Broad monitoring generates volume, and most of that volume is noise. Sites get redesigned. Wording changes without meaning changing. A job posting is a backfill, not an expansion.

If every flagged item reaches a human, they stop reading within a fortnight. If none do, you are trusting an automated judgement about what matters commercially — which is exactly the judgement a model is least equipped to make, because it depends on your strategy, not on the text.

The workable arrangement is a validation gate with a deliberately narrow throat:

authorised sources (you approve the list)
   → automated collection and change detection
   → evidence bundle assembled per detected change
   → material findings only → human validation
   → decision-ready report, every claim linked
     to preserved evidence

Two design points. First, you approve the source list — this prevents the system drifting into low-quality sources that generate noise. Second, the human validates rather than reads everything; their job is to judge commercial materiality, which is the thing only they can do.

And track the rejection rate. If a large share of flagged signals are rejected as immaterial, the detection thresholds are wrong. That number is the system's own quality measure, and it should be published internally.

What decision-grade looks like

A report is decision-grade when someone can challenge any material claim in it and be shown, within a minute, exactly what was observed and when. Not the summary. The artefact.

In our own Operating Labs, weekly monitoring coverage expanded from 12 to over 90 authorised sources per analyst without added headcount, and report assembly time fell from around two days to under three hours, with every material finding linked to preserved evidence. Internal lab conditions, not client results — but the second half of that sentence is the part that matters. Coverage without provenance is not an improvement.

“Coverage without provenance is not an improvement.”

When not to do this

If your market has three competitors you already know personally, and prices move once a year, a phone call is faster and better. Automated monitoring earns its cost when sources are many, changes are frequent, and the cost of noticing something late is real.

Start with the workflow that matters

If your market view is assembled from memory and screenshots, that is a workflow worth examining before scaling it. This is the discipline the AI Market Intelligence System is built around.