Skip to main content

Your monitoring dashboard can lie by omission

· 2 min read

A performance guard on a "latest status" query quietly turned a dead collector into a clean bill of health. Here is how the final review of env-sync Viewer caught it before anything shipped.

The setup​

env-sync Viewer has two parts. A scheduled collector records secrets-drift status for every target in the fleet. A read-only dashboard shows each target's latest state. The dashboard exists to tell a human when something is wrong.

The first version looked up each target's latest state only within the last 14 days. That seemed reasonable: there was no retention policy yet, and nobody wanted a full history scan on every page load. The line sat next to ordinary query code, far from the auth and access-control code that got the most scrutiny.

The question that found it​

The final adversarial review asked one question: what does this page show if the collector has been silently broken for a month?

It showed nothing. There were zero targets and zero stale warnings, just an empty, calm page. The grouping logic was correct. The bug was an assumption hidden in the WHERE clause: that a recent row would always exist. When a target went silent, it didn't look stale. It stopped existing.

The fix​

We removed the time bound. Two unbounded DISTINCT ON queries find each target's latest-ever snapshot and its latest-ever healthy snapshot. The display layer then compares the timestamps to now. A target that has ever been seen now always appears, and it ages into stale after STALE_AFTER_HOURS. A regression test against real Postgres seeds a target last seen 60 days ago and asserts that it is still on the page, marked stale.

The full scan is fine at current volume. When a retention policy arrives and the query needs a bound again, indexes come first. Silently dropping rows does not come back.

The lesson​

For any monitoring view, ask two questions: "what does it show when the thing is failing?" and "what does it show when the thing has been silent for a very long time?" A false all-clear is worse than having no dashboard, because people trust it.

Read more about staleness and the rest of the design in Architecture.