Skip to main content

Architecture

apps/collector (scheduled) Postgres, env_sync schema apps/web (Next.js)
execFile("env-sync diff ---> collector_runs <--- auth + allowlist gate
--format json") INSERT drift_snapshots SELECT server-side compare only;
(collector drift_snapshot_keys (viewer browser never sees a
role) role) fingerprint

D4: the collector shells out, it never imports​

The most important design decision is also the least visible one. The collector runs env-sync as a real subprocess (execFile) and parses its stdout. It does not import @env-sync/core or any other code from the env-sync repo, and it never will.

A library import (option "D1") was the obvious choice. env-sync's TypeScript core already exports the exact computeDiff function a collector needs, with full type safety across the boundary. It was rejected for one reason: the TypeScript implementation is explicitly transitional. A Go implementation already exists, and the plan is to cut over to it once it reaches fixture parity. A collector coupled to TypeScript internals would then be stuck on a deprecated package, or it would need a rewrite at the same moment as the cutover.

Instead, the collector (option "D4") depends on exactly one thing: the documented, versioned env-sync diff --format json contract (schemaVersion: 1, stable KeyStatus tokens, and a strict exit-code rule). Both the TS and Go implementations must satisfy that contract identically. A TS→Go cutover changes nothing in the collector. You put a different binary on PATH. The separate repository keeps this boundary real instead of just a line in a doc.

The cost is real, and we accept it: one subprocess per inventory entry, and contract drift is caught at parse time (schemaVersion ≠ 1 → the run is recorded as failed), not by the compiler.

How a run is classified​

The exit-code rule does the work. Exit 0 means the JSON document is trustworthy, and drift or per-target errors inside it are data. Any non-zero exit means stdout is not trusted at all.

CLI resultoutcomeSnapshot rows written
exit 0, no service or target errorssuccessone per target, one key row per key
exit 0, some services[].error or target errorpartialsame, and errors recorded
non-zero exit, invalid JSON, or unknown schemaVersionfailednone (only exit code + stderr)
exit 0 but the DB rejects the document (e.g. a CHECK violation)failednone (transaction rolls back, failed run written)

Also:

  • dotenv targets are dropped. A dotenv file exists only on the machine that wrote it, so a central collector cannot observe it.
  • Redaction is the CLI's job, because only the CLI holds the 1Password values it redacts against. The collector only caps error strings at 2000 characters, and that cap is not redaction.
  • Run and snapshot IDs are generated client-side. The collector's role has no SELECT, so it cannot use RETURNING.

Row-level security: app vs. everyone, not per user​

Everything lives in a dedicated env_sync schema, not public. public is the schema Supabase's PostgREST exposes by default. Access comes from two NOLOGIN group roles:

RoleCan doCannot do
collectorUSAGE on env_sync, INSERT on all three tablesread anything
viewerUSAGE on env_sync, SELECT on all three tableswrite anything
everyone else (PUBLIC, anon, authenticated, service_role)nothing (REVOKE ALL)—

All three tables have ENABLE and FORCE ROW LEVEL SECURITY, with one explicit policy per role. The SQL states one subtle point honestly: FORCE subjects the table owner to RLS, but it does nothing to a BYPASSRLS role such as Supabase's service_role. The REVOKE ALL is what stops that role. The same SQL file works on plain Postgres and on Supabase, because the Supabase-specific revokes are guarded by pg_roles checks.

Why not per-user policies? Every row is fleet-wide data with no natural owner, so a per-user policy would mean inventing ownership that does not exist. At the database layer, the boundary that matters is which process: the collector, the viewer, or anything else. Which human may look at the dashboard is decided one layer up, by the allowlist. That makes the allowlist a real security boundary, not a nicety. It is the reason the fix below mattered as much as an RLS bug would have.

The RLS suite runs against a real postgres:16. It re-applies the SQL after it grants ALL to fake Supabase roles, to prove the revokes strip real privileges. Then it drops every policy and shows that the positive cases fail, to prove the suite is not deny-only.

The access gate: authentication and an allowlist, checked twice​

The viewer checks two things, and both must pass. It runs the checks twice: in src/proxy.ts on every request, and again in requireViewer() in every page that reads data.

  1. Authentication: a @supabase/ssr cookie session, verified with auth.getUser(). No session → redirect to /login.
  2. Allowlist: VIEWER_ALLOWLIST (comma-separated, case-insensitive) and email_confirmed_at must be set. Not listed or not confirmed → 403 from the proxy, 404 from the page. An empty allowlist lets nobody in.

The email_confirmed_at check came from the final adversarial review. The first version matched on user.email alone. If the Supabase project allowed sign-up without email confirmation, anyone could register an allowlisted address they don't own and get in. An allowlist is only as strong as the identity claim beneath it.

Data minimization​

Per-key fingerprints are stored on purpose, because they show which key drifted. That also makes them the most sensitive data in the schema. Truncated, unsalted fingerprints of low-entropy values can be guessed, and equal fingerprints across services reveal reuse. So they are compared only on the server. The page is built entirely from Server Components that receive a FleetView type with no fingerprint field. Per key, the browser sees only match, drift, missing, present, or unverifiable. A test seeds known fingerprints and asserts that they appear in neither the view model nor the rendered HTML.

Staleness, and the bug that made silence look like health​

A target is stale when it has no status = 'ok' snapshot within STALE_AFTER_HOURS (default 26). The check is per target, not per run:

  • One target that fails for a long time doesn't make its healthy siblings stale.
  • A collector that keeps completing runs doesn't make a broken target look fresh.
  • A dead collector makes every target stale.

The first version had a bug here, and the final adversarial review caught it before the tool ever shipped. The query that finds each target's latest state was bounded to a 14-day lookback window, which looked like a harmless performance guard. But a target that was silent for longer than that window didn't show up as stale. It disappeared from the dashboard. If the collector was dead for a month, the page showed zero targets and zero warnings. An operator would reasonably read that as "all clear". It was the exact failure the tool exists to prevent.

The fix removed the bound. Two unbounded DISTINCT ON queries find each target's latest-ever snapshot and its latest-ever ok snapshot. The display layer then decides staleness. A regression test against real Postgres seeds a target last seen 60 days ago and asserts that it still appears, marked stale. The general lesson: for any monitoring view, ask what it shows when the monitored thing has been silent for a very long time, not only when it is failing loudly.

The credential model, named honestly​

The collector host holds real credentials for every platform env-sync supports:

  • 1Password Service Account: can read the plaintext value of every secret in its scoped vaults. The collector only chooses to store fingerprints.
  • Render: API keys are account-wide per team, with no read-only or per-service scope.
  • Vercel: project-scoped, but read-only access is not confirmed.
  • AWS SSM and GitHub: the only two with a real, enforced read-only option.

That is why the collector must run in a dedicated CI environment with protected branches and required reviewers, and never on a shared runner. One further reduction is available: one job per platform or Render team, each with its own inventory and narrower credentials.

Known gaps​

These are named as follow-ups, not solved:

  • No dead-man's-switch alert yet. The dashboard no longer hides a dead collector, because every target turns stale. But nothing pages anyone.
  • No retention policy. Storage grows as runs × targets × keys.
  • The allowlist is an env var. It moves to a table (still enforced in the app) if the viewer list grows.
  • No real-infrastructure run yet. The stack is verified against real local Postgres and fakes only. It has not run against a live Supabase project or real platform credentials.