Lab note · from FableWatch
Neon never scales to zero under a one-minute cron job
A Neon compute scales to zero only after five minutes with no activity, so a cron job that queries the database every minute keeps it running around the clock. On the Free plan that uses up the monthly compute allowance, and the compute is then suspended until the next billing period. FableWatch hit exactly this in July 2026. We rebuilt its live state on Vercel Blob with three rules: an ordinary tick makes no Blob writes or lists, anything that has to be atomic is a create-only object, and a failing store emails someone.
Symptoms
- In mid-July the database started refusing connections. The data was intact, just unreadable until the allowance reset on August 1.
- Nothing crashed. Reads were wrapped to fall back to an empty value, so pages showed zeros: the live "N watching" count returned
0. A signup API on the same deployment returned 503 on every request. - Nobody was alerted. The repo's own notes say the database "had been refusing connections for days before anyone noticed."
- The FableWatch waitlist wasn't affected. It had written to Blob from day one, with Postgres only as a best-effort mirror.
Why it happens
Neon's docs say a compute "scales to zero after an inactive period of 5 minutes," and that on the Free plan "this setting is fixed" (scale to zero). FableWatch's whole system ran off one endpoint, /api/cron/tick, called every minute: first by an external cron service, then from July 17 by Vercel Cron ("schedule": "* * * * *"). The status monitor inside that tick touched Postgres on every run (simplified from src/lib/monitor.ts before the migration):
// every tick, every minute
const state = await stateStore.upsert({
where: { id: "singleton" }, create: { id: "singleton" }, update: {},
});
// ... probe the model, compute the new readings ...
await stateStore.update({ where: { id: "singleton" }, data: persistData });
With a query every 60 seconds, the compute never sat idle for five minutes, so it never suspended. Neon meters the Free plan by compute time. At the time of writing, its plans page lists 100 CU-hours per project per month, "enough to run a 0.25 CU compute in a project for 400 hours/month." Our July notes recorded a different allowance, about 190 compute-hours. Either way, a compute that never sleeps runs about 730 hours in a month, well past both figures. The plans page also says what happens next: "your compute is suspended until the next billing period or until you upgrade."
The outage stayed silent because of the degradation code, not because of Neon:
export function dbReady(): boolean {
return Boolean(process.env.DATABASE_URL); // configured, not reachable
}
export async function getSubscriberCount(): Promise<number> {
if (!dbReady()) return 0;
try {
return await getDb().subscriber.count({ where: { status: "active" } });
} catch {
return 0; // an outage renders as "0 watching"
}
}
dbReady() checked whether the database was configured, not whether it was reachable. A suspended database therefore looked exactly like an empty one.
The fix
The migration shipped on July 21 (commit 77ac492). DATABASE_URL was removed from every Vercel environment, and live state moved to JSON objects in Vercel Blob. Blob bills differently from Neon. Per Vercel's pricing page, put(), copy() and list() are Advanced Operations. head(), and URL reads that miss the cache, are Simple Operations. The design rule sits at the top of src/lib/blobstore.ts: "The steady-state tick must cost ZERO advanced operations."
Reads on the tick go through the CDN
export async function getJsonCheap<T>(path: string, stalenessS = 30) {
const base = await baseUrl();
const bucket = Math.floor(Date.now() / (stalenessS * 1000));
const res = await fetch(`${base}/${path}?_t=${bucket}`, { cache: "no-store" });
return res.ok ? ((await res.json()) as T) : null;
}
Apart from one list() per cold instance to learn the store's base URL (skipped when it is set in an environment variable), this makes no SDK call. The result can be up to about 30 seconds stale, so it is only used where that is harmless: counters, indexes, and the timestamp gates below.
Periodic writers are gated
Hourly jobs still run from the one-minute tick. They check a timestamp cheaply and write only when the gate opens:
const at = await metaGetCheap("meta:credit-watch");
if (at && Date.now() - new Date(at).getTime() < CREDIT_CHECK_EVERY_MS) return null;
await metaSet("meta:credit-watch"); // one put() per hour, not per tick
The liveness stamp added later follows the same rule. It writes once every 15 minutes, about 96 writes a day, where an ungated stamp would cost about 1,440.
Create-only objects replace transactions
Postgres had provided the guards for one-shot actions. The old monitor claimed its "model is back" broadcast with a conditional updateMany, and only the tick that flipped online from false to true won. Blob has no transactions, so every such claim became an object that can be created only once:
export async function putJsonIfAbsent(path: string, value: unknown): Promise<boolean> {
try {
await put(path, JSON.stringify(value), {
access: "public", addRandomSuffix: false, allowOverwrite: false,
contentType: "application/json", cacheControlMaxAge: 60,
});
return true; // this caller created it
} catch (e) {
if (/exist/i.test(e instanceof Error ? e.message : String(e))) return false;
throw e;
}
}
Vercel's SDK docs say that by default "an error will be thrown if you try to overwrite a blob by using the same pathname." The codebase relies on that holding under concurrency, so that exactly one caller wins. The docs don't state that guarantee in those words.
The store is probed every tick
storeHealthTick() calls head() on one known object each tick. After five failures in a row it emails the operator, and it repeats the alert at most every six hours. A missing object still counts as healthy, because it proves the API answered.
How to check you've fixed it
- Test the gate, not the store.
tick-alive.tstakes itsreadandwritefunctions as parameters. The test passes awritethat records keys and a stamp from just inside the interval, then asserts that nothing was written (src/lib/tick-alive.test.ts). Any gated writer can be tested the same way. - Watch the meter. In Vercel's Observability section for Blob, an hour with no signups or alerts should add almost no Advanced Operations, only the gated writers. Simple Operations will keep rising with the per-tick
head()calls. - Break the store on purpose. With
BLOB_READ_WRITE_TOKENunset,storeHealthTick()counts a failure on every call, and the fifth call in the same process sends the email, provided the mail sender is configured. If that email never arrives, you are back to silent zeros.
Caveats
- Zero advanced operations is not zero metered operations. The health probe alone is one
head()a minute, about 43,000 Simple Operations a month, and the notification drain, when enabled, reads its open-event index with another. At the time of writing, Vercel's pricing table includes the first 10,000 Simple Operations on Hobby, and on Hobby going over the limit blocks Blob access for 30 days instead of billing. Which plan you are on decides whether this rule protects you. - No transactions means no read-modify-write. The repo's notes count four correctness bugs in 48 hours, all from concurrent writers editing the same mutable object. The code also notes that
head()metadata is eventually consistent, so reading an object back straight after a write proves nothing. Anything that must not be lost is now an append-only or create-only object. - Rarely changing content went into the repo as static JSON rather than living in a hot store.
- Versions at the migration: Next.js 16.2.9,
@vercel/blob2.4.0, Prisma 7.8.0 with@prisma/adapter-pg.
More lab notes
- Anthropic API key in an iOS app binary: move it to the KeychainA Swift string literal ships in the app binary, readable with strings. Store a user-supplied key in the Keychain with a ThisDeviceOnly class instead.
- AudioContext was not allowed to start: make the click the UIChrome starts an AudioContext made before any user gesture suspended. Create or resume() it in a click or keydown handler, like Meridian 7's power button.
- AXIsProcessTrusted false after rebuild: ad-hoc signing and TCCAn ad-hoc signature's designated requirement is the build's cdhash, so changed code no longer matches its Accessibility grant. Sign with a stable identity.
- CGWindowListCreateImage unavailable in macOS 15: use SCStreamCGWindowListCreateImage is deprecated in macOS 14 and a compile error from a macOS 15 target. Anchor still uses it; here is the SCStream replacement.