Ligang Yan颜力刚

A Vibe Coding Guide: Mistakes I Made and a Self-Check List for Your AI

What I learned from 787 commits written with Claude and some forty mistakes along the way. Every item comes with a self-check prompt: paste the whole post into your AI and it can go through your project item by item.

aivibe-codingclaude-codeengineeringcloudflareseo

中文版:Vibe Coding 指南:踩过的坑和给 AI 的自查清单

This is a project I worked on recently. In four months its repository got 787 commits. Leaving out merge commits, 547 of the remaining 566 carry Co-Authored-By: Claude, about 97%. 336 of 350 PRs were merged, and the days with any commits add up to only 40. The stack is a TypeScript monorepo: the backend is one Node process on Cloud Run; the frontend is a React single-page app with prerendering, served from Cloudflare Workers; a Worker in front of the API handles the allowlist, caching and the nightly jobs; the database is Postgres on Neon; and almost all traffic comes from Google Search.

This post collects the mistakes from those four months in one place, in two groups: how to work with the AI, and the traps in each platform. Choice of language and framework isn’t covered, since that depends on the project. I also left out the things everyone already knows (don’t commit secrets, parameterize your SQL) and kept only the ones that aren’t obvious.

Every item ends with an “AI self-check” block written for an AI. The way I picture this being used: you paste the whole post into your own AI and have it check your project against it. People can read it straight through too. The first half of each section is what happened to me; the self-check block after it is for the AI to run.

How to hand this to your AI

Paste the whole post into your coding agent (Claude Code, Cursor, Codex, any of them), or give it the link, and add one line: “follow the general instruction in the post”. The general instruction is this:

General instruction for the AI

This is a vibe coding guide. The narrative parts describe the author’s project and are background only: the time windows, config values and thresholds in them belong to the author’s project and are not requirements for mine. The only things to execute are the blocks titled “AI self-check”, 45 in total, numbered A1 to G8. Most have three parts: “Applies to / Check / Fix”. C4 and F7 constrain how you yourself debug and cite data; just follow them during this work. A11, D8 and G8 bundle several small checks; report them separately (A11-1, A11-2, …).

Go through them against my current project:

  1. First read the project structure, dependencies, deployment config and agent instruction files (CLAUDE.md, AGENTS.md, .cursorrules and so on), and decide whether each item applies. Groups D to G are platform-specific (Cloudflare, Cloud Run, Postgres and Neon, Google Search); skip a whole group if I don’t use that platform, with one line of reasoning.
  2. For the items that apply, follow “Check” through the code and config and cite specific files and line numbers. Don’t conclude from memory. Anything that needs the live environment, a cloud console or the production database to confirm: write down the command you would run or the page you would look at, mark it “needs human confirmation”, and don’t connect to it on your own yet.
  3. Summarize in one table: item number, verdict (fine / problem / needs confirmation / not applicable), evidence, suggested change, rough effort. Sort by impact, largest first.
  4. Show me the table and wait for my go-ahead before changing anything. Ask me separately about every change that deletes data, runs a migration, touches deployment config or changes permissions. For the big items (such as the planted-bug audit in B1 or the health-check script in C3), ask whether to do them and how far to go first.
  5. Once the fixes are in, turn every check that can be expressed as a lint rule or a test into one, so it fails automatically next time.
  6. At the end of the post there is a section “Rules to put in CLAUDE.md / AGENTS.md”. Take only the ones that apply to this project and merge them with the existing rules instead of pasting the whole block; rules that only concern one kind of file go into path-scoped rule files (see A1).

Now the actual content. Prefixes: A is collaboration and rules, B is testing, C is silent failures, D is Cloudflare, E is Cloud Run, F is Postgres and Neon, G is Google Search. Each item in D to G is labeled with the kind of mistake it was:

Label Meaning
platform The platform behaves counter-intuitively, and the docs don’t mention it or bury it
default The right answer was findable; the AI picked a default that looked reasonable
instruction A rule, doc, test or monitor in the repo locked the mistake in or kept it alive for days

A. Working with the AI: where rules live and what form they take

A1. The instruction file keeps growing

CLAUDE.md is put into every single turn of the agent’s conversation. For a long time I didn’t take that seriously, so the file kept growing.

On the day it was longest it was edited 16 times, and three of those commits were titled “update current phase”. The next day I went through it: more than half the content (directory layout, which files each service has) could be read straight from the code, and the “current phase” section was a snapshot of issue status that started going stale the moment it was written. That cleanup cut it from 12,445 characters to 3,801. The remaining rules moved down by directory into .claude/rules/*.md, each file declaring the paths it covers with paths:, so the agent only loads them when it touches those files. Nine days later it was back to 10,642 characters (by then in English, which takes more characters for the same content). Part of the growth was comments next to the dev commands that kept getting longer. They should have moved down too; I didn’t hold back.

How I decide now:

Where What
Top-level instruction file Things needed in every turn that can’t be read from the code: module boundaries, response format, branch rules, deploy windows
Path-scoped rules Conventions needed only when editing certain files: migration process, how to write tests, page conventions
Lint and tests Rules a machine can judge right or wrong
Issues Progress, roadmap, design discussion

AI self-check A1

  • Applies to: the project has CLAUDE.md, AGENTS.md, .cursorrules or a similar agent instruction file.
  • Check: measure its length; mark three kinds of content: things readable from the code (directory trees, file lists), status that will go stale (“current phase”, “TODO”, “this week”), and conventions that only concern one subdirectory.
  • Fix: delete the first kind, move the second to issues or the task tracker, move the third into path-scoped rule files (if the tool supports them). The top-level file should keep only the hard constraints needed in every turn.

A2. Write instructions in Chinese and the output drifts toward Chinese

On the day of the cleanup, CLAUDE.md was also switched from Chinese to English. The reason: Chinese instructions make the agent’s output lean toward Chinese, and this repo requires code comments, UI copy and developer docs in English. I measured tokens while I was at it: the same content is about 1,827 tokens in Chinese and 1,716 in English, only 6% apart, so saving tokens wasn’t a reason.

AI self-check A2

  • Applies to: the project has requirements on the language of code comments, copy or docs.
  • Check: whether the agent instruction files are written in the language the project expects as output; whether recent commits contain a language that shouldn’t be there.
  • Fix: write the instruction files in the target language; where another language must stay (content for Chinese readers, say), list it explicitly as an exception in the instructions.

A3. Stale docs send the agent the wrong way

On day three I deleted the whole docs/ directory: four documents, 1,084 lines, for one reason, which was that the docs kept disagreeing with the code. Conventions that still held moved into rule files; design discussion and progress went into issues from then on. The agent believes every word in the repo, and when a doc is out of date it acts on it anyway. Looking back after four months, the debugging sessions that went off track did so because of stale docs: the deployment doc said “CPU always allocated” long after the config changed, and the health-check notes named as the main suspect a function that had been fixed long before.

AI self-check A3

  • Applies to: every project.
  • Check: find sentences in the README, docs/, deployment docs and agent instruction files that describe concrete behavior (ports, timeouts, config keys, what a function does, where a service is deployed) and compare each one against the current code and config.
  • Fix: list the sentences that contradict the code, with evidence. Delete what can be deleted, keep the ones that only explain “why”, move progress notes to issues.

A4. If a machine can check a rule, don’t leave it only in a doc

Until five days before I wrote this, the repo had no linter. Architecture rules lived in prose and review: one server domain may not import another domain’s functions directly, code shipped to the browser may not touch Node built-ins, environment variables may only be read in one file, no console.* in the backend. The agent follows them most of the time, but “most of the time” means checking again in every review.

A PR that day turned them into 8 custom ESLint rules. The constraints we set on the rules turned out to matter more than the rules themselves:

  • No recommended rule sets. On ninety thousand lines of code that’s thousands of warnings at once, and nobody reads that many warnings.
  • Rules are only error or off, never warn.
  • A rule ships with zero violations: fix everything first, then turn it on.
  • Each error message names the rule document it comes from, so the agent can go read the source when it hits the error.
  • eslint-disable needs a stated reason, and review asks about it specifically.

The survey before turning them on found one real violation, plus a file that exported pure functions and network calls together.

AI self-check A4

  • Applies to: the project has architecture conventions written in prose (in instruction files, the README or comments).
  • Check: sort each convention into “machine-checkable” (import direction, banned APIs, a certain kind of call only allowed in one file) and “needs human judgment”. For the checkable ones, see whether a lint rule or test already covers them.
  • Fix: write custom lint rules for the uncovered ones (no-restricted-imports and no-restricted-syntax are often enough), all set to error, with the source of the rule in the message. Fix to zero violations before enabling.

A5. When the same fact is written in two places, add a test that compares them

Some things can’t be linted: the health-check script is bash, and its job list was copied from the TypeScript registry; .env.example has to match the environment variables the code actually reads. Each of these got a test that puts the two copies side by side. If one changes and the other doesn’t, the test fails.

AI self-check A5

  • Applies to: the same fact (job list, route table, environment variable list, ports, limit constants) is written separately in two or more places, especially across languages (shell, YAML, CI config and application code).
  • Check: list the duplicates and compare each pair for agreement right now.
  • Fix: generate from a single source where possible; where not, write a test that reads both sides and asserts they are equal.

A6. Build-time checks should only judge things you can measure

At one point I added a build-time wording check: scan the generated HTML and fail the build on banned words. I deleted it four days later. The post-mortem had three points: the word list came from intuition, not one entry from a real misuse; singular versus plural decided whether a page passed; and it couldn’t see client-rendered pages at all, so “build passed” got read as “the pages are clean”, which is worse than having no check. The CI performance budget added the same week stayed, because what it judges is a number you can measure.

The performance budget had its own failure once: an environment variable wasn’t set in CI, the bundler inlined it as an empty string, concluded from that the login SDK could never be used, and dropped it entirely. The budget measured 640 kB; what actually shipped was 724 kB. CI now sets a placeholder URL that is never requested.

AI self-check A6

  • Applies to: CI has build-time checks, budgets or quality gates.
  • Check: does each gate judge a measurable quantity, or a meaning that needs judgment? Does its actual coverage match the coverage people assume (for instance, it only scans static HTML but is treated as “whole site checked”)? Are the environment variables in the CI build the same as in the production build, and has an empty variable caused code to be tree-shaken away so that the CI bundle differs from production?
  • Fix: turn judgment-based gates into a review checklist; for gates with partial coverage, state what they cover in their name and output; give CI every environment variable that affects bundling, with placeholder values if needed.

A7. Let the LLM make the judgment, and store the result as a file

Some work needs judgment: classifying a few hundred entities one by one, picking related items. My approach is to have the AI in a local session make the calls and write a generated file, one line of reasoning per entry; a person reads the diff before merging. The server runs no LLM at all and only reads that table. If the table and the master data disagree, a test fails.

AI self-check A7

  • Applies to: the project calls an LLM in a request path or scheduled job for judgments whose results are fairly stable, such as classifying, tagging or picking.
  • Check: whether these judgments are recomputed on every run; whether the results are reproducible and reviewable by a person; whether the system still works without the LLM.
  • Fix: generate a data file offline with a reasoning field, put it under version control and have a person review it; read only the file at runtime; add a test that the file covers every entity it needs to.

A8. API tokens for agents should have narrow permissions

When accounts came back, I wanted external agents to be able to write posts. The identity service is hosted Neon Auth, which can’t act as an OAuth authorization server, and browser JWTs expire in 15 minutes, so instead of MCP I issued personal API tokens over plain REST:

  • The database stores only a SHA-256. The token is 32 random bytes, so bcrypt and other password-hashing functions aren’t needed: they protect weak passwords against brute force, and random bytes don’t have that problem.
  • Tokens have a fixed prefix, and the middleware picks the verification path from it. If it tried both paths, a pile of junk bearer tokens would turn into a pile of database queries.
  • The user id only comes from the credential. No route ever reads a user id from the body, query or headers.
  • Admin routes reject tokens and accept only browser sessions. A token given to an agent can’t publish and can’t start jobs.
  • The newsletter has its own scope: an agent holding it can draft, self-check and send test emails to itself, but there is no endpoint for sending to the list. The person has to be the one who clicks send.
  • The writing endpoint takes the whole post as Markdown, no field-level patches, because a post is one complete output from the model anyway; dryRun lets it self-check first.

AI self-check A8

  • Applies to: the project issues API tokens to scripts, agents or third parties.
  • Check: whether tokens are stored in plain text; whether invalid tokens can be amplified into database queries on the verification path; whether any route reads a user id from the request body or parameters; whether a token can call irreversible operations such as publish, delete, mass-send or permission changes; whether a token can mint a token with higher privileges.
  • Fix: store only hashes; distinguish credential types by prefix; take the user id only from the authentication result; require an interactive session or a separate scope for irreversible operations, off by default.

A9. Write each validation rule once

The length limits for titles and descriptions used to live in two files, one for the build and one for the content spec. A 162-character description passed in one place and got truncated in the other. The site build and the submission API now share the same validation code.

AI self-check A9

  • Applies to: the same validation exists in both frontend and backend, or in both a submission endpoint and the build.
  • Check: find limits, regexes and enums defined more than once and compare them.
  • Fix: move them into a shared module that both sides import; add a test asserting both sides reference the same constant.

A10. Be more careful deleting code than writing it

Version 2.0 is the biggest commit in the repo: 196 files, 17,863 lines deleted, 28 tables dropped. A v1.0.0 tag went on before the deletion, so anything deleted by mistake could be recovered. The migration SQL that drops the tables was generated but not run; the PR said “this deletes real data, a person decides”, and the person ran pg_dump first. Earlier, when the old frontend (281 files) was removed, three agents searched for leftovers separately and a fourth agent’s only job was to find their mistakes. Big refactors were accepted on build output: when five oversized UI files were split up, all 5,817 build artifacts were byte-for-byte identical before and after (with hashes in filenames masked). Even so, when the blog moved to edge rendering, testing before launch turned up four problems the unit tests couldn’t catch. One was that the RSS feed would go from 12 posts to 0 after the switch.

AI self-check A10

  • Applies to: a large deletion, rename, refactor or destructive migration is coming.
  • Check: whether there is a tag or branch from before the deletion; whether the destructive migration runs separately from the code change, with a backup first; whether there is a way to prove behavior didn’t change (build-output comparison, snapshot tests, recorded API replay).
  • Fix: tag first; generate migrations but don’t run them, leave that to a person; pick an artifact that can be compared and diff it before and after; after deleting, search again for references to removed symbols, paths and environment variables.

A11. A few small rules

AI self-check A11

  • Git: fetch before creating a branch, and branch from the remote main branch; before adding commits to an existing branch, confirm its PR hasn’t been merged. This repo had that rule from day one, and before long a set of stacked PRs still got merged into another feature branch, stranding the commits until cherry-pick rescued them.
  • Dev servers: when clearing processes on a port before starting, kill the process group. A supervisor like wrangler dev will immediately restart a child you kill on its own. The agent’s shell is non-interactive, so the old and new runs may share a process group; have a fallback that goes up only one level.
  • External input: don’t use type assertions to turn external data into domain types. Here an error field was asserted to be a string when it was actually a {code, message} object, so any 4xx blanked the page. Check every assertion on HTTP responses and third-party SDK return values and replace them with runtime validation.
  • CI all red: first check whether a runner was assigned at all (fails within seconds, no logs). That happened here three and a half days in a row, and it was an account problem. Don’t keep re-running; run lint, type checks and tests locally and say so in the PR.

B. Testing: all green doesn’t mean useful

B1. Planted bugs: 44% caught

A few days ago I did a test audit. There were 19,502 tests at the time, all passing. The method was crude: plant 341 realistic bugs in the source by hand (change > to >=, drop a term, flip a sign, swap the divisor for last period’s), run the relevant tests after each one, and see whether they go red. 151 were caught, 44%. The core calculation module was at 30%.

The ones that slipped through clustered around a few patterns, nearly all of them habits the AI falls into when writing tests:

  • Expected and actual values come from the same object, like expect(total).toBe(r.a + r.b); if the function computes b as 0, the test still passes.
  • The production function used as the reference answer: expect(header).toBe(cacheHeader()).
  • The variable under test never changes in the fixture. If a quantity is equal in every period, any bug involving it can never be detected.
  • Testing only the shape: toBeDefined(), > 0, a wide range from 30 to 500 when the actual value is 68.41.
  • Assertions inside a loop that never run. Four multilingual-link tests iterated over a list of pages; after the pages moved, they looped over zero pages and stayed green.
  • if (…) expect(…): when the branch isn’t taken, the test passes without checking anything.

The audit also fixed 9 production bugs. In one, an upstream data source had changed the spelling of an enum value more than a year earlier, and every new record since then had been silently dropped while the tests stayed green. The test count went from nineteen thousand to eighteen hundred because eighteen thousand per-row it.each cases were merged into a few aggregate tests: collect every failure into a list of [object, reason] and assert at the end that the list is empty. After the new tests were in, the same 341 bugs were planted again: 337 caught, 99%.

AI self-check B1

  • Applies to: every project with tests.
  • Check: pick the 5 to 10 most important functions (calculations, permissions, billing, data transforms). Plant 3 to 5 realistic bugs in each: flipped comparison, dropped term, negated sign, off-by-one at a boundary, wrong field. Run the tests after each one, record whether they went red, then revert. Also search for the six patterns listed above.
  • Fix: report the surviving bugs and the matching test gaps. New tests must use independently computed expected values with the arithmetic in a comment; each threshold gets one case on the boundary and one on each side; iterating tests assert that the number of items actually iterated is greater than zero. Plant the bugs again afterwards to confirm the tests go red.

B2. Tests and monitors written by the AI lock in its assumptions

When the code and the tests come out of the same line of reasoning, the tests prove nothing. Several times in this project, the test or the monitor itself treated a bug as correct behavior: a test asserted the API domain’s robots.txt was Disallow: / (see G1); a production probe asserted unknown paths return 200 (see G2); the health check used the date in the sitemap to decide whether a rebuild had happened (see G3); the _redirects count check counted the wrong way (see D1).

AI self-check B2

  • Applies to: every project, especially parts where the AI wrote the tests and the implementation together.
  • Check: for every assertion about external platform behavior (status codes, headers, limits, config values), ask where the expected value came from. If the only source is “that’s what the code does now”, flag it. For monitors and health checks, ask: if the current behavior is itself a bug, would this check notice?
  • Fix: take expected values from the platform’s documentation, an error message or a real measurement, and cite the source in a comment; monitors should assert what ought to happen, for example that unknown paths return 404.

C. Silent failures: the most expensive mistakes don’t raise errors

C1. The job reported success and wrote 2%

On the first full nightly run, the database hit its storage limit and every write in the fifth batch failed. The job returned ok: true, every Workflow step was green, and the summary said 2 of 104 succeeded. Job summaries now have a partial status: if the failure list isn’t empty or fewer than 90% succeeded, it doesn’t count as success.

AI self-check C1

  • Applies to: the project has batch jobs, scheduled jobs or queue consumers.
  • Check: what decides whether a job succeeded? Is it “no exception means success”? When some items fail (try/catch per item and carry on), can you tell from the return value and the logs?
  • Fix: return total, success count and a failure list in the job summary; define a threshold for partial success and mark the status partial or failed below it; make sure the scheduler and the alerts understand that status.

C2. Failures got hidden

On the night login launched (see F1), JWT verification failures were only logged at debug, which production doesn’t output; when the token list request failed, the page showed “No tokens yet”, and that empty state fooled even the AI once. Another time, Drizzle printed every parameter along with a failed statement; Cloud Logging truncated it, and the last line with the actual error (database full) was cut off. I only saw it after reproducing in psql.

AI self-check C2

  • Applies to: every project.
  • Check: search catch blocks for places that swallow errors, log only at debug, or render a failure as an empty state (“no data”, “list is empty”). Check whether error logs include long parameters or payloads that would get the root cause truncated.
  • Fix: server-side problems (downstream unavailable, misconfiguration, invalid secrets) log at warn or above with the key context; the caller’s own errors can stay at debug. The frontend distinguishes “failed to load” from “actually empty”. Error logs keep only the first line of the root cause and the identifiers needed, with a length cap.

C3. Health checks have to run outside the thing they check

Later I wrote a 528-line check.sh that checks, one by one, the edge entry point, the freshness of every data bundle in the cache, the nightly job history, Cloud Run logs, database migration history and usage, printing OK, WARN or FAIL per line. A companion skill has the AI run it and explain only the lines that aren’t OK, with a next step. The first run found three things: one group of cache entries had only 1/24 of its entries present, four jobs hadn’t run in eight days, and one monitor was hitting the origin every 20 seconds. It started as a scheduled task on my desktop and succeeded exactly once; later it moved to the cloud and runs four times a day, alerting through two channels. The bash script was kept, because it’s the only check that doesn’t run on the system it checks.

AI self-check C3

  • Applies to: projects with a production service.
  • Check: is there one command that answers “is production OK” in one go? Does it check liveness, or data freshness and what the jobs actually produced? Where does it run, and if the checked system is down, can it still run and alert?
  • Fix: write a read-only health-check script that prints OK, WARN or FAIL plus a reason per item, covering entry-point reachability, the latest timestamp of key data, the last result of each scheduled job, error log counts and quota usage. Run it on a schedule on separate infrastructure, and keep a way to run it locally by hand.

C4. The first diagnosis is usually plausible, and wrong

IndexNow kept returning 429; the first diagnosis was “too many URLs”, and after cutting the volume by 90% it was still 429 (see G4). GA4 session counts didn’t add up; the first diagnosis was “small-sample thresholding”, and it was actually double counting (see G7). After fixing the sitemap dates, the first measurement said “12 URLs a night”, but the local environment was pointing at a port where nothing was running. Each time, what overturned the first diagnosis was a small experiment that could tell the causes apart: send just 1 URL, compare page views, measure again in production.

AI self-check C4

  • Applies to: debugging any production problem. This one constrains how the AI itself works.
  • How: before fixing anything, write down at least two candidate causes, design the smallest experiment that tells them apart (a different source, a single item, a different environment), run it, and fix according to the result. Verify the fix with the same experiment.

D. Cloudflare

D1. The dynamic-rule limit in _redirects counts from the first placeholder (platform)

_redirects on Workers static assets allows 2,000 static rules and 100 dynamic rules (ones with :name or *). Rule 7 in the hand-written file had a placeholder, and the build appended 1,554 static 301s after it. The build script counted for itself: 13 dynamic rules, fine. Then one day the deploy failed: Line 147: Maximum number of dynamic _redirects rules limit of 100 exceeded. It turned out that from the first placeholder on, no rule can use the hash lookup any more, and Cloudflare counts all of them as dynamic.

Also, _redirects on Workers static assets runs whether or not a file matched, so the /* /index.html 200 that was common in the Pages days hides every prerendered page.

AI self-check D1

  • Applies to: deployed on Cloudflare Workers static assets or Pages, with a _redirects file.
  • Check: find the first rule containing : or * and count the rules from there to the end of the file (including ones appended at build time); the total number of static rules; whether there is a catch-all like /* /index.html 200.
  • Fix: move all placeholder rules to the end of the file, ideally generated by the build; in the build, count “rules after the first placeholder” and fail above 100; remove the catch-all and use not_found_handling instead (see G2).

D2. Workflow’s built-in schedules can quietly not fire (platform)

To get around the free plan’s limit of 5 crons per account, the nightly jobs moved to Cloudflare Workflow schedules. For five days in a row, four groups of jobs never created an instance: no errors, not one request reached the backend, and 71 cached data bundles went more than 26 hours without an update. I switched back to plain Cron Triggers that only create Workflow instances. The instance id is built from “schedule + the minute it was due”, so duplicate deliveries still produce one instance, and three failed creations send a Telegram message.

AI self-check D2

  • Applies to: Cloudflare Workflows or any platform’s built-in scheduling.
  • Check: is there detection for “should have run but didn’t”? Does it rely on platform scheduling or your own cron entry point? Would a duplicate trigger start two copies?
  • Fix: use the plainest cron entry point and have it only create instances; build the instance id from the schedule name and the due time so it’s idempotent; retry and alert on failed creation; add a separate check comparing “last run time” to the expected interval.

D3. Workflows replay all of run() when resuming (default)

The timer was outside step.do, so it measured only the last replay: a job that ran about 21 minutes in five batches was recorded as 126 milliseconds, and the public status page showed 0 seconds for every job. The docs do say code outside a step runs repeatedly; the AI still wrote the timing like ordinary imperative code.

AI self-check D3

  • Applies to: Cloudflare Workflows, Temporal, Durable Functions or any other durable execution framework that replays.
  • Check: whether code outside steps does anything with side effects or nondeterminism: timing, counting, sending notifications, writing to external storage, generating random numbers, reading the current time.
  • Fix: move it inside a step or return it as a step’s result; measure time inside each step and add it up.

D4. Slow origin: 524 after 125 seconds, while the origin actually finished (platform)

The nightly job is a Worker calling an endpoint on Cloud Run, and the job takes a bit over three minutes. Cloud Run’s log shows a 200 after 192 seconds; the Worker got a 524 at 125.4 seconds, and the cache warm-up after it was skipped. Everyone knows about proxy timeouts, but it’s easy to miss that they also apply to subrequests a Worker makes to an origin. The fix: send 200 headers as soon as the job starts, write a newline every 10 seconds as a heartbeat, and write the JSON at the end (JSON allows leading whitespace).

AI self-check D4

  • Applies to: requests that may take longer than 100 seconds behind a Worker, a CDN or any reverse proxy.
  • Check: list the slowest endpoints and how long they actually take; the caller’s timeout; whether the caller treats “origin still running” as a failure after a timeout.
  • Fix: make long jobs stream with a heartbeat, or make them asynchronous: submission returns a job id and the caller polls for the result.

D5. “Partially successful” deploys (platform)

Going from 4 crons to 6 exceeded the free plan’s 5. wrangler printed Trigger configuration … was only partially updated: the script was live, the schedules weren’t registered. Everything looked deployed, and those two jobs were never scheduled again until the next deploy surfaced it. A related detail: if the config leaves out triggers.crons entirely, the old schedules in production stay as they are; to clear them you have to write an explicit empty array.

AI self-check D5

  • Applies to: Cloudflare Workers cron triggers, or any deployment config with account-level quotas.
  • Check: the number of crons in the current config against the plan’s limit; whether recent deploy logs contain anything like partially updated; whether schedules removed from the config are actually gone in production.
  • Fix: add a test asserting the cron count stays within the plan limit; have the deploy process treat “partially updated” as a failure.

D6. A shared egress IP runs into a third party’s per-IP rate limit (platform)

IndexNow submissions from the Worker got 429 every night, every time. A third-party API that rate-limits by IP sees Cloudflare’s shared egress IPs when called from a Worker, and other people had used up the quota long before. Details in G4.

AI self-check D6

  • Applies to: calling third-party APIs from environments with shared egress such as Workers, Vercel Edge or Lambda.
  • Check: whether those APIs rate-limit by source IP (read the docs, or look for 429s out of proportion to your request volume).
  • Fix: move those calls to a service with its own egress IP; record a 429 as partial success and alert, and don’t retry from the same egress.

D7. Bundling for the edge runtime resolved the browser build of a dependency (platform)

After the blog moved to rendering in a Worker at the edge, the bundler resolved a Markdown dependency to its browser build, which calls document.createElement at module top level, and the Worker crashed on startup. Changing resolve conditions globally broke the prerendering on the Node side, so the Worker got its own bundling config with the workerd and worker conditions. The local dev server doesn’t execute the Worker at all, so nothing showed up locally; later I added a dev command that runs under wrangler dev.

AI self-check D7

  • Applies to: the same code is bundled for two or more of browser, Node and an edge runtime (Workers, Deno).
  • Check: the export conditions (browser, worker, workerd, node) in each target’s bundling config; whether the output references document, window or Node built-ins; whether the local dev workflow actually executes the edge runtime code.
  • Fix: a separate bundling config per runtime; a startup smoke test (one request under wrangler dev or equivalent); a local dev command that runs in the real runtime.

D8. A few small ones

AI self-check D8

  • HEAD requests: do cache decisions and cache headers only recognize GET? Uptime probes and crawlers use HEAD a lot, and all of it would hit the origin. Treat HEAD like GET (use GET when going to the origin).
  • _headers: at most 100 rules; anything beyond is dropped silently. Count the rules and assert in the build.
  • Files without an extension (such as /.well-known/api-catalog) are served as application/octet-stream by default; declare the type in _headers.
  • A Worker’s plain-text variables are overwritten by vars on the next deploy; values typed into the dashboard by hand belong in Secrets.
  • pnpm --filter <package> deploy runs pnpm’s built-in deploy command and skips the script of the same name in package.json; write run deploy.
  • The Worker name must match the one that actually exists in the dashboard, or wrangler deploy creates a new Worker with no routes.

E. Cloud Run

E1. Once the response is sent, the CPU is gone (platform, later instruction)

With min-instances at 0 and CPU allocated only during requests, anything done after the response is sent can be frozen at any moment. An early piece of code that “returns first, then writes to the database in the background” showed up as Connection terminated unexpectedly from the database. The deployment doc kept saying “min 1 + CPU always allocated” long after the config changed; the doc wasn’t updated, and that stale doc misled several more debugging sessions. The same correction also found that the default 300-second request timeout cut off the nightly jobs; it’s now 1,800 seconds.

AI self-check E1

  • Applies to: Cloud Run, Lambda, Cloud Functions or any platform that allocates CPU per request.
  • Check: search for code still running after the response is sent: un-awaited Promises, setTimeout, event-triggered background writes, asynchronous log or telemetry uploads. Compare the platform’s request timeout with how long the longest job actually takes.
  • Fix: either await before returning or hand the work to a real queue or job service; set the request timeout to at least twice the longest job; state the CPU allocation mode in the deployment doc.

E2. A new revision cuts off a running batch (platform)

On the first full nightly run, the fifth batch was interrupted by two backend deploys. Workflow retried as designed, but that run was wasted. The instruction file now says: no merging PRs that deploy the backend between 01:00–01:45 and 03:00–04:30 UTC, and every PR description the agent writes repeats that window.

AI self-check E2

  • Applies to: there are scheduled batch jobs, and merging to main deploys automatically.
  • Check: the batch time windows; whether deploys interrupt in-flight requests; whether jobs can resume from where they stopped.
  • Fix: write the no-deploy windows into the instruction file and the PR template; make batches idempotent per batch and resumable.

E3. Cache versions should follow the API contract, not deploys (default)

After an API field was renamed, the new frontend read old data bundles from KV, and a panel on the home page was empty for about an hour. The fix at the time was to version cache entries with Cloud Run’s K_REVISION and treat anything written by an old revision as a miss. As a result, every deploy invalidated all ~680 data bundles, even for a one-line CSS change, and it took 34 hours to notice. And with browser caching at 5 minutes and edge caching at 1 hour, this approach didn’t really guarantee anything in the first place. Now it’s a manually maintained contract version, bumped only when the frontend can’t read the new data.

AI self-check E3

  • Applies to: a cache layer (CDN, KV, Redis) caches API responses, and frontend and backend deploy separately.
  • Check: whether the cache key or invalidation condition includes a deploy id, git sha or build time; whether the new frontend can read old cache entries after an incompatible API change.
  • Fix: introduce an explicit contract version, bumped only on incompatible changes and included in the cache key; compatible new fields don’t invalidate the cache; add a test that stops anyone from wiring the deploy id back in.

E4. /healthz is taken by Google’s frontend (platform)

That path returns 404 straight from Cloud Run’s edge; the request never reaches the container and nothing appears in the logs. Renamed to /health.

AI self-check E4

  • Applies to: services on Cloud Run or behind other Google frontends.
  • Check: whether health checks or probes use /healthz or another reserved path ending in z.
  • Fix: rename to something like /health and update the probe config to match.

E5. Custom domains, logging and build context (platform)

A Cloud Run custom domain can’t sit behind the Cloudflare proxy: managed certificates require a direct connection, Cloud Armor needs a load balancer, and Host rewriting on Cloudflare is an Enterprise feature. The final setup puts another Worker in front of the API that allowlists paths and adds a secret header when forwarding to *.run.app; the origin rejects every direct request without that header. Plain-text logs all get severity DEFAULT in Cloud Logging; output JSON with a severity field. When deploying straight from the repo, the build context is the directory containing the Dockerfile, so in a monorepo the Dockerfile has to sit at the root.

AI self-check E5

  • Applies to: deployed on Cloud Run.
  • Check: whether the origin can be reached directly, bypassing the CDN; whether production logs are structured JSON with a severity; whether the Dockerfile location and build context can see the monorepo’s root files.
  • Fix: have the CDN or edge layer add a header only it knows, and verify it at the origin; log JSON lines with severity; put the Dockerfile at the root of the build context.

F. Postgres and Neon

F1. Launch night for login: four PRs fixing four silent failures (platform + default)

That night, over about four and a half hours, in order:

  1. The beta SDK’s types didn’t include the getJWTToken() the docs described. The AI decided “the types are missing it, the docs are right” and added a runtime typeof fn === "function" check instead of a type assertion. But the client is a Proxy that turns any unknown method name into an HTTP path, so typeof is always "function". The call became GET /get-jwt-token: 404.
  2. Switched to client.token(); still no usable token.
  3. Bypassed the SDK and requested /token directly. The JWT is at the top level of the response, not inside {data, error}.
  4. The browser finally sent a JWT, and the API returned 401 for everyone. The cause was new URL("/.well-known/jwks.json", base): base has a path (…/neondb/auth), and a relative path starting with / drops it. Along the way it also turned out the JWT’s iss is the bare domain without the path.

Every step took a few extra rounds because the failures were hidden (see C2). Afterwards a test was added that asserts the JWKS URL is not equal to what new URL() produces, with a comment saying the line looks far too much like a harmless simplification.

AI self-check F1

  • Applies to: every project, especially ones integrating an auth service or a beta SDK.
  • Check: search for new URL("/…", base) with a leading slash and confirm the base has no path or that dropping it is intended; search for typeof … === "function" checks on SDK methods and confirm the object isn’t a Proxy; check the log level and user-visible behavior of every failure branch in the auth flow.
  • Fix: handle slashes explicitly when joining URLs with a path, and add a test with a base that has a path; on Proxy-style clients only call methods that exist in the types; log the server-side part of auth failures at warn.

F2. Reading the database across clouds: every row is egress (default + platform)

The database is in AWS us-east-1 and the app is on GCP, so every row read is billed as public internet egress. To decide whether data was fresh, the first cache layer did a select * over two to four thousand rows and looked only at the first and last dates. The free plan’s 5 GB a month was gone in a week. It was changed to query max, min and count first and then fetch incrementally. After upgrading to a paid plan it was still 3 GB a day. There was no hourly usage API and pg_stat_statements wasn’t installed, so in the end I summed bytesRead on the database connection’s socket and added a db_kb field to every access log line. Once the numbers were there it was obvious: in one day the API read about 3 GB from the database and sent out only about 140 MB; one data bundle read 555 KB and sent 268 KB, because the same ~110 KB snapshot was read three times; one join query read 309 KB and sent 1 KB.

AI self-check F2

  • Applies to: managed databases, especially when the database and the app aren’t in the same cloud or region.
  • Check: which cloud and region the database and the app are each in; search for select * and for queries that read whole rows just to get a few fields or make a decision; whether the same data is read more than once in one request; whether there is any way to measure bytes read from the database per request or per endpoint.
  • Fix: turn decision queries into aggregates; select only the columns needed; cache already-read data within a request; add per-request database bytes read to the access log (summing from the driver’s socket works), and measure before optimizing.

F3. The database was full and the job reported success (default + platform)

On the first full night the database hit its 512 MB storage limit, and the job reported success anyway (see C1). 63 MB of the storage was an index that exactly duplicated the primary key, there since the scaffolding on day one.

AI self-check F3

  • Applies to: Postgres.
  • Check: query pg_indexes for redundant indexes whose columns are identical to the primary key or another index; the size of each table and index; how far you are from the storage limit.
  • Fix: drop redundant indexes (after confirming no constraint depends on them); add a cleanup policy for the fastest-growing tables; add storage usage to the health check.

F4. Don’t use a driver written for the edge in a long-running process (default)

The WebSocket pool from @neondatabase/serverless failed every handshake on Cloud Run, and every login returned 500. That driver is meant for edge runtimes; a long-running Node process is fine with an ordinary TCP pool. The pool also needs keepAlive, an idle timeout and pool.on("error"): the pooler kills idle connections, and one unhandled idle-client error can bring down the whole Node process.

AI self-check F4

  • Applies to: Node services connecting to Postgres.
  • Check: which driver is used and whether it was designed for edge or serverless environments; whether the pool sets keepAlive and an idle timeout; whether the pool’s error event is handled; whether the database’s pooled connection string is used.
  • Fix: use a TCP driver in long-running processes; add the settings above; create only one pool for the whole project.

F5. Telemetry writes kept the database from scaling to zero (platform)

The frontend wrote Web Vitals to our own database, and during the day every write woke up a compute node that had scaled to zero, adding about 1 compute hour a day. Vitals now go to GA4.

AI self-check F5

  • Applies to: the database bills by compute time and can scale to zero.
  • Check: which writes are triggered by anonymous visitors (analytics events, telemetry, counters), and how often.
  • Fix: move those writes to a dedicated analytics service or log pipeline, or batch them and write less often.

F6. Migrations, a shared database and unique constraints (platform + default)

ORM-generated migrations don’t always run: one generated migration put a new primary key constraint before the new column, and it failed. Renames need an interactive terminal, which the agent’s shell doesn’t have, so they had to be hand-written as ALTER … RENAME. With only one database and no test database, a temporary test script cleaned up with DELETE … WHERE account_id='me' AND date IN (…) and nearly deleted real data.

Another point many people get wrong (including the rules we first wrote): when a unique constraint includes a nullable column, Postgres treats NULLs as distinct from each other, so two rows of (a, NULL) don’t conflict, ON CONFLICT never fires, and the upsert degrades into duplicate inserts. Primary keys aren’t affected, since primary key columns are NOT NULL anyway.

AI self-check F6

  • Applies to: Postgres with any migration tool.
  • Check: whether anyone has read the recently generated migration SQL; whether it contains DROP, RENAME or type changes; whether tests and temporary scripts connect to the production database, and whether their cleanup conditions can touch real data; whether unique constraints used for ON CONFLICT include nullable columns.
  • Fix: generate, read, then run migrations; have temporary scripts write with a one-off marker (such as smoke-test-<random>), or run their assertions inside a transaction and roll back; for unique constraints with nullable columns use NULLS NOT DISTINCT (Postgres 15+) or give the column a non-null sentinel default.

F7. The prices the model knows are out of date (default)

When we were discussing whether to upgrade, the AI quoted Neon’s paid plan as $19 a month; Neon had moved to usage-based pricing back in late 2025. Prices, quotas and plan names in training data may all be out of date.

AI self-check F7

  • Applies to: every decision involving a third-party service’s quotas, prices or limits.
  • How: cite the source and date whenever quoting a price, quota or limit; anything not checked against the official page on the spot is marked “unverified”.

G1. Disallow: / on the API domain turned a page into a soft 404 (platform + default + instruction)

The API isn’t a website, so the AI wrote Disallow: / in the API domain’s robots.txt, plus a test asserting it. But when Google’s renderer renders a page, it obeys the robots.txt of every domain the page requests. The data request was blocked, the page showed “failed to load”, and the same day Search Console reported a page as a soft 404. It now allows the public API paths and keeps X-Robots-Tag: noindex: being crawlable and being indexable are two different things.

AI self-check G1

  • Applies to: pages that fetch data in the browser from another domain (API, CDN, image service), where search indexing matters.
  • Check: whether those domains’ robots.txt disallows paths the pages need to request.
  • Fix: allow crawling of the paths needed to render the pages; control indexing of responses you don’t want indexed with X-Robots-Tag: noindex.

G2. SPA fallback: every unknown path returns the home page with status 200 (default + instruction)

When moving to Workers, the AI chose not_found_handling: "single-page-application" so deep links would work. That’s the natural choice for an SPA, but this site has about 4,500 prerendered pages and gets most of its traffic from search. So misspelled paths, old uppercase URLs and client-side routes like /sign-in were all copies of the home page with canonical: / and status 200. The health check and the production probe both asserted “unknown paths return 200”, so the monitoring treated the bug as normal. It wasn’t until the day before I wrote this that it turned up in Search Console under “Crawled – currently not indexed”: a /sign-up that no longer existed had 39 impressions.

AI self-check G2

  • Applies to: single-page apps with pages meant to be indexed.
  • Check: request a path that doesn’t exist and look at the status code and the canonical returned; the status code and robots directives of client-side routes (login, account pages) when visited directly; whether monitoring asserts that unknown paths return 200.
  • Fix: return a real 404 for unknown paths (it can be a noindex page that boots the app); have client-side routes each serve a 200 shell with noindex; make monitoring assert 404.

G3. The sitemap’s lastmod was wrong four times (default + instruction)

  1. It started as the build time, then changed to each page’s modification date, but kept a fallback branch to the build time.
  2. About 200 new pages without a modification date were added, and the build time quietly came back.
  3. After a redesign, the AI moved lastmod forward by hand to “get crawlers to recrawl”. Google explicitly advises against this.
  4. It switched to the data bundle’s asOf, which is the nightly job’s own timestamp, so 1,042 of 1,247 URLs claimed to have changed every day.

Google ignores a lastmod that is always “today”. A page’s date is now the date of the newest dated item on the page. The health check used to decide “was there a rebuild last night” from the newest date in the sitemap; once the dates were honest, any day without content changes raised a false alarm, so a separate /build.json was added.

AI self-check G3

  • Applies to: sites with a sitemap.
  • Check: generate the sitemap and compute the share of URLs whose lastmod is today; trace where lastmod comes from: the content’s modification time, or a build, job or deploy time; whether there’s a fallback to the current time; whether another system relies on lastmod to tell whether a build happened.
  • Fix: take lastmod from the last date the page’s actual content changed; omit it when unsure, never fill in the current time; tell whether a build happened from a separate build-info file.

G4. IndexNow kept returning 429, and the first diagnosis was wrong (platform)

The same lastmod fed IndexNow, so every night 1,127 URLs were submitted, all 429. The AI decided “the volume itself is unreasonable” and fixed lastmod; the volume dropped to around a hundred, still 429. The experiment that actually separated the causes was one line: submit 1 URL with curl from my machine, get 202. IndexNow rate-limits by source IP, and the Worker’s outbound requests go out through shared IPs. Now the edge collects URLs and Cloud Run sends them. Separately, Google has no general URL submission API (the Indexing API only accepts job postings and livestreams), so the only signal you can give Google is lastmod.

AI self-check G4

  • Applies to: IndexNow or any other submission API.
  • Check: which egress the submissions go out through; recent response codes; whether the submitted URLs actually changed (see G3).
  • Fix: submit only URLs whose content actually changed; submit from a dedicated egress IP; don’t retry on 429, log it and alert.

G5. canonical is only a hint, and Google ignores it when the content differs (platform)

To “avoid diluting the main page”, a set of tab pages all had their canonical pointing at the main page and were left out of the sitemap. Search Console’s response was “Google chose a different canonical”. Now each tab page has its own canonical, is prerendered with its content, is in the sitemap, and is dated by its own content.

AI self-check G5

  • Applies to: several URLs pointing to the same canonical.
  • Check: whether the main content of those URLs is really the same. If not, the canonical will be ignored.
  • Fix: only true duplicates (print versions, embed versions, different parameter order) point to a canonical; pages with different content get their own canonical or are merged into one page.

G6. The HTML crawlers get has to contain the content (default + instruction)

The prerendering approach changed back and forth four times in under a month: first prerendering with data injected at build time; the next day it was reverted to pure client-side fetching, with the PR saying “data pages are no longer an SEO surface” and no reason given; then data pages got empty shells; then it turned out that 1,190 of about 1,200 URLs in the sitemap had no content in their HTML, and it went back to prerendering with data. Along the way one design doc listed “the build must not depend on the API” as a design constraint, a constraint created by that very revert a few days earlier. A related problem: the page served prerendered HTML first, but replaced it with an error state when the client request failed, and since Google indexes the rendered DOM, a single API timeout could turn into a soft 404.

AI self-check G6

  • Applies to: pages meant to be indexed whose main content depends on a client-side request.
  • Check: curl 20 random URLs from the sitemap and see whether the HTML contains the page’s main content; whether existing prerendered content gets replaced when a client request fails; whether constraints in design docs have a source.
  • Fix: render the main content on the server or at build time for pages that should be indexed; when a client request fails, keep the existing content instead of replacing it with an error state.

G7. Summing GA4 sessions across pages (default)

The first version added up sessions per page and got 636; the per-channel total was 258. The AI attributed the gap to GA4’s “small-sample thresholding” in a code comment. But page views matched exactly on both sides, which rules out thresholding: a session that visited several pages was counted once per page. The same day also ran into several console hurdles: the service account was already a viewer, but the Analytics Data API still had to be enabled separately in the project; and the API wants the numeric property id, not the measurement id starting with G-.

AI self-check G7

  • Applies to: pulling data from GA4 or another analytics API for reports.
  • Check: whether non-additive metrics such as sessions or users are summed across pages or other dimensions; which kind of id the data API is given; whether the fetch window covers the period in which data is still revised (about 2 days for GA4, about 3 for Search Console).
  • Fix: query non-additive metrics on their own dimension; use the numeric property id; refetch over the window and upsert.

G8. A few small ones

AI self-check G8

  • The service account for the Search Console API has to be added as a user on the property.
  • Text in FAQ and other structured data must match the visible text on the page word for word.
  • og:image and article structured-data images should be PNG, JPG or WebP, not SVG, and at least 1,200 pixels wide.
  • In robots.txt a crawler matches only the most specific group and inherits nothing from *. When writing a group for a specific crawler, repeat the general Disallow lines.
  • Markdown or plain-text copies generated for a page should point back to it with a Link response header (rel="canonical").
  • hreflang must be reciprocal and include a self-referencing entry.
  • The same path with and without a trailing slash must not both return 200.
  • The brand suffix on titles should be added by one function, not written by hand on some pages and added automatically on others.

Looking back

Laid out side by side, the most expensive of these forty-odd mistakes were almost all silent failures: a job that showed success while writing 2%, a deploy that was green while the schedules weren’t registered, 401s logged only at debug, a page saying “no data” when the request had failed. Most of them took one or two PRs to fix. The time went into noticing them.

A week before writing this I did a product review, and its conclusion is the first paragraph of the roadmap: the engineering and SEO foundations are solid, but recent effort has gone mostly into infrastructure, and the few things that decide whether this project survives have been underinvested. Most of this post is the itemized version of “effort has gone mostly into infrastructure”.

Finally, a set of rules you can put in your own instruction file. They match the self-checks above: a self-check is done once, while a rule in the file is read in every turn. Take the ones that apply rather than pasting the whole block; as A1 says, the longer the instruction file, the more every turn costs.

Rules to put in CLAUDE.md / AGENTS.md

  • The instruction file holds only constraints needed in every turn that can’t be read from the code; progress goes in issues; when docs and code disagree, the code wins and the disagreement gets pointed out.
  • Conventions a machine can check become lint rules or tests; rules are error-only and ship with zero violations.
  • Expected values in tests must be derived independently, with the arithmetic or source in a comment; before merging, deliberately break the code under test and confirm the test fails.
  • Scheduled jobs and batches must distinguish success, partial success and failure; server-side errors log at warn or above; the frontend distinguishes “failed to load” from “empty”.
  • Don’t keep working after the response is sent; long jobs use a streaming heartbeat or an async job.
  • Cache versions follow the API contract; deploy ids, git shas and build times never go into cache keys.
  • Destructive operations (deleting data, migrations, permission changes, mass sends) are prepared, not executed, and handed to a person to confirm.
  • When debugging, propose at least two hypotheses and test them with the smallest experiment that tells them apart before changing anything.
  • Cite the source and date when quoting prices, quotas or limits; mark anything unverified as such.
  • Paths with no content return a real 404; sitemap dates come only from content changes.