My vibe coding setup

The rules, guardrails, and verification habits that make building software with an agent reliable rather than lucky.

Updated 6 Sept 2026 · Current as of Claude Opus 5, Claude Code

Rules live next to the code

  • One rules file per repo, checked in
  • Conventions, commands, and the traps that bite twice
  • Cite real files and line numbers, not principles

If a rule matters, it lives beside the code it governs.

An agent that has to rediscover your conventions every session will get them wrong roughly as often as a new contractor would. A short rules file in the repository — naming conventions, the actual build and test commands, the gotchas that have already caused a bad afternoon — removes most of that. Point at concrete files rather than stating principles; specifics survive, adjectives do not.

Load context on demand

  • A short always-on router file
  • Deeper modules loaded only when relevant
  • Each names the trigger for loading it

Everything in context is paid for on every single turn.

Loading every document into every session is expensive and, past a point, counterproductive — the signal you need competes with everything you did not. Keep the always-loaded layer small and make it a routing table: here is where routing lives, here is where styling lives, load that one when you touch this.

Guardrails the model cannot skip

  • Hooks run in the harness, not in the model
  • Deterministic checks beat polite instructions
  • Enforce the rules you keep having to repeat

An instruction is a request. A hook is a fact.

Anything phrased as "always remember to…" will eventually be forgotten, because it competes with everything else in the prompt. Anything enforced by the tool that runs before or after an action simply happens. Move your non-negotiables — formatting, secret scanning, write protection on system paths — out of the prompt and into the harness.

Evidence or it is not done

  • No completion claim without tool output
  • Tests, diffs, screenshots, build logs
  • "Should work" is not a status

Name the check that would prove you wrong.

The most expensive failure mode in agentic development is not a bug, it is a confident report of success that nobody verified. Require the evidence to be part of the deliverable: the command that was run and what it printed. A claim with no possible falsifier is not a claim.

When you cannot verify, say so

  • An unavailable checker means defer, not assume
  • State plainly what was checked and how
  • An admitted gap beats a confident guess

The gap you name is the one that does not bite you.

Sometimes the verifier genuinely is not available — no browser in the environment, no access to the staging system. The correct response is to finish everything else, then say exactly which behaviour remains unverified and by what method it should be checked. Reporting an unverified claim as verified is how a defect reaches a client demo.

Plan, then execute in isolation

  • Agree what done looks like before code
  • Work on a branch or a worktree, never the default
  • Fresh context per task, reviewed before the next

Review between steps, not only at the end.

Long agentic runs drift. Splitting work into tasks that each end in a review gate keeps the drift bounded and makes it obvious which step introduced a problem. An isolated workspace means an abandoned experiment costs a deletion rather than an untangling, and a plan agreed up front means the review has something to judge against.

What would make your
work a little easier?

A recurring task, a half-formed idea, a system that could work better. That’s a good place to start.

Let’s talk it through