Tag: Code review

  • What we hand to AI, and what we keep

    What we hand to AI, and what we keep

    An agent will happily scaffold a checkout flow for a product nobody has decided to build. It will not object, and the code will look fine. That is the core problem with deciding where AI belongs: the output looks finished whether or not the thinking behind it was sound.

    At Artasce, we run AI-first operations. Agents are part of how we build web, mobile, brand and e-commerce products. We’ve also had to be specific about where they stop. Here is how we divide the work.

    The rule underneath the split

    Agents get work where the answer can be checked. People keep work where the answer is a judgment.

    A test either passes or it doesn’t. A refactor either preserves behavior or it doesn’t. Whether a screen feels right, whether a name will hold up, or whether a feature deserves to exist can’t be checked that way. Someone has to decide, and that someone is accountable for it.

    What agents handle well

    Scaffolding. Project structure, routing, config, component shells, API client boilerplate. This work is well-understood and tedious, and mistakes show up quickly. We point agents at our conventions and let them lay the foundation.

    Repetitive refactors. Renaming a prop across a codebase, migrating a deprecated API, moving hard-coded values onto design tokens. The task is mechanical and the scope is clear, so agents do it well and the diff is easy to review.

    First drafts. Documentation, release notes, alternate headline options, rough copy for a page. A draft gives a person something to react to. We treat it as raw material and never as a final.

    Test coverage. Agents are good at enumerating edge cases a tired developer skips, such as empty states, malformed input and boundary values. We still read the tests. A test that asserts the wrong thing is worse than no test, because it creates false confidence.

    What stays with people

    Product judgment. What to build, what to cut, and what to say no to. An agent optimizes for the request in front of it. A person weighs the request against the business, the user and the next six months.

    Visual taste. Spacing, weight, rhythm, restraint. Agents can apply a system, and our Crystal Design System gives them tokens and rules to apply. Deciding when a layout needs to break from the pattern is a human call.

    Naming. Products, features, components, even functions. A name is a small decision that every future reader inherits. Generated options can seed a conversation, but they don’t settle it.

    The final review. This one gets its own section.

    The review step between machine output and a client

    Nothing an agent produces reaches a client directly. Every piece of output has a named person responsible for it, and that person reviews it before it moves on.

    For code, the reviewer reads the diff, not the summary the agent wrote about the diff. They run the tests, then check that the tests test the right thing. They look for changes outside the requested scope, since agents sometimes tidy things nobody asked them to touch.

    For design and copy, the reviewer checks the work against the brand: the type scale, the palette, the voice. They also check facts. A confident sentence with an invented number in it is the most dangerous kind of draft, so every claim gets traced to a source or removed.

    The reviewer can approve, edit or reject. Rejecting is a normal outcome, and we don’t treat it as a failure of the process. It is the process working.

    Responsibility is the point of this step. If the person signing off can’t explain why the output is right, it isn’t ready.

    A checklist for your own workflow

    Run a task through these questions before you hand it to an agent:

    • Can the result be verified? If a test, a type check or a clear spec can confirm it, it’s a good candidate.
    • Is the scope bounded? Tasks with clear edges suit agents. Open-ended ones need a person to set direction first.
    • Is it reversible? Prefer work you can review in a diff and roll back. Be slower with anything that touches production data or money.
    • Does it need taste or context? If the right answer depends on knowing your customer, your brand or your history, keep it with a person.
    • Who owns the result? Name one human before the work starts. If you can’t, don’t delegate it.
    • What does review look like? Decide in advance what the reviewer will read, run and check. Skimming isn’t review.
    • What happens if it’s wrong and nobody notices? The higher the cost, the more of the work should stay human.

    If a task passes most of these, hand it over. If it fails the first or the fourth, keep it.

    The line will move

    Tools improve, and this division will shift with them. The questions above should hold up better than any specific list of tasks, because they ask about the work rather than the tool.

    Quality has never come from a single step. It comes from a person who cares about the result and has the authority to say it isn’t good enough yet. We use agents to give that person better material and more time to decide.

    If you’re working out where AI fits in your own product or team, we’re happy to talk it through.