What Precedent is

AI coding agents write code fast, and they copy a lot. Often the function an agent needs already exists somewhere in the project, but the agent never saw it, so it writes a new one. After a few weeks the codebase is full of near-identical functions that slowly drift apart.

Precedent is a small tool I am building to catch this while the agent is still working. It is at the proof-of-concept stage and handles only Rust for now.

The idea:

  • When a session starts, Precedent records what the code looks like.
  • When the agent finishes its turn, Precedent compares the new code with everything that was already there.
  • If the agent wrote something that already exists, Precedent tells the agent where it is, and the agent decides what to do about it.

The goals:

  • Fast. A check takes milliseconds, so it can run after every turn.
  • Quiet. No warning when the agent only moves code, edits code that was already duplicated, or changes nothing.
  • Local. Finding the duplicates uses no AI model and sends no code anywhere.
  • The agent decides. Precedent only points at the existing code. The agent's own model makes the call.

How it plugs into oh-my-pi

oh-my-pi is the coding-agent harness I use. Precedent connects to it through a small extension that only starts the Precedent program and passes its answer back. All the logic lives in Precedent itself.

 session starts ───► Precedent records the code as it is now
                     and watches files in the background
        │
        ▼
 agent works: reads, edits, writes files
        │
        ▼
 agent says it is done ───► Precedent checks: did the agent add code
                            that already existed?
                                 │                      │
                              no │                      │ yes
                                 ▼                      ▼
                         turn ends as usual    agent gets one more turn,
                                               with a note naming the
                                               existing code
                                                        │
                                                        ▼
                                               agent reuses it, merges it,
                                               or leaves it; turn ends

The note the agent gets looks like this:

precedent: 1 duplicate introduced this session
1. src/shipping.rs:103-119 `shipping::validate_order`
   ≈ src/orders.rs:42-58 `orders::validate_order` (pub, crate bench_repo)
   overlap 100% of 187 tokens · type-1 (exact copy)
   → call `crate::orders::validate_order`, or extract a shared helper

Precedent speaks up at most once each time the agent stops. If anything goes wrong (the program is missing, too slow, or its index is broken) the session is designed to carry on as if Precedent were not there. The extension is being built right now, by the same automated pipeline that built the rest of Precedent.

Why this matters: review rounds that never end

Precedent is itself written by an automated pipeline. One model writes a step, another model reviews it, the first one fixes what the reviewer found, and the reviewer looks again. Many steps did not settle. In 7 of the first 17 steps the reviewer still had findings after three rounds of fixes, and I had to step in by hand.

When I went through the review findings, a large share were about duplication: "this is a second copy of X", "this test repeats the same setup", "this restates a rule defined elsewhere". A simple keyword search finds 96 such findings out of 525, and 49 out of the 137 the reviewer marked as serious. That is more than a third.

If the model writing the code and the model reviewing it have different ideas of what counts as duplicate code, they can go round in circles. So I measured how much two models agree.

The test

I ran Precedent on two of my own Rust projects, both written mostly by AI agents: agentblog (about 200,000 lines) and nucleus-mcp (about 60,000 lines). From each I took a random sample of 60 pairs of similar code that Precedent reported.

Two models labelled every pair without seeing each other's answers: Claude Opus 5.5 and Claude Sonnet 5.5, both with high thinking effort. Each could read the whole project. For every pair they chose one label:

  • actionable: a reviewer would ask to merge the two into one function
  • boilerplate: ordinary code that is normally written out like this
  • coincidental: looks similar, does different things
  • intentional: a deliberate copy, such as test data

To measure agreement I used Cohen's kappa. It is a score that ignores the agreement you would get by chance: 1 means the two always agree, 0 means they agree no more than chance would. I had set 0.6 as the bar to pass.

Result: they don't agree

ProjectAgreement (kappa)Pairs they disagreed on
agentblog0.719 of 60
nucleus-mcp0.4422 of 60

On nucleus-mcp, 18 of the 22 disagreements were about the same thing: one 13-field block of code that is repeated in the parser of every programming language the project supports, nearly 80 times in total.

Opus: "the same 13 fields … repeated about 75 times across the parsers, and DartParser::make_symbol already shows the shared helper this repository would want." Its label: actionable.

Sonnet: "Both are the same struct literal … Each parser fills it with its own node kinds and doc extractor." Its label: boilerplate.

Both answers are reasonable. That is the point. Whether repeated code should be merged is a judgement call, and two models from the same company make it differently.

My first explanation was wrong

The labelling instructions gave "match arms" as an example of boilerplate, and the repeated block sits inside match arms. I assumed Sonnet had latched onto that phrase. I removed it and ran both models again. Nothing changed.

So I ran the whole labelling several more times: same pairs, same instructions, once through Anthropic's own API and once through OpenRouter.

Runagentblognucleus-mcpOpus calls the block actionable
original run0.710.4418 of 18
same instructions again, Anthropic–0.750 of 18
same instructions again, OpenRouter0.550.3918 of 18
"match arms" removed, Anthropic–0.686 of 18
"match arms" removed, OpenRouter0.560.4514 of 18

What this shows:

  • Sonnet is consistent. It called the block boilerplate in every run, 90 labels out of 90.
  • Opus changes its mind between runs. The models labelled 10 pairs per session. Within a session, Opus gives every copy of the block the same label. Across 15 sessions it said actionable 9 times. It behaves like a coin flip per session, and it leans towards actionable when it goes looking and finds the existing helper.
  • The agreement score itself is unstable. With identical instructions, agentblog went from 0.71 (pass) to 0.55 (fail).
  • The provider made no visible difference. Anthropic and OpenRouter runs varied as much among themselves as between each other.
  • Part of the problem was my sample. 28 of the 60 nucleus-mcp pairs came from one group of similar code, so a single disputed pattern decided the score. Counting one pair per group, agreement rises to about 0.89.

Small decision models

I also tried two fast models that return a probability instead of text: TypeSafe's Jev (through OpenRouter), and jev-open, my open-source take on the same idea, running a 4-billion-parameter Qwen model on my own machine.

  • As labellers they were weak. Jev agreed with the large models much less than the large models agreed with each other (kappa 0.22 on agentblog). jev-open, asked to pick one of the four labels, picked "actionable" 597 times out of 600.
  • As rankers they were good. Asked a plain yes-or-no question ("would a reviewer ask to merge these?"), they gave higher scores to the pairs the large models called actionable. On a ranking score where 0.5 is guessing and 1.0 is perfect, Jev reached 0.86 and 1.00 on the two projects, jev-open 0.79 and 0.93.
  • They are cheap and fast. Jev cost about $0.00003 per decision. jev-open answers in about a tenth of a second. A label from a large model cost about $0.016 and took minutes per batch.
  • With a cut-off set from labelled examples, jev-open could drop about a third of the findings without losing any pair that most of the large-model runs called actionable.

What I changed

In Precedent, the agent's own model already makes the final decision. A false alarm costs it a few hundred words of reading. A missed duplicate is gone for good. So:

  • The labelling models no longer have to agree. Their agreement is recorded, not used as a pass or fail bar.
  • A pair counts as worth showing if any model run called it actionable. Every filter between Precedent and the agent must keep at least 95% of those pairs.
  • Similar findings are grouped: the agent sees one item, with the other places listed under it. In the samples above, the 120 pairs came from only 49 groups.
  • A small decision model can then filter what is left, using a cut-off set from labelled examples and never its first-choice answer.

Caveats

This is 120 pairs from two projects, labelled by two models from one vendor, with the models' labels as the reference. No human has labelled the disputed block yet.