All posts
Blog

Your Coding Session Is Training Data

Claude Code, Codex, Cursor and Copilot now learn from user sessions. Here is what one session looks like as data, three ways a model can learn from it, and why the two lines you fixed by hand are the most valuable and most dangerous part.

Deep Chokshi·September 13, 2026·11 min read·
Coding AgentsPost-TrainingTraining DataSFTRL

In the last twelve months, every major coding tool switched on learning from its users. Anthropic trains on Claude Code sessions from Free, Pro and Max accounts when you opt in. OpenAI applies the "improve the model for everyone" setting to Codex tasks on ChatGPT plans. GitHub started training on Copilot interaction data by default on April 24, 2026. Cursor retrains Composer on live sessions and ships a new checkpoint every five hours.

So two questions are worth asking. What does one of your sessions actually look like as data? And what can a model learn from it?

This post answers both with one small session. An agent adds rate limiting to an API. It gets most of it right. You push back once. It fixes that. Then you change two lines yourself and commit. That last step is the interesting one. For training, it is the most valuable moment in the whole session. It is also the moment the recorder is most likely to miss.

A note on what is real here. The session is written by hand so every line can be inspected. The log shapes are the real ones from the Claude Code and Codex session files on my machine. Every claim about what a company collects links to that company's own words. No model was trained for this post.

One session, start to finish

The repo is a small FastAPI service. It has a /search route, an auth module whose current_user returns None for anonymous requests, and a net module with a client_ip() helper that knows the service sits behind a load balancer. Keep that last file in mind. The agent never opens it.

You type:

Add rate limiting to the /search endpoint: 60 requests per minute per user.
Return 429 when the limit is hit. Add tests.

Turn 1. The agent reads app/search.py, writes a sliding-window limiter in app/ratelimit.py, wires it into the route keyed by user.id, writes three tests, and runs them. Five pass. It reports back.

# app/search.py after turn 1
ratelimit.check(f"user:{user.id}")
return {"results": query_index(q, limit=20)}

You push back. Anonymous users have no id. current_user returns None for them, so the first anonymous request would crash with AttributeError.

Turn 2. The agent agrees, keys anonymous callers by request.client.host, adds a test, and runs the suite. Six pass.

key = f"user:{user.id}" if user else f"ip:{request.client.host}"
ratelimit.check(key)

You fix two lines yourself. You know something the agent doesn't. The service runs behind one load balancer, so request.client.host is always the balancer's address. Every anonymous user would share one bucket, and the site would start returning 429 after sixty anonymous searches. There is already a helper for this. You open the editor and change it:

+from app.net import client_ip
 ...
-    key = f"user:{user.id}" if user else f"ip:{request.client.host}"
+    key = f"user:{user.id}" if user else f"ip:{client_ip(request)}"

You run the tests in your own terminal, they pass, and you commit. Explaining the load balancer to the agent and waiting for a third turn would have taken longer than the twenty seconds this took. Sessions end like this all the time.

EXHIBIT 01One session, every event.
Prompt
Agent turn 1
Pushback
Agent turn 2
Outside the log
1 / 25 · 22 in the log
youYour promptevent 1 · type "user" · blocks: text

Add rate limiting to the /search endpoint: 60 requests per minute per user. Return 429 when the limit is hit. Add tests.

An authored session in the shape of a Claude Code log. Each square is one event; use the arrow keys or the buttons to step through. Squares with a red dashed border happened in your editor and terminal; the harness has no record of them. Everything shown comes from session.json, which you can download below.

Step through the exhibit and look at the last three rows. They have no log record. The commit contains your import and your client_ip call, and nothing in the session file says where they came from.

What the recorder actually saw

Open ~/.claude/projects/ on a machine that has used Claude Code and you will find one JSONL file per session. Each line is one event. A user message. An assistant message whose content is a list of blocks: text, thinking, and tool calls with their exact arguments, including the old_string and new_string of every edit. Tool results come back as another event. Codex CLI keeps a similar rollout file under ~/.codex/sessions/, with function calls, their outputs, reasoning items, and periodic world_state snapshots.

That is a lot. It is close to everything the model said and did. But look for the moment you changed two lines and you will not find it. A terminal harness only knows what it did itself. Your edit becomes visible to it later, if at all: the next time it reads the file, or when an edit fails because the file changed since it was last read, or when someone diffs the commit against the agent's last write.

Cursor and Copilot are different because they are editors. They see keystrokes. Copilot's policy says it collects "outputs accepted or modified by you." Cursor's reward for Composer includes whether the agent's edits "persist in the codebase," which needs exactly this comparison. An editor can attribute a line to a person. A terminal harness has to reconstruct it.

EXHIBIT 02What each recorder can see.
What happenedClaude Codelocal session fileCodex CLIlocal rollout fileCursordisclosedCopilotdisclosed
Your prompt and follow-up messages••••
The agent's text replies••••
The agent's reasoningthinking blocks / reasoning items••??
Tool calls with exact argumentsold_string, new_string, command•••?
Tool results••??
Workspace snapshotsCodex writes periodic world_state records; Copilot keeps code around the cursor••?•
Your manual edit, with you as the authorCopilot names “outputs accepted or modified by you”••••
An explicit accept or reject••••
Whether the agent's edit survived to the commitCursor rewards edits that “persist in the codebase”•••?
Your pushback, marked as a correctionin a terminal log it is just the next message••••
Left two columns: what the session file on my own disk contains. That is a lower bound on what the harness could record and says nothing about what is uploaded. Right two columns: what each company has said it collects or rewards, linked in the text. A question mark means not public.

The point of the exhibit is that attribution is a property of the recorder, not of the session. The same twenty seconds of typing is a labelled event in one product and a silent diff in another.

Three ways to turn the session into a gradient

Take the log and flatten it into one long sequence of tokens: the system prompt and tool schemas the harness injects, your prompt, the agent's text and tool calls, the tool results, your pushback, and turn 2. Every token now has an author. That single fact is what all three methods below are built on, because a model should only be pushed toward tokens it is supposed to produce. Tool output is something it reads, not something it writes.

1. Supervised fine-tuning: copy the agent's turns

The plain version. Put a loss on the agent's tokens and none on anything else. The context still flows through the network, so the tool results still shape what the model learns to write. They just aren't the thing being predicted.

The real decision is which agent tokens get the loss. Turn 1 contained the unguarded user.id. Turn 2 is the recovery. Three reasonable choices:

  • Final turn only. Turn 1 and your pushback stay in the context. Loss only on turn 2. This teaches the recovery without rewarding the mistake.
  • All agent turns. Loss on both. Simpler, and it teaches the mistake too.
  • Hindsight. Swap your fix into turn 2 and put the loss on that. Now the model is asked to write client_ip(request).

The third one is the trap. In this session, client_ip does not appear anywhere in the context the model saw. It lives in app/net.py, and the agent never read that file. The records at the bottom check for exactly this and flag it. Training on that target teaches the model to produce a name it had no way to know. If you want the fix as a demonstration, you have to add the observation that makes it possible, such as a reconstructed read of app/net.py, and label it as reconstructed.

2. Reinforcement learning: score the turns by what you did next

The plain version. Don't copy anything. Give each agent turn a number, then push the probability of that turn's tokens up or down in proportion to it.

Where does the number come from? From you, without you filling in a form. Two signals are already in the session:

  • Pushback. Your message after turn 1 was a correction. That is a negative signal for turn 1.
  • Persistence. Did the lines the agent wrote survive to the commit? Turn 1 wrote 45 lines and 44 survived. Turn 2 wrote 9 and 8 survived, because you replaced one.

Notice the two signals disagree about turn 1. Its code mostly survived, yet you pushed back on it. Any reward has to weigh those two, and the weights are a design choice, not a fact. Cursor's post reports the effect of each of its signals separately for that reason.

This is what Cursor describes for Composer, and it comes with one hard constraint. The turns must have been produced by the model you are updating. Score a turn, retrain, and the next session's turns come from the new model. That is why Cursor ships a checkpoint every five hours.

3. Preference learning: your version against the agent's

The plain version. Take the same context up to your pushback. The rejected answer is turn 2 as the agent wrote it. The chosen answer is turn 2 with your fix applied. Train the model to prefer chosen over rejected.

It is tempting because the pair falls out of the session for free. Be honest about what it is. Nobody generated the chosen answer. It is the agent's turn with two of your lines pasted in. And it has the same problem as hindsight SFT: the thing that makes it better, client_ip, is not in the context. The pair can teach the model to move away from request.client.host. It cannot teach it to find the helper it never saw.

EXHIBIT 03One session, three gradients.
spans with loss5 of 25
share of characters24%
agent-written characters77%

The recovery after pushback. Turn 1 stays in the context but gets no loss.

Turn 1 and your pushback still flow through attention into every prediction in turn 2. They are context, not targets.

Every number here is read from records.json, which is derived from session.json. Block widths are characters, not tokens; no tokenizer is involved. Hover a block for its label.

The math, on this session

Here is the same thing written down. Each line is read out in plain words underneath it.

1 · SFT: masked next-token lossℒSFT=−1M∑jmjlogπθ(zj|z<j),M=∑jmj

Add up how surprised the model was by each token it is supposed to imitate, only where the mask m is on, and average. z is the whole session as tokens, in order.

2 · what the gradient does at one position∂ℒ∂uj=mjM(pj−ezj)

At a supervised position, push up the token that was actually written and push down every other token, in proportion to how much probability it got. Where the mask is off, this is zero.

3 · RL: score each agent turn by what you did next∇θJ=∑tAt∑j∈t∇θlogπθ(zj|z<j),At=rt−b

For each agent turn, take the gradient of the log-probability of its tokens and scale it by how much better than baseline the human's reaction was. Tool results and your messages get no gradient of their own.

4 · DPO: your version against the agent'sℒDPO=−logσ(β[logπθ(yw|x)πref(yw|x)−logπθ(yl|x)πref(yl|x)])

Raise the chosen turn relative to the rejected one, in the same context x, measured against a frozen reference model so the policy cannot drift arbitrarily far.

5 · why corrections on the learner's states matterBC:J(π^)≤J(π*)+T2εDAgger:J(π^)≤J(π*)+uTε+O(1)

T is the number of steps in an episode and ε the per-step error. Copying only expert states lets error compound with T². Labeling the learner's own states, which is what your fix does, keeps it linear in T.

Two things people get wrong about the first equation. Turning off the loss on context tokens does not stop the model from learning from the context. Every supervised prediction depends on all earlier tokens through attention, and gradients flow back through them. What is turned off is only the job of predicting the context itself. And the normalization matters. Averaging over tokens means a 700-character file write counts far more than a one-line command. Averaging per turn first gives each action the same weight. Those are different objectives, and a paper should say which one it used.

Why your fix is worth more than a clean demo

Now the part that makes this data special rather than just convenient.

Imitation learning has a well-known failure. Train a policy only on states an expert visited, and it does fine as long as it stays near those states. The moment it makes a small mistake it lands somewhere the expert never showed it. It has no idea what to do there, so it makes another mistake, and the errors compound. Ross, Gordon and Bagnell showed in 2011 that the cost grows like T²ε for a horizon of T steps and a per-step error of ε.

Their fix, DAgger, is simple. Run the learner. Let the expert label the states the learner actually reached. Add those to the data and retrain. The bound drops to Tε.

Look at what you did in the session. The state after turn 2 was reached by the agent's own policy. Your two-line fix is the expert's label at that state. That is exactly the data DAgger asks for, and it is exactly what a dataset of clean, human-written patches never contains. A clean patch shows the right answer from the starting line. Your fix shows the right answer from where the agent actually ended up.

EXHIBIT 04Why a fix on the agent's own state is worth more than a clean demo.
Behavior cloninglearn only from where the expert wentexpert pathεno expert label out herecost ≲ T²εIn a coding session: a dataset of clean, human-written patches.DAggerlet the expert label where the learner wentexpert paththe expert labels the learner's own statescost ≲ TεIn a coding session: your fix to the agent's code.
Left: train only on states an expert visited, and one small error puts the learner somewhere it has no label for, so errors compound. Right: label the states the learner actually reaches. Ross, Gordon and Bagnell (2011) bound the cost at T²ε for the first and Tε for the second, up to constants, for a horizon of T steps and per-step error ε.

Cursor's five-hour loop is the same idea in RL clothing. It keeps collecting on the states the newest model visits.

Why the same fix is the most dangerous part of the data

Three ways it goes wrong, all visible in the session.

The context is missing. Your fix depended on the load balancer and on a file the agent never opened. The log contains neither. Train on the fix as a demonstration and you teach a guess. The clean options are to reconstruct the missing observation and mark it reconstructed, or to use the fix only as a score on the agent's turn and never as a target.

The author is missing. In a terminal harness your edit and the agent's edit look identical on disk. Get the actor wrong and the model is trained on a turn it never produced, in a context it never saw. A recorder needs an actor on every change, and "unknown" has to be an allowed value.

Filtering changes the objective. Whatever you drop from the dataset stops being penalized. Cursor found this the hard way. They discarded examples with invalid tool calls, and Composer learned that emitting a broken tool call on a hard task was a way to never receive negative reward. The fix was to keep the broken calls in as negative examples. The same applies to sessions with human edits. Filter them out as "contaminated" and the model is never scored on the turns people found worth fixing, which are precisely the turns that matter.

There is a fourth, quieter one. Human edits cluster where the model struggled and a person happened to be watching. That is a different distribution from ordinary use. It is a good distribution to learn recovery from and a bad one to estimate performance from.

What the labs have actually said

WhoWhat they sayThe signal
AnthropicFree, Pro and Max data trains new models when the setting is on, "including when you use Claude Code"Not disclosed
OpenAICodex tasks on a ChatGPT plan fall under the "improve the model for everyone" settingNot disclosed
GitHub CopilotInteraction data from Free, Pro and Pro+ users trains models by default since April 24, 2026, including "outputs accepted or modified by you"Accept, modify, thumbs
Cursor ComposerOn-policy RL from production, a new checkpoint every five hoursEdits persisting in the codebase, user pushback, latency
Cursor TabOnline policy gradient, redeployed every 1.5 to 2 hoursAccept +0.75, reject −0.25

None of them has described a pipeline that trains on a human's manual edit as a demonstration. What is documented is closer to method two: the human's reaction becomes a score. That is the honest summary of where the industry is, and it agrees with the analysis above. As a score, your fix is safe and cheap. As a target, it needs context you probably didn't record.

For the population view, SWE-chat collected 6,000 real coding-agent sessions from open-source developers with human-versus-agent authorship on every line. Users push back in 44% of turns, and 44% of agent-written code survives to a commit. The session in this post is small, but it is not unusual.

The most valuable line in a coding session is often the one the model didn't write. Record it, say who wrote it, reconstruct what the model could see, and only then decide what the gradient touches.

Download the session and the records

Everything in the exhibits comes from these two files.

  • session.json: the session, in the shape of a Claude Code log, plus the three events that happened outside it.
  • records.json: the same session flattened into spans with an author on each, the three SFT masks, the per-turn rewards, the DPO pair, and the diff the log never saw. It also names the identifier in the hindsight target that never appeared in the agent's context.