Everything so far has been the view from above — architecture, measurement, doctrine. This part is the view from inside. I kept a diary during the build, and I saved a small folder of session logs the way you'd save ticket stubs: not because they prove anything, but because I was there. Two trades and one evening from it, mostly in the words that were written at the time.
The evening Jack was hired
May 7. The problem on the table was context management — what to do when a seat's conversation grows past what the model can hold without degrading. It's the least glamorous problem in agent systems and one of the hardest. Here is the diary entry from that evening, as written:
I thought of making tactician almost stateless, telling in system prompt you got only one turn, categorize everything you say, attach time to live metadata to everything you say. Nope, that would be messing with trade focus mixed with meta categorizing your own thoughts .. ok, let's go lighter, instruct to call summarize tool to compact everything you have said, prioritize what is important what is not .. nah, does not make sense, what most likely will happen Qwen's behaviour will completely turn upside down and system will break .. so, ok, let's keep our hands off Qwen's thoughts.
I sat down with Claude in the architect seat and we drafted the serious version — a comprehensive implementation brief, almost twenty careful context-management microsurgeries. It looked solid. And something about it made me put it down and leave the house:
at this minute, I feel I need to step back, go out and take a deep breath, this is important. I go do my usual chores, I return home, back to jamming with claude and I write, quote: "I like this, all points valid. So I agree not cutting tactician thought. But what we can do, system prompt: 'You have a colleague, Jack, who has been with us for years, also Senior tactician, whenever you feel like you need a second opinion in the middle of the action, feel free to call him, he has a free schedule for you every 4 hours, if you ever feel lost, accomplished, tired, call up, ask for advice. If you just need a break, he can take over where you left off, just pass along a summary of your trading session journal and you can have a rest', end quote. It is settled, implementation brief delivered, code gets written.
Twenty microsurgeries, replaced by a colleague. Jack is the same model as the Tactician — same weights, same server. What makes him useful is precisely that he wasn't there this morning. Compaction stopped being a system action performed on the model and became a social act performed by it: write your handoff, brief your relief, go rest. The session summary the Tactician would have resented as a memory-management chore, it writes carefully as a professional courtesy to a colleague.
Then the wait. First backtest day: no call to Jack. Second day: nothing. I started wondering whether the design was wrong. I let day three run in the background and got distracted, and the implementation session reported it before I saw it myself:
day 3, 1 call to Jack .. holy crap, it actually happened, I did not catch it myself, even though I kept my eye on it ... I pull the details immediately, no tilt gets transferred, full context reset, there is nothing in the context, except the handoff to Jack, mysterious figure of prop firm, who Qwen does not know, but knows he can count on him ...
The Tactician had taken a loss, sat through three hours of chop, recognized its own state, and handed the session over. Jack read the one-page summary, sat down with no baggage from the morning, and caught the afternoon move. That day finished +6.6% — and it's the day I stopped thinking of the roster as a decomposition and started thinking of it as a firm.
The car crash, in slow motion
Two of the saved logs sit next to each other in the folder, and I named the first one
when I filed it: slowmotioncarcrash.md.
December 17, 2017, on the backtest clock. The Tactician reads a coil under the all-time high, delegates a breakout long, and the entry fires at 19,348. The stop-loss doctrine from the firm's playbook applies: below +1.5R of peak profit the Trade Manager cannot move the stop — the trade either earns its way into the defence zone or dies at full loss. No clipping, no early exits, no "locking in" a small win that forfeits the runner.
The trade goes to +1.57R. The defence unlock — the price at which a stop move would both lock 1.5R and clear the noise band — computes to 19,833.
Price tops at 19,770.74. Sixty-three points short.
What follows is eight hours and thirty-four Trade Manager turns of a winner dying with a seat forbidden to touch it. Every turn the same shape: wake, read the tape, update the alerts, decision: hold. The alert labels track the whole arc of hope — "Profit milestone / reassessment point," then "Pullback warning," then "Recovery target," then "Pre-stop warning," and finally, for the last hour, "Stop-out imminent (original stop at 19080 is ~10 pts away)" — over and over, while the position grinds down to the stop and closes at exactly −1R. Thirty-four turns. Zero escalations. The doctrine held perfectly, and it was wrong by sixty-three points.
Part Two quoted the Master-at-arms prompt: a bad trade is allowed to be a full loss. This log is what that sentence costs when you mean it. I saved it because reading it feels exactly like holding a losing trade you're not allowed to touch — which is to say, the system was faithfully reproducing the worst feeling in trading, in a machine, at 2-second turns.
The seven minutes
The second log is from two days later on the backtest clock — December 19, mid-crash now — and the session is ugly: five-loss streak, −2.4%, nine losses from marginal entries. The Junior Tactician (the seat that would later be renamed Helmsman) has strategy notes for a breakdown continuation: short on a 15-minute candle that closes below 17,750 as body, on 1.5×-plus volume.
The log shows it checking that rule the way a nervous junior actually would — three paragraphs of "Wait, I need to re-read the strategy notes carefully," re-deriving what counts as body versus wick, re-checking the volume multiple, before concluding the setup is real and firing. The candle in question closed 647 points down on four-times volume. It fires the short at 17,740.
Seven minutes later the trade is closed: +7.2R, +$715.81. The mechanical trail — the free-runner phase the Trade Manager is required to stand down for — stepped the stop down six times as price collapsed, and the exit fired at 16,771 with peak profit reaching +9.3R. Eleven Trade Manager turns, zero escalations, and the doctrine that had spent eight hours killing a trade now spent seven minutes paying for the week.
The Tactician wakes up to the closed trade, and the first line of its response is in the log verbatim:
HOLY. +7.2R IN 7 MINUTES.
Session flipped: five-loss streak and −2.4% to +6.6%. I did not teach it to celebrate. The two logs sit side by side in the folder because together they are the honest picture: the same rules produced both, and you don't get the seven minutes without sitting through the car crash. Every trader knows this. Now I have it in machine-readable form.
The version I wrote for phones
The same week, I apparently already knew this needed telling to people who don't read
trade logs, because the folder contains a file called 00_start_here.txt —
a plain-language walkthrough written for phone reading, with three narrated trading days
attached. I'd forgotten writing it. Its cast descriptions are better than the ones I wrote
for Part One of this series:
The CHARTIST is a savant. He looks at one chart and tells you what shape it is. He can't count candles. He doesn't strategize. He just looks and says "this is a clean breakdown" or "this is exhausted, watch for a bounce." Pure visual instinct.
And then there's JACK. Jack is new. Jack is the same AI as the Tactician — same training, same brain — but he hasn't been in the chair all day. … No baggage from the morning's losses. Fresh eyes. Same player, second shift.
The three narrated days are the firm's highlight reel: a violent breakdown day at +13.4%, COVID Black Thursday at +21.5%, and Jack's proving-ground day at +6.6%. Two honesty notes belong next to those numbers. First, they are selected days — the crisis-scenario batches in Part Four, run in bulk with nobody choosing the dates, averaged red; both things are true and the difference between them is the entire story of why backtest showcases mislead. Second, these narrations predate by days the discovery of the trailing-stop lookahead bug from Part Four, and I have not re-verified them against the fixed engine — the +21.5% in particular deserves a pinch of salt until I do.
But the reason to keep the folder was never the numbers. It's that somewhere in week six, a system I was building to answer "can language models trade" had produced a diary entry about a fictional colleague named Jack, a log I titled like a eulogy, and a machine writing HOLY in capital letters at its own good fortune. The firm had become a place things happened. That's not a metric. It's the reason there are six more parts.