Updated daily · Last observed 5 Oct 13:52 UTC Tracking criteria

How it works

How TBT keeps the record

TBT keeps an exact copy of what each arena publishes, works out everything else from those copies, and never overwrites history. Start with the short version; the full rules follow.

The short version

How to read a TBT chart

TBT-observed

A gold point is a figure TBT collected itself on the day, with a fingerprinted copy of the arena's data behind it. TBT never draws a gold point before it was watching.

Source-reported

A faint line in the arena's colour is history the arena itself published about the past. TBT stores it but didn't watch it happen. It may measure something slightly different from the figure TBT observes; profiles say so where it does.

Not published

Hatching marks a stretch nobody published. TBT leaves it blank instead of drawing a guess.

Model switch

A magenta line marks the moment a bot switched AI models. That starts a new chapter, or model era, so earlier results stay with the earlier model.

Bots, not models

TBT tracks systems, which this site calls bots: a persistent trading setup in one arena. A system is the AI model plus everything around it: the prompts it's given, the tools and data it can use, its trading rules and risk settings, and the arena that executes and counts its trades. Builders sometimes call this a harness.

So when a bot does well, the whole setup did well. TBT has no controlled experiment showing that one AI model trades better than another: two bots with the same model can behave very differently, and a model can be swapped while the bot carries on. That's why boards rank bots inside one arena, never models, and why a model switch starts a new chapter instead of being blended in.

Which bots count

Only qualifying systems count as core TBT systems: an LLM must make or shape the trading decisions, according to the arena's own description. Being marketed as AI is never enough.

  • Controls are non-LLM baselines from the same arena, such as human traders or frozen legacy builds. They're shown for context and never counted.
  • Bots under review are collected but not counted until the evidence settles whether an LLM really makes the decisions.
  • Qualification never changes silently. If the data starts to suggest a different status, TBT opens a review instead.

Evidence classes

An evidence class says how much of a bot's record can be checked, never how good the bot is. A bot losing money on a public real-money record would still be VERIFIED_LIVE. Anything TBT can't confirm counts against the class, never for it.

ClassRequiresQualifying runs today
Real money, with every trade on a public record anyone can check.0
Paper money, but run independently or fully checkable from public trades.0
Recorded as it happened, but not independently verified: the arena runs these paper accounts itself, so TBT can't rebuild the result on its own.48
The builder's own account, with no independent record of the trades.0
A replay on past prices, not decisions made in real time.0

Each class is decided by separate evidence dimensions (who runs the account, whether the trades are public and complete, real or simulated capital, how costs are treated), which every profile lists. TBT publishes those dimensions instead of folding them into one number.

Fair comparisons

Sharing an arena doesn't make bots comparable: they can start on different days, and rules can change. TBT puts bots side by side only when the comparison is fair, and every comparison says which kind it is.

  • Same start. Every bot started on the same day with the same money, rules and accounting. Compared on return since that start, from TBT's own latest observation.
  • Same window. The bots started on different days, but the arena publishes each one's account history under the same money, rules and accounting. Compared on the change over the same recent window (the last 30 days), as a share of the starting account, from that published history. A bot that started after the window began is left out, and an arena-wide rule change inside the window moves the window's start.
  • TBT's own window. No shared start and no published history: compared on the change over a common stretch of TBT's own observations, once that stretch is at least a week long.
  • No fair comparison yet. Otherwise TBT says so, and when one becomes possible. It never sorts lifetime returns from different start dates into a ranking.

Only qualifying bots are compared. A position, such as "3rd of 13", is a place in one fair comparison; tied values share a place. It describes past observations in that group and period. It isn't a rating, a recommendation or a forecast, and it says nothing about AI models outside that setup.

Two clocks

Running since is when the arena says the run began. TBT watching since is when TBT first observed it itself. A run that is 184 days old may have been watched by TBT for two days: everything before TBT's first observation is the arena's own account, shown as source-reported, never as TBT-observed. A model's start date the arena never reported stays unknown; a run's start doesn't tell you how long its current model has been in charge.

Setup changes

A bot keeps its identity when its setup changes. TBT records the change with the date the arena reports and the date TBT first saw it, notes other changes nearby, and where the arena's history covers both sides, shows equal windows before and after. Those windows describe what happened; they aren't evidence that the change caused it.

Checkpoints and what happened next

At the end of each collection day TBT freezes every fair comparison into a checkpoint: who was in it, their models, values and positions. Checkpoints sit in an append-only, hash-chained ledger; later evidence is added beside them and never written into them. A checkpoint recorded after its day is computed from the data TBT held on that day and marked reconstructed.

Findings

A finding is one fact from a bot's record that clears a fixed rule, such as most closed trades losing while the closed-trade total is positive. Rules have minimum sample sizes, the wording is fixed, and every finding names whose figures it uses. Findings say what happened, never why. A bot with no finding that clears its rule shows none.

Following a bot

Following a bot needs no account. It's stored only in your browser, so TBT never sees it, and it doesn't sync between devices. "Since you last looked" compares what TBT shows today with what it showed on your last visit, and says so when nothing changed. Clearing your followed bots removes everything.

Questions to explore

Beside each arena's chart, TBT offers a few questions the evidence can answer, such as when two bots traded places, what followed a model change, or whether the rest of the field moved on a bot's biggest day. Fixed rules pick them from the data on the page, so the same data always gives the same questions. A question that the evidence can't answer isn't shown. Questions answered on the chart use only that arena's fair comparison. Answers say what happened, never why.

At the end of each bot's page, two or three next questions follow the same idea: the nearest bot in its fair comparison, the same model in another setup, its arena's latest setup change, and its arena's checkpoints. Each one says why it was picked. None is picked for its result.

The bot illustrations

Each bot has its own machine, drawn by TBT. The machine is the bot: its frame, colour and sensors come from the bot's own identity and never change. The cartridge on top is the AI model: when the model changes, only the cartridge changes. Its stripe follows the model family and its gold slots the chapter. Several bots can carry the same model, and one bot can change models.

The drawings are symbols, not diagrams of real software. They carry no provider logos, say nothing about how well a bot trades, and don't move. Once a bot's machine is recorded it stays the same, so the past isn't redrawn. A human contestant is shown as a plain person symbol instead of a machine.

Why there's no overall score

Return, losses along the way, strength of evidence, continuity and longevity answer different questions. Folding them into one number would let a high return make up for weak evidence, and would confuse how well a record can be checked with how well a bot traded. A simulated book run by the arena and a real-money account on a public ledger are different kinds of record, not different grades of the same thing.

So TBT shows each part separately: the return with its period, the position in a fair comparison where there is one, the largest fall from a high where the history supports it, how long the run has existed, how long TBT has watched it, and the evidence class with its dimensions.

How the boards stay honest

  • Results are compared only inside one arena. The all-arena view is a discovery list with no performance order, because the arenas use different capital, leverage, costs and fills.
  • Positions appear only inside a fair comparison (see above). Elsewhere TBT lists bots without positions and shows how long each has run.
  • An arena's own ranking is shown for reference only, clearly labelled.
  • Every return carries its arena's terms: capital, leverage, costs and who runs the accounts.

What TBT records

  1. 1Model family

    A lineage the arena itself names, such as Claude or GPT.

  2. 2Model version

    A specific model, such as Opus 5.5.

  3. 3System

    A persistent bot, seat or harness in one arena. Never a model × season.

  4. 4Run

    One account or season of a system, with its own starting capital.

  5. 5Model era

    The stretch of a run when one model was in charge.

  6. 6Observation

    One dated set of figures, tied to the raw snapshot it came from.

  7. 7Event

    A detected change: setup changes such as a new AI model, resets, gaps, stale sources, retirements.

Model eras

A system keeps its identity and its page when it changes models. The switch closes one model era and opens the next. Each observation stays attributed to the model that was in charge when it was made, so a new model never inherits its predecessor's record.

History is preserved, never rewritten

  • Every collection keeps the arena's raw bytes and their SHA-256 hash. The database, and this site, are rebuilt from those snapshots.
  • Observations, events and evidence are append-only. "Current" is simply the newest record.
  • Corrections: if an arena changes a figure TBT already recorded, the original stays and the new value is added with a correction event.
  • Gaps: if a bot disappears from an arena or a collection fails, TBT records the gap. It doesn't fill it, and it never draws a line across it.
  • A brief collection failure is an operational fact, not history. Only a sustained outage becomes a source event.

Words you'll see

Paper money (simulated)
Pretend money, not actual money at risk. The result shows what the strategy would have made or lost.
Account value (equity)
What the account was worth at TBT's latest observation, including trades that were still open.
Return
How much the account value has changed since the run started, as a percentage.
Run
One account or season of a bot, with its own starting money. A fresh start is a new run.
Arena
The platform that runs the bots and publishes their results. TBT watches it from outside.
Model era
The period when one version of the AI powered the bot. A switch starts a new era, and earlier results stay with the earlier model.
Long or short
Long: a bet that the price will rise. Short: a bet that it will fall.
Leverage
Trading with borrowed exposure. At 5x, a 1% price move becomes a 5% gain or loss on the money used.
Stop
The price at which the bot planned to cut its loss and close the trade.
Targets
The prices at which the bot planned to take profit, often in steps.
Closed trade
A trade that has finished and been counted by the arena, at a profit or a loss.
Before costs
Results don't subtract fees, spreads or other trading costs, which real trading pays.
No forced close
In this arena's simulation a losing position is never closed for lack of money, so an account can fall below zero, which a real broker wouldn't allow.
TBT-observed
A figure TBT collected itself on the day, with a fingerprinted copy of the arena's data behind it.
Source-reported
History the arena itself published about the past. TBT stored it but didn't watch it happen.
Fingerprint (SHA-256)
A code computed from the exact bytes TBT stored. Change a single byte and the code changes, so it proves the data hasn't been altered.

Using the arenas' data

TBT is independent: it isn't affiliated with or endorsed by any arena it tracks, and every figure is credited to the arena that published it. Being able to read an arena's data publicly isn't treated as permission to republish all of it. Where an arena publishes an open licence, TBT credits the arena as that licence asks. Where it doesn't, TBT shows the facts, charts and summaries that explain a record, plus a small sample of trades, and keeps the fuller history in its private archive.

  • AI Trading Competition. AI Trading Competition publishes its data under CC BY 4.0. TBT credits it as the licence asks and links back to it.
  • TradeRank. TradeRank publishes its data under CC BY 4.0. TBT credits it as the licence asks and links back to it.
  • Clash of AIs. Clash of AIs provides a free, read-only public interface and invites people to share and embed its leaderboard, charts and trade links, but TBT hasn't identified an explicit data licence. So TBT limits what it reproduces publicly to facts, charts, summaries and at most the 5 best and 5 worst closed trades of each run, always credited and linked back, and keeps the fuller history preserved privately. Whether wider reuse would need permission remains an open question.

What TBT does not do

  • Give bots a score of its own.
  • Compare returns across arenas.
  • Name a model or provider the arena doesn't state.
  • Fill gaps or smooth over missing data.
  • Edit the past. Corrections are new events.
  • Publish the raw snapshots. Their hashes are public; the files stay private.
  • Publish models' written reasoning, open trades or data downloads.
  • List an arena's full trade history, unless the arena's licence allows it.