The scorecard & verdicts

How to read the verdict and the checks beneath it — and why running the same idea again raises the bar it has to clear.

A run produces a profile, not a score out of 100. The scorecard shows which checks passed, which failed, by how much, and — where the evidence is clear — what that pattern of failure usually means.

The verdict

The banner at the top of the scorecard is one of four:

  • Cleared — no veto failed and no graded check failed.
  • Cleared with one gap — no veto failed and exactly one graded check failed. The subtitle names which one.
  • Has not cleared — a veto failed, or two or more graded checks failed.
  • Not enough trades to grade — the run missed the trade-count floor and produced too few trades for any other graded check to return a result. This is not a pass and not a graded failure: there was nothing to grade. A strategy that never opened a position lands here. More history, or a less restrictive entry, is the next step — re-running the same spec against the same data is not.

The subtitle beside the verdict names the check that decided it — the veto that failed, or the gate that did not clear. It is the single most load-bearing line on the page.

What is on the scorecard, in order

  1. The verdict and its one-line reason.
  2. Attribution — a plain-language account of what the evidence says, with a link to a historical case where the same pattern appeared, when there is one. This layer is a Pro and Research feature; on the Free tier you see the raw checks and a note that Pro adds the explanation.
  3. The outcome panel — the gross and net result and, where there are trades, the equity replay, its annualized return, and a dashed line showing what simply buying and holding the traded symbol over the same window would have done. A “Cleared” verdict does not weigh the strategy against that line; the panel is where you do.
  4. Vetoes and graded profile — every check as its own row. Each row expands into the check’s definition, its floor, and how the observed value was computed. The definitions live here, in the product, not behind marketing copy.
  5. Deflated-Sharpe accounting — which trial this is in your history, and your significance after adjusting for every trial you have spent. See the glossary for what the deflated Sharpe is.
  6. Trade diagnostics — a breakdown of where the result came from across the trade log.
  7. Prop challenge (Pro and Research) — how often this trial’s own trades would have passed a prop firm’s evaluation rules. It runs only when you ask it to; see below.
  8. Details (collapsed) — the run record: trial number, the spec hash, the floors as locked at submission, run duration; plus the interpreted intent and, on Pro and Research, a chat grounded in this run’s own data.
Check the spec hash
The details section shows the hash of the exact spec that was run. It is there so you can confirm the scorecard belongs to the strategy you think you submitted — not a stale copy, not a half-applied edit.

Attribution is diagnosis, not prescription

Which explanation applies is decided by deterministic code matching the shape of your scorecard against a table of known patterns — not by a language model. The model only fills your numbers into the matched template and adjusts the phrasing. It never chooses the pattern, never names a parameter to change, and never ranks what to try. Generated text is filtered against a banned-phrase list before you see it; anything that reads as advice is blocked.

“No matched pattern” is not a failure
When your scorecard does not clearly match one of the characterized patterns, the attribution says exactly that and shows you the raw checks. A fabricated single-cause story for a mixed profile would be worse than an honest “the evidence does not point to one thing”. The raw scorecard is the answer in that case.

What “Cleared” does and does not mean

Cleared means the strategy survived every check the platform could apply, on this data, at this sample size, against floors that were fixed before the run. It is not a prediction that the strategy will make money, and it is not permission to stop testing. The next filter is time: forward-registering the strategy tracks it on data that did not exist when you built it, which is the one test no backtest can fake.

Cleared is also a statement about statistical significance, not economic significance. Every check asks whether an edge is real; none asks whether it is large enough to be worth trading, or whether it beat the far simpler alternative of just holding the asset. The outcome panel now puts a buy-and-hold line and an annualized return next to the strategy’s own so you can see that for yourself. If you want the platform to enforce a magnitude bar, commit a minimum net-of-cost profit factor in your validation criteria — it adds one graded check, the economic significance row, on top of the fixed stack.

Every change you make and re-run is a new trial with its own scorecard, and the significance bar rises each time — see Getting started on why refining is not free.

Prop challenge simulator

On Pro and Research, the scorecard carries a Prop challenge panel. It asks one question of a finished trial: at a given risk per trade, under a prop firm’s evaluation rules, how often would this strategy’s own trades have passed? It reads the trades already stored with the trial. It is not a new trial and not a validation check — nothing it computes is recorded to your history or the deflated-Sharpe count, and it uses no quota. It runs only when you press Simulate.

  • Historical start dates — the rules replayed from every firm-day reset in the data window that leaves room for a full attempt. This is a historical pass frequency, not a probability: neighbouring start dates share most of their trades, so the panel also shows roughly how many independent windows the history holds, with a rough interval.
  • Monte Carlo — a block bootstrap of the trial’s own trading weeks, at 1- and 2-week blocks; the lower of the two is reported. It assumes the trial’s edge and its mix of market regimes persist.
  • Zero-edge baseline — the same starts and paths with each trade’s average net R removed. The gap between it and the real figure is how much of the pass rate comes from edge rather than luck.
  • Risk sweep — the pass rates across a grid of risk levels, in ascending order. Nothing on it is marked as a level to use; the only highlight is the risk you chose.

Wherever a firm’s rules leave room for interpretation, the simulator takes the reading that lowers the pass rate. Trades are sized from the starting balance with no compounding; only net-of-cost R is used; each trade’s worst adverse move is assumed to come before any gain; touching a limit counts as breaching it; losses are booked on the latest firm day they could fall on and profits on the earliest; floating profit never counts toward a target. An attempt still undecided when the data runs out counts against the pass rate. A trial that did not clear validation can still be simulated, and the panel labels it as unvalidated.

What it does not model
News-trading rules, lot and position caps, weekend or overnight bans, funded-stage payouts, and slippage beyond what the backtest already prices. One symbol per simulation; pair-mode trials are not supported. Each preset is modelled on a firm’s published rules as of the date shown beside it — check the firm’s current rules before relying on one.