Verified indicator search: monte-neo discover¶
discover looks for a trading rule in a space of causal formulas and then asks the verifier whether the winner is real.
Searching many formulas always finds something that looks good on past data; the point of this command is that the
search is counted and tested, not just run.
monte-neo discover --ohlcv data.csv --out found/ --budget 2000 --seed 1
Options: --budget (candidates), --seed, --null-runs (shuffled markets, default 39), --lockbox (share of the last
bars kept for one final look, default 0.2), --cost-bps (round trip, default 5), --side long_short|long_flat, --evolve N (generations of mutating the best candidates; every child is a counted trial and the search null repeats the same procedure), --recheck DIR (run the recorded search again on --ohlcv; exit 4 unless the journal, the winner and the certificate id match).
Exit code 0 means found, 1 nothing found, 3 bad input. The MCP tool is discover_indicator.
What it does¶
- Causal formulas only. Candidates are trees of columns (open, high, low, close, volume), trailing windows and
arithmetic. They are never built from text with
eval; the source of the winner is generated from the tree. - A gate before the search. Every generated source passes the same truncation test the verifier uses. Nine known leaking formulas (canaries) are run through the gate first; if one passes, the run stops.
- Effective trials. Candidates whose returns correlate above 0.8 are one idea. The number of clusters is the
n_trialsused for the Deflated Sharpe. - Search null. The whole search is repeated on shuffled copies of the market. The best Sharpe on the real data must beat the best Sharpe on those copies (p ≤ 0.05).
- Reality Check, SPA and PBO on the cluster representatives (PBO above 0.5 is a reason against).
- Lockbox. The last part of the data is touched once, for the winner only.
- Certificate. The winner is verified like any strategy, with the effective trials counted.
found is true only when every step agrees.
Output files¶
strategy.py (the winner), result.json (statistics, reasons, certificate, config), search.jsonl (every candidate
with its Sharpe, hashed in journal_sha256), certificate.json, report.html.
Calibration¶
The calibration page is generated by scripts/discover_calibration.py: how often a search
names an indicator on random walks (with the best Sharpe a plain search would have reported on the same data), and how
often it finds a planted edge of a given strength.
What to expect¶
On noise the search finds nothing in about 11 of 12 runs, although the best naive Sharpe on noise is above 1.5. An edge planted in the data is found at costs up to about 1 bp and not found at 5 bp: the honest answer when costs eat the edge. The search finds structure that exists in the data you give it; it does not promise profit out of sample.