Everything the product can do — and why each ability is needed
Not a single function appeared here “because everyone has it”. Each one has a concrete error it catches, and a concrete number it uses to do it.
If a backtest can lie — why do one?
For the same reason a doctor orders a test, even though tests sometimes err. The alternative is not “exact knowledge”, it is an opinion. And opinions err more often, only silently.
An experiment is a question the market can answer “no” to
If your idea is formulated so that “no” is impossible in principle — “buy when the market shows strength” — that is not an idea, it is a belief. The first thing the product does is help turn a belief into a question: which exact event, on which instrument, over what horizon, by how much better than an ordinary day.
How decisions are usually made
And why this does not work even for smart people
What an experiment gives you
Four things you cannot get any other way
From an expensive mistake
Checking an idea costs a few seconds of machine time. Checking the same idea with live money costs months and a drawdown on which you will close it anyway.
But it does not give a guarantee
An experiment narrows, it does not prove. Even a perfectly run one raises the odds of a hypothesis from ~2 % to ~25 % — and that is still “probably not”.
One experience means nothing
What matters is the whole series. Twenty experiments in a row will almost certainly produce one “lucky” one — and that is the one you will remember if nobody is counting the other nineteen.
That is where everything else follows from. Experiments have become almost free: used to be, checking an idea took an evening, now it takes a second, and with AI agents you can launch them by the thousand. The cheaper the experiment, the more meaningless its result is on its own — and the more important it is that someone counts how many there were. That is what the product does.
Six lines between your idea and the verdict
Overfitting is when you, without meaning to, tuned the rules to a specific slice of history. No single technique catches this, so there are several lines and each catches its own class of error.
The bounds are computed once — at the moment the hypothesis is registered — and turn into concrete dates. After that history can grow as much as it likes: your holdout will not move because of that. If the slices were set as percentages, every data top-up would silently open the sealed part.
We count the attempts for you
Every run increments the counter automatically, and you cannot zero it. Similar variants — a 0.4 % stop and a 0.41 % stop — count as almost one: what matters is not how many times you pressed the button, but how many independent ideas you tested.
effective N through eigenvaluesWe raise the bar to your number of attempts
For each metric we compute how much the best result from the same number of random, meaningless rules would have shown. That is the threshold. Your result has to be above it — not above zero.
deflated Sharpe · √(2·ln N)We cut history by dates and seal the last slice
Training, validation, and holdout are separated at the moment the hypothesis is registered. At the joints a strip of bars is dropped so that a trade opened in training does not close already in validation.
purge & embargo · fixed boundsWe compare with a thousand fake histories
We shuffle history a thousand times, breaking the order of events but keeping its overall properties, and run your same rules on it. Then we count how many fake runs did better than the real one.
block permutation testWe break the strategy on purpose
We double the commission. We shift the entry one bar later. We change parameters to neighboring ones. We compute the commission at which the strategy goes to zero, and compare it with your rate.
stress tests and a robustness profileWe wait for data that was not there
The rules are sealed with a hash at registration — you cannot say after the fact “I meant a different stop”. After that the strategy simply goes forward in time, and the result accumulates on its own.
blind forward trackFree search is separated from the verdict
Each next report is stricter than the previous one and rests on it. Passing the first and failing the second is the most common outcome, and that is not a product error.
Research
Is there an effect at all. Two distributions — after the event and across all bars, the difference from the base rate and its confidence interval, the effect by horizons and market regimes.
There is no Sharpe here, and no equity curve: there are no rules yet, nothing to compute from. The trial counter is silent.
Run
Does the effect survive costs. An equity curve with a drawdown ribbon, a “gross − costs = left” breakdown, four fragility tests, a monthly map, trades.
A run does not issue a verdict: it does not know how many times you searched. Every launch is plus one trial.
Verdict
Is the result distinguishable from luck. Forward versus the expectation from the backtest, eight checks with thresholds, and a reproducibility stamp.
The stamp is a data snapshot, a rules hash, and history bounds. Without it the verdict is a picture; with it — a document.
An open catalog of what did not work
Nobody publishes negative results — that is why the internet looks like everything works. Every grave has a date, an epitaph with the cause of death, and the number that killed it.
Golden cross
MA 50 × MA 200 · US equities · daily · 2005–2025
Died of no effect. The difference from an ordinary day is +0.004 % — the confidence interval covers zero with a wide margin. Plenty of signals, no substance.
Mean reversion on five-minute bars
BTC/USDT · 5 minutes · 2019–2026
Died of costs. Before commission +0.031 % per trade, after commission and spread −0.004 %. The effect is real, but it fits entirely inside the cost of getting in and out.
RSI below 30 on daily bars
7 indexes · daily · 2010–2026
Died of over-searching. Sharpe 1.12 against a threshold of 1.58: to get that number the author tried 34 combinations of period and level. On two variants the threshold would have been 0.9.
MACD divergence
EUR/USD · hourly · 2012–2026
Died of fragility. It held exactly until the entry was shifted by one bar: +14.2 % a year turned into −1.8 %. The strategy lived on specific minutes, not on a market property.
Previous-day level breakout
ES · 15 minutes · 2018–2026
Died of one year. The entire eight-year result is collected in 2020: without it the return is −2.1 % a year. This is not a strategy, it is a memory of one event.
Morning range breakout
ES · 5 minutes · 2021–2024 · forward is running
Still alive. It did not pass history — 0.94 against a threshold of 1.47 — but the author left it on a blind forward. Twenty months later the actual result is tracking close to the expectation.
The cards are layout examples. The real graveyard will fill with the first checks.
So you do not bury twice
Before spending a week on an idea, look whether it already lies here — with a date, a trial count, and a cause of death.
So you can see the overall score
If four hundred people have checked the “golden cross”, your four hundred and first check is not independent. The graveyard is the trial counter for the whole product.
So we can be challenged
Every card has a data snapshot, the rules, and history bounds. Anyone can repeat the check and show that we were wrong.
Graveyard rules
- They are buried automatically. A completed check is published on its own. You cannot choose what to show — otherwise the graveyard would turn into a shop window.
- The name can be hidden, the result cannot. The author is published anonymously on request; the idea, the data, and the verdict are always published.
- We bury our own too. Checks we run ourselves go into the same catalog on the same terms.
- A grave can be opened. If the market has changed or there was an error in the check, the idea is restarted — the old card stays, a new one is added to it.
- An epitaph is required. One line about why the idea died. Without a cause of death the card is not published.
One history for everyone, fixed by a snapshot
Crypto
Spot and futures, minute candles from each pair’s listing.
Forex
Majors and metals, tick history more than twenty years deep.
Futures and indexes
Core CME contracts, minute and five-minute data.
US equities
Adjusted for delisting and historical index compositions.
A snapshot instead of “latest data”
Every hypothesis has a history-snapshot identifier recorded. Data is topped up — the experiment does not change because of that. New bars help the forward, and a revision of old ones creates a new snapshot version, while the old one stays available: a year-old result is reproduced verbatim.
Costs are set by you
Commission, spread, and the slippage model are parameters, not engine constants. By default they are pessimistic: better to reject a good idea than to let a bad one through. Separately we compute the critical commission — the one at which the strategy goes to zero.
How indicators actually work
Not “ten best RSI settings”, but what each indicator physically measures, what its lag is, and on which markets it cannot work in principle. The materials are written so that after them you invent more meaningful hypotheses, rather than look for ready-made ones.
What an indicator is
Any moving average is a filter with lag. A breakdown of what the lag equals and why it is not removed by tuning the period.
Where “levels” come from
What in levels is a property of the market, and what is a property of human vision, and how to tell one from the other by measurement.
Why reversal patterns do not work
A breakdown in numbers: how often the pattern occurs, what happens after, and where the effect goes after costs.
The section is filled together with the blog.
Counting the number of attempts is something one platform in twenty can do
This is not an advertising line, it is the result of an independent open ranking of mathematical rigor. The two columns on the right decide everything else.
| Platform | Rigor | Accounts for attempts | Blind forward | Price |
|---|---|---|---|---|
| normaltrading in development | no score | ✓yes | ✓yes | free tier |
| Minerva US equities and ETFs only, 2016–2026 | 7.9 | ✓yes | ✕no | $948–2 388 / year |
| StrategyQuant X Monte Carlo is there, no adjustment for N | 6.8 | ✕no | ✕no | $1 290–2 900 one-time |
| QuantConnect requires Python or C# | 6.6 | ✕no | ✕no | free · from ~$60 / mo |
| Wealth-Lab | 6.4 | ✕no | ✕no | ~$400 / year |
| MetaTrader 5 | 5.5 | ✕no | ✕no | via broker |
| TradeStation | 5.4 | ✕no | ✕no | via broker |
| TrendSpider no walk-forward, no Monte Carlo | 4.9 | ✕no | ✕no | $59–233 / mo |
| Composer 3 000+ ready-made strategies to choose from | 4.7 | ✕no | ✕no | free · $32 / mo |
| TradingView no overfitting protection, acknowledged in the documentation | not ranked | ✕no | ✕no | free · $13–200 / mo |
Scores are an independent open ranking of mathematical rigor. Prices are public plans as of September 2026. We deliberately do not give ourselves a score: the product has not yet gone through an external evaluation, and a self-score is worth nothing.
How we differ from Minerva
The only platform that counts attempts — and its math is stronger than ours. But neither it nor anyone else in retail has a blind forward period, and its authors honestly publish that on real data their score showed no link to future returns. Plus: we have crypto, forex, and futures, not only US equities; a threshold on every metric, not one overall score; a graveyard and a library.
How we differ from the wave of AI products
One of them searches millions of precomputed strategies and ranks them by Sharpe — in the announcement not a word about overfitting or out-of-sample testing. The cost of an experiment for AI agents falls almost to zero, and that makes counting attempts not an option but the only way to keep meaning in the result. We do not pick a strategy for you: you bring the idea.
Send an idea you have long wanted to test
While the product is in development, we run ideas by hand and send back a full report — with every threshold and every explanation. Free, no strings. The only ask in return is to tell us what in the report stayed unclear.
You will describe the idea in the reply email — in words, no formulas.