Six predictions, taken from three papers the OpenAIRE server chose for us, committed to git before a single market bar was read.
Three reproduce. Two leave no trace. One we refuse to answer.
2 predictions left no measurable trace · 3 reproduced · 1 was declared untestable before the run, and stayed that way
Saying that a paper cannot be tested on the data at hand — rather than answering anyway — is part of the result, not a gap in it. So is publishing the two predictions that failed next to the three that held.
Why let a graph choose?
Thousands of published claims describe how prices move. Anyone testing them on their own data picks which ones to test — and picks after having a rough idea of what their data will say. The result looks like a test and behaves like a search.
So we handed the choice over. OpenAIRE selects the papers, on a criterion stated before the query runs and logged verbatim. We never see a result and then decide which claim was worth testing. The graph is not used here as a search box — it is used as the selection mechanism itself, which turns “which papers” from a judgment call into an auditable artifact.
And once it has chosen, the graph is asked a second question that matters more than it sounds: an accessible paper is not an accessible dataset. OpenAIRE models that difference through typed relations, so we can ask each publication whether any data or software was ever linked to it.
| Publication | Linked dataset | Linked software |
|---|---|---|
| Lou, Polk & Skouras 2019 10.1016/j.jfineco.2019.03.011 | none returned | none returned |
| Korajczyk & Sadka 2002 10.2139/ssrn.305282 | none returned | none returned |
| Sadka 2006 10.1016/j.jfineco.2005.04.005 | none returned | none returned |
Read that precisely. Six calls to
openaire_explore_research_relationships on 20 August 2026 returned zero
relations — so what we can state is that no linked dataset was returned by the queried
OpenAIRE endpoint. That is not the claim “the authors published no data”, and we do not
make it. What it does establish is why the rest of this page has to exist: for these three
papers, re-running the original analysis on the original data is not an option available to
anyone. Testing the prediction on independent data is the only route left.
How the bench works
- 1 · DiscoverWe named an influence class; the OpenAIRE MCP server named the papers inside it, ordered by field-normalised citation impact.
- 2 · InspectAsk the graph whether each paper has linked data or software — is the original evidence reachable at all?
- 3 · FreezeOne testable claim per paper, with its expected sign and a verdict rule, committed to git before computing.
- 4 · MeasureRun against proprietary intraday bars already collected for another purpose — 2 866 stocks, 347 793 stock-days.
- 5 · JudgeReproduces · no trace · out of reach. Decided by a rule written in advance.
- 6 · PublishMethod, code, raw MCP calls and aggregates — everything except the market bars, which are licensed.
Step 2 is the whole point. A bench that writes its
predictions after seeing the numbers proves nothing. The predictions live in commit
ed1a739; the program that tested them was written after it.
The tug of war
2 of 4 leave no traceA tug of war: Overnight versus intraday expected returns
“We document strong overnight and intraday firm-level return continuation along with an offsetting cross-period reversal effect, all of which lasts for years.”
Two distinct claims live in that sentence: continuation (a stock's night predicts its next night; its day predicts its next day) and cross-period reversal (a night predicts the opposite of the day that follows). We tested them separately, because they can fail separately — and they did.
Spearman rank correlation computed within each stock, then the median taken across stocks — so that a handful of violently volatile tickers cannot speak for the whole market. 2 866 stocks with ≥ 60 usable sessions each, opening price ≥ $2. Verdict rule, fixed in advance: correct sign and ≥ 55% of stocks carrying it. Anything else leaves no trace.
Reading the split honestly. The cross-period reversal is the sturdier half of the paper: it survives a protocol far coarser than the original — plain serial correlation over six months, instead of decile-sorted portfolios over decades. Continuation does not survive that coarsening. That is not a refutation. A portfolio-level effect can be real and still be invisible to a stock-level correlation over 126 sessions. What we can say is narrower and worth saying: of the two effects, only one is robust enough to show up in a small, short, independent sample.
And then the costs eat it
reproduces, from the other endAre Momentum Profits Robust to Trading Costs?
“The price impact models imply that abnormal returns to portfolio strategies decline with portfolio size. We calculate break-even fund sizes that lead to zero abnormal returns.”
Their claim is not that the effect is illusory — it is that it dies above a certain amount of capital, because a large order moves the price against itself. We sit at the opposite extreme of that curve: $500 per position. No price impact whatsoever. And the effect dies anyway.
| Buying an 8% intraday drop, exit at close — 549 cases | Median |
|---|---|
| Net return after broker commissions | +0.09% |
| Same rule, random-entry control on the same stock and day | −0.54% |
| Edge over chance | +0.63 pt |
| Bid–ask spread actually paid — measured on 173 stocks, 1 011 660 bars, not included in the line above | −0.24% to −0.48% |
| Return once the spread is paid | −0.15% to −0.39% |
The control is the load-bearing part. In a rising market any entry makes money; only the gap between the rule and a coin flip means anything. The spread range spans the median spread across all trading minutes (0.24%) and the spread actually standing at the two moments this rule trades — early session and close (0.48%). Commission is $0.005 per share, $1.00 minimum per order, two orders, on $500 committed.
reproduces — the edge over chance is real and sizeable. But an edge over chance is not money: the control loses too, and the rule's own return is a rounding error away from zero before the spread and below zero after it — under either spread assumption, and without a single share ever being moved by our own order.
What the bench adds to the paper. The cost curve has two ends. Korajczyk & Sadka measured the one where cost rises with size, through price impact. At $500 the cost falls with size, because a $1.00 minimum commission per order weighs proportionally more the smaller the trade. Abnormal returns vanish at both ends. The interesting question is not whether a strategy works — it is the band of capital in which it can work at all, and that band is bounded from below as well as from above.
The one we refuse to answer
out of reach, declared in advanceMomentum and post-earnings-announcement drift anomalies: The role of liquidity risk
“Unexpected systematic variations of the variable component rather than the fixed component of liquidity are shown to be priced within the context of momentum and post-earnings-announcement drift portfolio returns.”
Sadka's liquidity measure is built from the unexpected component of order flow. Our bars are trade prints. We hold a bid–ask spread — which is a different object, the fixed component his paper explicitly sets aside as the one that is not priced. Substituting one for the other would produce a number, and the number would be meaningless.
This verdict was written into the frozen hypotheses before the run, not chosen afterwards to explain away a disappointing result. A bench that always returns an answer is a bench that has stopped checking whether it can.
What this does not show
- Not a replicationThe original studies use decile-sorted portfolios over decades of CRSP data. We use within-stock rank correlations over six months. Signs are comparable; magnitudes are not.
- One regimeFebruary to August 2026. Nothing here speaks to any other period.
- Not tradableH3 and H4 reproducing does not make them tradable — Paper 02's arithmetic is precisely what stands in the way.
- Returns are ceilingsSlippage is modelled at zero and limit orders are assumed filled the moment the bar's high touches them.
- Stale opensThe first printed price of a session is sometimes stale, which adds noise to the overnight leg. A source of noise, not a directional bias.
Reproducibility
| File | What it is |
|---|---|
HYPOTHESES.md |
The six predictions with their expected signs and the verdict rule, committed before the measurement program existed. Later corrections are appended, never rewritten over. |
bench_reproducibility.py |
The measurement. No network, no model, no agent — reads compressed bars, emits JSON. |
mcp_calls.jsonl |
Every OpenAIRE MCP call with its exact arguments and its raw response, so the paper selection can be replayed. |
lou_polk_skouras.json |
Aggregate output: medians, quartiles, stock counts, verdicts. |
mcp_calls_datasets.jsonl |
The six data-availability calls, arguments and raw responses, so the “none returned” above can be checked rather than believed. |
provenance/ |
The exact content of the pre-measurement commit, with its SHA-256 — and a plain statement of what that does and does not prove to a third party. |
Everything above is public. Repository: github.com/Alexry375/reproducibility-bench — method · provenance · MIT (code) · CC BY 4.0 (text, results).
What is not published, and why. The underlying minute and five-minute bars come from a commercial broker feed whose redistribution is contractually forbidden. Only aggregates, code and method are released. This is a licensing boundary, not a convenience: the bench is reproducible in method by anyone holding equivalent data.
Why an MCP server changed the shape of this
The papers above were not retrieved by keyword search and then ranked by a
human hunch. We named a topic and an influence class;
openaire_find_by_influence_class returned the work inside that class, ordered by
field-normalised, time-adjusted citation impact. The one choice we made is a single line and
it sits in the call log; which papers came back is not ours. The selection step carries a
stated, auditable criterion instead of a taste. That is the part worth keeping: not that an agent
can read papers, but that “which papers” stopped being an unexamined
choice.
openaire_find_by_influence_class(influence_class="C3", query="intraday momentum", page_size=5) openaire_get_research_product_details(identifier="10.1016/j.jfineco.2019.03.011") openaire_explore_research_relationships(doi="10.1016/j.jfineco.2019.03.011", target_type="dataset")