Reproducibility Bench OpenAIRE AI Hackathon 2026 · track: analyse

Reproducible in method, not in data: the bars are a commercial feed and stay unpublished. 2 866 US equities · ~126 trading sessions · 347 793 stock-days · CC BY 4.0

Six predictions, taken from three papers the OpenAIRE server chose for us, committed to git before a single market bar was read.

Three reproduce. Two leave no trace. One we refuse to answer.

2 predictions left no measurable trace · 3 reproduced · 1 was declared untestable before the run, and stayed that way

Saying that a paper cannot be tested on the data at hand — rather than answering anyway — is part of the result, not a gap in it. So is publishing the two predictions that failed next to the three that held.

The question

Why let a graph choose?

Thousands of published claims describe how prices move. Anyone testing them on their own data picks which ones to test — and picks after having a rough idea of what their data will say. The result looks like a test and behaves like a search.

So we handed the choice over. OpenAIRE selects the papers, on a criterion stated before the query runs and logged verbatim. We never see a result and then decide which claim was worth testing. The graph is not used here as a search box — it is used as the selection mechanism itself, which turns “which papers” from a judgment call into an auditable artifact.

And once it has chosen, the graph is asked a second question that matters more than it sounds: an accessible paper is not an accessible dataset. OpenAIRE models that difference through typed relations, so we can ask each publication whether any data or software was ever linked to it.

PublicationLinked datasetLinked software
Lou, Polk & Skouras 2019 10.1016/j.jfineco.2019.03.011 none returnednone returned
Korajczyk & Sadka 2002 10.2139/ssrn.305282 none returnednone returned
Sadka 2006 10.1016/j.jfineco.2005.04.005 none returnednone returned

Read that precisely. Six calls to openaire_explore_research_relationships on 20 August 2026 returned zero relations — so what we can state is that no linked dataset was returned by the queried OpenAIRE endpoint. That is not the claim “the authors published no data”, and we do not make it. What it does establish is why the rest of this page has to exist: for these three papers, re-running the original analysis on the original data is not an option available to anyone. Testing the prediction on independent data is the only route left.

Method

How the bench works

  • 1 · DiscoverWe named an influence class; the OpenAIRE MCP server named the papers inside it, ordered by field-normalised citation impact.
  • 2 · InspectAsk the graph whether each paper has linked data or software — is the original evidence reachable at all?
  • 3 · FreezeOne testable claim per paper, with its expected sign and a verdict rule, committed to git before computing.
  • 4 · MeasureRun against proprietary intraday bars already collected for another purpose — 2 866 stocks, 347 793 stock-days.
  • 5 · JudgeReproduces · no trace · out of reach. Decided by a rule written in advance.
  • 6 · PublishMethod, code, raw MCP calls and aggregates — everything except the market bars, which are licensed.

Step 2 is the whole point. A bench that writes its predictions after seeing the numbers proves nothing. The predictions live in commit ed1a739; the program that tested them was written after it.

Paper 01

The tug of war

2 of 4 leave no trace

A tug of war: Overnight versus intraday expected returns

Lou, Polk & Skouras · Journal of Financial Economics, 2019 · DOI 10.1016/j.jfineco.2019.03.011

272 citationsinfluence C3popularity C2 impulse C2found via find_by_influence_class

“We document strong overnight and intraday firm-level return continuation along with an offsetting cross-period reversal effect, all of which lasts for years.”

Two distinct claims live in that sentence: continuation (a stock's night predicts its next night; its day predicts its next day) and cross-period reversal (a night predicts the opposite of the day that follows). We tested them separately, because they can fail separately — and they did.

What the paper predictsWhat the bars say
night → next night, positive H1 · continuation
no trace median ρ +0.001751.2% of stocks carry the sign, against a 55% threshold fixed in advance
day → next day, positive H2 · continuation
no trace median ρ −0.013045.7% of stocks carry the sign — the median itself leans the other way
night → same-day session, negative H3 · cross-period reversal
reproduces median ρ −0.034664.5% of stocks carry the sign
day → following night, negative H4 · cross-period reversal
reproduces median ρ −0.051571.6% of stocks carry the sign

Spearman rank correlation computed within each stock, then the median taken across stocks — so that a handful of violently volatile tickers cannot speak for the whole market. 2 866 stocks with ≥ 60 usable sessions each, opening price ≥ $2. Verdict rule, fixed in advance: correct sign and ≥ 55% of stocks carrying it. Anything else leaves no trace.

Reading the split honestly. The cross-period reversal is the sturdier half of the paper: it survives a protocol far coarser than the original — plain serial correlation over six months, instead of decile-sorted portfolios over decades. Continuation does not survive that coarsening. That is not a refutation. A portfolio-level effect can be real and still be invisible to a stock-level correlation over 126 sessions. What we can say is narrower and worth saying: of the two effects, only one is robust enough to show up in a small, short, independent sample.

Paper 02

And then the costs eat it

reproduces, from the other end

Are Momentum Profits Robust to Trading Costs?

Korajczyk & Sadka · SSRN Electronic Journal, 2002 · DOI 10.2139/ssrn.305282

419 citationsinfluence C3popularity C3 found via find_by_influence_class

“The price impact models imply that abnormal returns to portfolio strategies decline with portfolio size. We calculate break-even fund sizes that lead to zero abnormal returns.”

Their claim is not that the effect is illusory — it is that it dies above a certain amount of capital, because a large order moves the price against itself. We sit at the opposite extreme of that curve: $500 per position. No price impact whatsoever. And the effect dies anyway.

Buying an 8% intraday drop, exit at close — 549 casesMedian
Net return after broker commissions+0.09%
Same rule, random-entry control on the same stock and day −0.54%
Edge over chance +0.63 pt
Bid–ask spread actually paid — measured on 173 stocks, 1 011 660 bars, not included in the line above −0.24% to −0.48%
Return once the spread is paid −0.15% to −0.39%

The control is the load-bearing part. In a rising market any entry makes money; only the gap between the rule and a coin flip means anything. The spread range spans the median spread across all trading minutes (0.24%) and the spread actually standing at the two moments this rule trades — early session and close (0.48%). Commission is $0.005 per share, $1.00 minimum per order, two orders, on $500 committed.

reproduces — the edge over chance is real and sizeable. But an edge over chance is not money: the control loses too, and the rule's own return is a rounding error away from zero before the spread and below zero after it — under either spread assumption, and without a single share ever being moved by our own order.

What the bench adds to the paper. The cost curve has two ends. Korajczyk & Sadka measured the one where cost rises with size, through price impact. At $500 the cost falls with size, because a $1.00 minimum commission per order weighs proportionally more the smaller the trade. Abnormal returns vanish at both ends. The interesting question is not whether a strategy works — it is the band of capital in which it can work at all, and that band is bounded from below as well as from above.

Paper 03

The one we refuse to answer

out of reach, declared in advance

Momentum and post-earnings-announcement drift anomalies: The role of liquidity risk

Sadka · Journal of Financial Economics, 2006 · DOI 10.1016/j.jfineco.2005.04.005

661 citationsinfluence C3popularity C2 found via find_by_influence_class

“Unexpected systematic variations of the variable component rather than the fixed component of liquidity are shown to be priced within the context of momentum and post-earnings-announcement drift portfolio returns.”

Sadka's liquidity measure is built from the unexpected component of order flow. Our bars are trade prints. We hold a bid–ask spread — which is a different object, the fixed component his paper explicitly sets aside as the one that is not priced. Substituting one for the other would produce a number, and the number would be meaningless.

This verdict was written into the frozen hypotheses before the run, not chosen afterwards to explain away a disappointing result. A bench that always returns an answer is a bench that has stopped checking whether it can.

Limits

What this does not show

  • Not a replicationThe original studies use decile-sorted portfolios over decades of CRSP data. We use within-stock rank correlations over six months. Signs are comparable; magnitudes are not.
  • One regimeFebruary to August 2026. Nothing here speaks to any other period.
  • Not tradableH3 and H4 reproducing does not make them tradable — Paper 02's arithmetic is precisely what stands in the way.
  • Returns are ceilingsSlippage is modelled at zero and limit orders are assumed filled the moment the bar's high touches them.
  • Stale opensThe first printed price of a session is sometimes stale, which adds noise to the overnight leg. A source of noise, not a directional bias.
Artefacts

Reproducibility

FileWhat it is
HYPOTHESES.md The six predictions with their expected signs and the verdict rule, committed before the measurement program existed. Later corrections are appended, never rewritten over.
bench_reproducibility.py The measurement. No network, no model, no agent — reads compressed bars, emits JSON.
mcp_calls.jsonl Every OpenAIRE MCP call with its exact arguments and its raw response, so the paper selection can be replayed.
lou_polk_skouras.json Aggregate output: medians, quartiles, stock counts, verdicts.
mcp_calls_datasets.jsonl The six data-availability calls, arguments and raw responses, so the “none returned” above can be checked rather than believed.
provenance/ The exact content of the pre-measurement commit, with its SHA-256 — and a plain statement of what that does and does not prove to a third party.

Everything above is public. Repository: github.com/Alexry375/reproducibility-benchmethod · provenance · MIT (code) · CC BY 4.0 (text, results).

What is not published, and why. The underlying minute and five-minute bars come from a commercial broker feed whose redistribution is contractually forbidden. Only aggregates, code and method are released. This is a licensing boundary, not a convenience: the bench is reproducible in method by anyone holding equivalent data.

Postscript

Why an MCP server changed the shape of this

The papers above were not retrieved by keyword search and then ranked by a human hunch. We named a topic and an influence class; openaire_find_by_influence_class returned the work inside that class, ordered by field-normalised, time-adjusted citation impact. The one choice we made is a single line and it sits in the call log; which papers came back is not ours. The selection step carries a stated, auditable criterion instead of a taste. That is the part worth keeping: not that an agent can read papers, but that “which papers” stopped being an unexamined choice.

openaire_find_by_influence_class(influence_class="C3", query="intraday momentum", page_size=5)
openaire_get_research_product_details(identifier="10.1016/j.jfineco.2019.03.011")
openaire_explore_research_relationships(doi="10.1016/j.jfineco.2019.03.011", target_type="dataset")