Publishing a null result — and correcting one of our own figures
Our forecast research demo Xommodo is complete, and the final report is out. It answers the original question with no. Re-running the numbers for that report, however, turned up something that occupied us more than the null result itself: a figure we published here in August was wrong. This post corrects it and explains why, of all things, the error is the transferable part.
The correction first
The August update states that the reported total return had fallen from +135.0 % to +127.3 % once roll costs were included, accompanied by the sentence “accuracy takes precedence over the prettier number”.
The sentence is right, the number is not. +127.3 % was not the corrected value but a wrong one, slightly corrected. Roll costs accounted for roughly 8 of roughly 67 percentage points of deviation. The reliable value for the same date is +67.9 %. In the genuine forward window, the simulation also trailed behind buy-and-hold — by 10.3 percentage points over 37 trading days.
The old post stays up and now carries a dated correction notice. The wrong figure is not removed: whoever corrects a publication should leave what stood there before open to inspection.
How a verified chain could carry a wrong figure
Xommodo had safeguards for exactly this: every forecast was written into a SHA-256 hash chain before the outcome was known and anchored in the Bitcoin blockchain via OpenTimestamps. That chain was never broken. It remains verifiable to this day, and its record is accurate.
The error lay in three gaps beside it:
- What was secured was not what was published. The quality gate that the chain recorded alongside measured a different strategy from the one whose curve stood on the page. A model change improved one quantity and let the other collapse — so the gate stayed green, entirely correctly.
- What was measured was not what was shown. The published curve was recomputed every night, but from stored decisions. A changed configuration came through; a changed model could not. The display stayed internally consistent and became wrong on the outside — the least favourable combination, because internal consistency is precisely what a check controls.
- The guard could not take hold. A test was meant to prevent exactly this class of error. It searched the source code for a string — and the string was there. It ran green the whole time.
What follows for verification systems
Tamper-evident proof establishes what was recorded, and when. It does not establish that a metric derived from it still belongs to it, and it says nothing about which quantity a release check measures. Anyone building auditability should design for both gaps from the outset:
- Reconcile derived metrics against a fresh derivation — regularly, automatically, with an alarm above a threshold. We did not have this reconciliation; it would have been cheap.
- Check what the gate is actually looking at. A green signal is worth only as much as the agreement between the quantity measured and the quantity shown.
- Test behaviour, not source code. A test that searches for a string checks the spelling, not the effect.
And the null result
The original question — can weak, publicly available leading signals predict commodity prices? — is answered with no. Across ten pre-registered runs with 7,825 tested combinations, no candidate survived the pre-specified criterion. The final test missed its previously set threshold by 0.023; the decision rule, likewise written down in advance, prescribed discontinuation for that case, and so it was decided.
The report also contains two methodological findings that reach beyond the project: a placebo control we had deliberately built in passed the Bonferroni correction — which invalidated our first screening phase — and a synthetic control subsequently separated a broken method from a merely miscalibrated one.
Read the full report on xommodo.com →
The report on xommodo.com is the German version. The English version of record, with DOI and the full reproduction package — raw forecasts, ledger with timestamp anchors, pre-registrations — appears on Zenodo.
Xommodo was a research demonstration, not investment advice. All return figures quoted come from simulations or from a nine-week forward window and are no indicator of future results.