Sundial vs Laplace
Sundial is a generative time-series model; we draw samples per step and score the empirical predictive. Run raw only in this study; the collaborative arms were built for the three strongest models first.
Resources: GitHub · Model card · Paper (arXiv 2502.00816)
Live snapshot Derived from the week-long round-robin study; counts grow as coverage deepens. Everything scores against Laplace on the identical series and windows.
Standalone
The model's own predictive, run zero-shot, scored per series against Laplace by a one-step-ahead (k=1) Diebold–Mariano test on the log-score differential.
| stratum | n | win / draw / loss vs Laplace | med ΔLL | CRPS ratio | cov₀₀ |
|---|---|---|---|---|---|
| economic change-series, business-daily | 1328 | -2.77 | 1.078 | 0.64 | |
| economic change-series, weekly | 2720 | -2.56 | 1.022 | 0.68 | |
| economic change-series, monthly (annual cycle) | 2720 | -3.02 | 1.014 | 0.56 | |
| M4 hourly, strongly seasonal | 414 | -3.41 | 1.088 | 0.52 | |
| asset prices and returns, daily | 1392 | -2.35 | 1.095 | 0.59 |
Median per-series Δ log-likelihood in nats (negative is worse than Laplace); CRPS ratio to Laplace (above 1 is worse); raw central-90% coverage (0.90 target).
Star plot
Sundial standalone against Laplace, on the same six regime axes as the site's
standalone radar. Each radius is the log-likelihood
ratio, (wins + ½·ties) / n scaled so an even split with
Laplace sits on the dashed 1.0 ring; outward beats Laplace more often, inward less. The
M4-hourly set splits into soft and hard waveforms by corpus order, matching that radar.
Collaborative use
The recalibration (@lap) and portfolio (&lap) arms were built for the three strongest models first, so Sundial is scored standalone here. The sidecar wraps any per-step predictive, so these arms can be added without retraining. See the sidecar pattern.
Protocol
Fixed 128-length context, rolling one-step test window, no fitting, each model in its own environment. Strata split the cached FRED universe and the M4-hourly set by frequency and regime. Full method on the sidecar pattern page and in the methodology.
Architecture and methodology
Sundial is a decoder-only transformer trained without discrete tokenisation. It embeds continuous-valued patches and learns a generative TimeFlow head by flow-matching, so instead of a parametric density it produces samples of plausible futures. This study draws those samples (the base-128m model) and scores their empirical predictive; it is pretrained on the trillion-point TimeBench corpus.