A history of the probability integral transform (with some bias/variance tradeoff)
Push an observation through its own cumulative distribution and it becomes uniform. Pull a uniform back through an inverse cdf and it becomes anything you like.
That pair of moves is about a century old and it keeps finding new work: significance testing, simulation, dependence modelling, forecast evaluation, generative models, and lately streaming prediction and anomaly detection. Here is the family tree as we understand it. Some of the later history is told firsthand, which is a source and a bias at the same time. Corrections welcome: open an issue.
Try it: every distribution is a costume
Samples fall from the left distribution, pass through its own cdf to become uniform, then out through the right distribution's inverse cdf. Swap either costume; the middle histogram never changes. (The sampler itself is the reverse move: uniforms pulled through an inverse cdf. The two moves of the whole timeline, live.)
This is why continuous univariate random variables are inherently the same object: any two are joined by a monotone map through the flat middle. For the time-series version, where F is a live forecast and the map becomes a causal bijection on whole paths, see the Rosenblatt demo at skaters.microprediction.org.
Monge poses the transport question: move a pile of earth to a target shape at least cost. Two centuries on, the probability integral transform turns out to be a transport map, and its one-dimensional form the optimal one.
The cdf becomes a working object. In Statistics by Intercomparison Galton plots the cumulative rank-to-magnitude curve of a trait and names it the ogive. Everything below is a use of this curve and its inverse.
The chi-squared goodness-of-fit test, the first general test of whether data fit a hypothesized law. Every EDF test below is an attempt to beat it by working on the transformed scale instead of binned counts.
Flood-frequency analysis plots ordered observations at positions approximating F, the plotting-position idea that puts the empirical cdf on paper a generation before the transform is named.
Tables for Statisticians and Biometricians. The cdf becomes computable: using F in practice meant looking it up, and Pearson’s laboratory industrialized the tables that made the ogive a working tool rather than a picture.
Fisher's square-root approximation: for a chi-squared X, sqrt(2X) is roughly normal with mean sqrt(2k-1), the earliest cheap Gaussianizer, a poor man's PIT by moment matching.
The integrated-squared-deviation statistic omega-squared, computed on PIT-transformed data, the ancestor Cramer 1928 and von Mises 1931 gave and that Anderson and Darling merely re-weight.
The cube-root Gaussianizer: (X/k) to the one-third power of a chi-squared is close to normal, the famous fixed-exponent stand-in for the exact PIT.
The implicit era. A p-value is the PIT of a test statistic under the null. Fisher never quite says so, but his machinery leans on it: significance levels compare across experiments because null p-values are uniform, and his rule for combining independent tests is a fact about transformed uniforms, minus twice the sum of the log p-values is chi-squared1. His 1930 fiducial argument runs the transform in reverse.
The transform gets a proper job. Pearson's Biometrika paper puts the probability integral to work on goodness of fit, transforming a sample to uniforms and testing it there.
The empirical cdf earns trust: uniform convergence to the population law, proved twice in the same 1933 volume of the Giornale dell’Istituto Italiano degli Attuari. The license to treat Galton’s ogive as an estimate; every EDF statistic below leans on it.
The first big structural payoff. The sup|Fn − F| statistic is distribution-free: after the PIT every continuous distribution is the same one, so a single table serves them all. It becomes the KS test.
The square-root transform stabilizes the variance of Poisson counts, an early member of the transform-to-tractable family the PIT perfects.
Smirnov derives the limiting distribution of omega-squared and publishes the tables that make EDF testing practical, and the two-sample version of Kolmogorov’s statistic carries his name.
Neyman builds his smooth goodness-of-fit tests directly on PIT-transformed data, testing uniformity against smooth alternatives.
Egon Pearson writes the paper that treats the transformation as a subject in its own right and gives it its name2.
Dependence built on the uniform scale. Hoeffding standardises each margin to uniform and studies the dependence that remains, a copula in all but name, in a Berlin institute series.
The inverse direction becomes an industry. Convalescing over solitaire, Ulam wondered whether dealing the hands out and counting beats combinatorics, and thought of neutron diffusion. A 1947 von Neumann letter, circulated as Los Alamos report LAMS-551, draws neutron free paths as the log of a uniform, the first surviving use of inverse-cdf sampling3.
The variance-stabilizing 2 sqrt(x + 3/8) for Poisson counts, another cheap monotone map toward a tractable distribution.
The transform meets estimation and loses its innocence. David and Johnson show that plugging sample-estimated parameters into the cdf destroys the iid-uniform property.
The paper that named the Monte Carlo method, the home von Neumann inverse-transform sampling lives inside.
Doob conjectured and Donsker proved that the empirical process of PIT-transformed data converges to a Brownian bridge free of the law, which is why every EDF statistic has one universal null distribution and the result Durbin extends to estimated parameters.
The double-square-root refinement for Poisson counts, better near zero, a sharper variance stabilizer.
A second David and Johnson paper, on the PIT when the variable is discontinuous: at an atom F(X) is no longer uniform, and the fix is randomization inside the jump or a deliberate surrogate, the mid-PIT that arrives sixty years later. The innocent transform acquires boundary conditions.
Fréchet poses the given-margins problem and proves the sharp bounds on a joint distribution with fixed marginals, now the Fréchet–Hoeffding bounds, the frame Sklar's theorem later fills.
The general statement X = F−1(U) is written down in Various techniques used in connection with random digits, alongside the rejection method as the fallback when the inverse is cumbersome. Every simulated random variable since is the inverse PIT. It is even named the inverse probability integral transform.
Three pages in the Annals. Chain the transform through conditional distributions, ut = F(xt | x<t), and any joint law becomes iid uniforms, a causal bijection between an arbitrary process and pure noise. The universal whitener.
Weighting the transformed process. Anderson and Darling substitute u = F(x) and weight the uniform empirical process by 1/(u(1−u)) to sharpen the tails, and supply usable significance points in 1954.
Normal scores: replace ranks by standard-normal quantiles and test there. The rank-to-Gaussian move, the PIT’s flat middle re-dressed as a bell, and a bridge from distribution-free ranks to Gaussian machinery that Chernoff and Savage would later show gives up nothing.
A second way to become uniform: do not transform the value, transform the clock. A submartingale splits uniquely into a martingale plus a predictable increasing process; for a counting process that increasing process is the compensator, the model-implied amount of event time elapsed. Later point-process transforms do not push the value through a cdf, they push calendar time through this compensator, turning raw observations into standard noise, Brownian or Poisson, with one last PIT to make it uniform.
First rigorous asymptotics for distance-based EDF tests when parameters are estimated, the theory companion to the estimated-parameter problem.
Minimax rational approximations to standard functions including the inverse normal cdf, the ancestor of every later inverse-normal fit used to turn uniforms into Gaussians.
The finite-sample license: an exponential bound on the sup-error of the empirical cdf, sharpened by Massart in 1990 to its clean constant. Confidence bands for F itself, and non-asymptotic control for everything built on the transformed scale.
The transport reading. Knothe builds a triangular map carrying one multivariate law onto another by sequential conditioning, the same object as Rosenblatt's transform, now called the Knothe–Rosenblatt rearrangement. The PIT is a transport map, and in one dimension it is literally the optimal one4.
Computing Φ−1 was hard enough that Box and Muller built a method to sample normals without it.
Sklar's theorem: every joint distribution is a copula applied to its margins, unique when the margins are continuous, so dependence becomes an object you model apart from the margins. The three-page note appeared in Paris, in French, at Fréchet's request, three theorems stated without proof, and coined "copula" from the grammatical term for a linking word.
The copula goes underground for two decades. Schweizer and Sklar develop copulas inside probabilistic metric spaces, where they mostly stay until the 1980s.
Low-discrepancy sequences fill the cube more evenly than random points, and mapped through the inverse cdf they Gaussianize into quasi-Monte Carlo.
Bivariate normals are asymptotically independent in the tail, the observation that seeded the tail-dependence coefficients and the hunt for copulas that keep their tails.
Distribution-free independence tests built from sample distribution functions: Hoeffding’s margin-standardisation made an operating programme, dependence as an object on uniform margins ahead of the copula boom.
U-squared, the rotation-invariant EDF statistic for circular and periodic data, the omega-squared idea on the circle.
The Poisson process gets its martingale certificate: if a unit-jump counting process has N(t) minus t as a martingale, it is unit-rate Poisson. The jump analogue of Levy characterizing Brownian motion as the canonical continuous martingale on its own clock.
The inverse PIT reaches finance. Hertz's Harvard Business Review paper on risk analysis in capital investment samples uncertain inputs from specified distributions, Monte Carlo by inverse transform, in corporate decision-making.
The adaptive power family (x^lambda - 1)/lambda chosen by likelihood to Gaussianize, the parametric generalization of the fixed-exponent root transforms.
The polar method, Box-Muller without the trigonometry, another route from uniforms to normals.
The continuous-path PIT. A continuous local martingale is Brownian motion evaluated at its own quadratic-variation clock. Calendar time is the wrong coordinate and accumulated variance is the canonical one: the scalar PIT says use the probability scale, this says use information time.
The order-statistic normality test: correlate the ordered sample with expected normal order statistics. Quantile structure as a test statistic, and for decades the default answer to “is it normal?”
A multivariate exponential from a fatal-shock model, whose dependence is the singular Marshall-Olkin copula, a genuine construction predating Sklar’s uptake.
The Sobol sequence, the workhorse low-discrepancy set, fed through the inverse cdf for quasi-Monte Carlo.
Monte-Carlo significance points for the KS normality test when mean and variance are estimated, the practical fix for the estimated-parameter problem.
Residuals for any parametric model. A General Definition of Residuals pushes each observation through its own fitted distribution, and in survival analysis the Cox–Snell residual is a PIT residual, unit exponential after a minus log when the model is right5. Model criticism by transformed residuals starts here, and is what the parade does one tick at a time.
The transform gets its picture. The Q–Q plot reads calibration off a straight line, the diagnostic descendant of the PIT that every forecaster still squints at.
Quantile mapping is born. Correcting one distribution to match another by pushing it through Fsource then Ftarget−1 is inverse-PIT composition, and climate and hydrology have leaned on it ever since.
The PIT learns to move time. For a point process with conditional intensity, transform event times by the compensator, the integral of that intensity, and under a correct model the result is unit-rate Poisson. The gaps in compensator time are iid exponential, and 1 minus their exponential is iid uniform, Rosenblatt for arrivals rather than values.
Weak convergence with estimated parameters, the definitive limit theory for the empirical process when the cdf is fitted, not known. The estimated-parameter problem carries his name, and every model that scores its own residuals inherits it.
EDF statistics for goodness of fit: the significance points and power comparisons that made D, W-squared, U-squared and A-squared standard equipment on transformed data.
Histogram equalization, formalized in image processing in the mid-1970s (Hummel and others), maps each pixel through the image's own cumulative histogram. That is the forward PIT applied to intensities, and the output comes out flat for exactly the reason the transform makes uniforms.
Histogram specification proves the monotone gray-level map carrying one image histogram onto any target is unique, which is inverse-PIT composition, the matching complement to equalization.
Kimeldorf and Sampson reinvent copulas independently as uniform representations, unaware of the metric-space lineage, a sign of how obscure the idea still was.
The score the forecasting world lives by is born far from it, in Management Science. The CRPS is the Brier score of the forecast cdf integrated over every threshold, and its reliability component reduces to rank-histogram, that is PIT, information.
Options: A Monte Carlo Approach prices derivatives by simulating asset paths from the risk-neutral law, inverse-transform sampling on Wall Street twenty-three years before the copula work. This is where the inverse PIT reaches finance in earnest.
Algorithm AS 111, the rational approximation to the normal quantile that inverse-transform pipelines shipped for a generation, later given a better tail by Moro and superseded by Wichura.
Random time changes enter multivariate counting-process statistics: subtract the compensator for a martingale, or run on compensator time for Poisson noise. The hazard scale becomes the PIT scale for lifetimes and recurrent events.
The empirical copula, the rank-based nonparametric estimator of the dependence function, copulas built straight from the PIT’s empirical margins.
Latin hypercube sampling: stratify each uniform into equal-probability bins, one draw per bin, then push through the inverse cdf. Stratified inverse PIT, and a staple of computer experiments ever since.
Bayesian model criticism. Box judges a model by where the data fall in its predictive distribution, the prior-predictive PIT check.
Schweizer and Wolff build measures of dependence on the copula alone, invariant to monotone transforms of the margins.
Probability plotting renders the PIT graphically: order the failure times, assign plotting positions approximating F, and a correct model falls on a straight line.
Rank-correlation induction: impose a target dependence on samples with arbitrary margins by reordering them, dependence carried by the ordering of the uniforms, a copula idea in simulation clothing sixteen years before the credit copulas.
The well-calibrated Bayesian: a coherent forecaster must expect to be calibrated, the binary-event foundation the PIT extends to full distributions.
The KS test becomes an everyday tool in astronomy for comparing luminosity functions and source counts, with the standing caution, from Peacock, that its distribution-free property fails in more than one dimension.
The forecast-evaluation program, stated in full and mostly overlooked. The prequential approach has a section literally titled Probability Integral Transform, proposing that a sequential forecaster be judged by whether Un = Fn(Xn) looks like an iid uniform sample, citing Rosenblatt. This is what modern density-forecast checking does, fourteen years early6.
Posterior predictive checks: simulate replicated data and ask whether the observed data are typical, the PIT pointed at the fitted model rather than a fixed one.
The Joy of Copulas finally sells the subject to statisticians, and it starts spreading beyond the metric-space corner.
The fourth root of a chi-squared is close to normal, not uniform: Y = X^c is Weibull with shape 1/c, and the zero-skew shape sits near 3.60, so c near 0.277, with one quarter the convenient round number beside it. The obscure sharper sibling of Wilson-Hilferty, and the transform spectral analysts use to normalize periodogram ordinates, which are asymptotically exponential.
Non-Uniform Random Variate Generation, the canonical reference for turning uniforms into everything else. He calls the inverse-cdf recipe the inversion method, and the book is the field’s dictionary.
Algorithm AS 241 gives the inverse normal to sixteen digits, closing the accuracy problem that had motivated Box and Muller thirty years earlier. It is the qnorm behind most software.
A simple proof of the multivariate random time-change theorem, the clean statement that a point process rescaled by its compensator is unit-rate Poisson, the theorem Ogata turns into earthquake residual analysis the same year.
ETAS residual analysis rescales earthquake times by the integrated conditional intensity so a correct model yields a unit-rate Poisson process tested for uniformity, the PIT for point processes.
The smooth-test revival puts Neyman's 1937 PIT-based framework back to work, showing many classical goodness-of-fit statistics are smooth tests or their components.
The optimal transport map for quadratic cost exists, is unique, and is the gradient of a convex function, the general-dimension theorem whose one-dimensional case is exactly the monotone PIT.
Identifies and fits bivariate Archimedean copulas through the Kendall distribution of the generator, the standard selection tool.
Turns prequential calibration into tests of sequential forecast validity, the testing theory between Dawid 1984 and later PIT tests.
The empirical PIT on trading desks. J.P. Morgan's RiskMetrics Technical Document popularises historical-simulation VaR, which reads a loss quantile straight off the empirical cdf.
Drops Brenier’s moment assumptions and frames the optimal map as the monotone measure-preserving map, the clean generalization of the monotone rearrangement the PIT instantiates.
The Beasley-Springer inverse normal with a better tail, the Gaussianizer that became the finance and quasi-Monte Carlo default.
Randomized quantile residuals: the fitted cdf at each response, then the normal quantile, with randomization for discrete data, exactly-uniform-then-normal residuals, the discrete PIT that GLM diagnostics rest on.
Posterior predictive p-values, the everyday Bayesian model-check, a PIT statistic on realised discrepancies.
Weather forecasting reinvents the PIT as the rank histogram, also called the Talagrand diagram: where the observation falls among the sorted ensemble members is uniform when the forecast is calibrated. Flat is the goal.
Joe builds multivariate distributions from bivariate blocks, the first pair-copula construction, which is Rosenblatt's transform used to build distributions rather than test them.
Forecast evaluation reaches econometrics. Diebold, Gunther and Tay make the PIT histogram the standard check for density forecasts, crediting Rosenblatt.
Frees and Valdez carry copulas to actuaries, the route by which they reached credit modelling.
Conditional-coverage and independence tests for interval forecasts, the calibration test underneath VaR backtesting.
A copula construction in credit: a hybrid default simulation at Morgan Stanley that nested both a copula and a doubly stochastic intensity, of which the copula approach later made famous in publication is the special case7.
The multivariate extension: apply the Rosenblatt ordering to joint density forecasts and check the transforms for uniformity.
An Introduction to Copulas, the textbook that made the subject teachable and its tail-dependence and rank-correlation vocabulary standard.
Decomposes the CRPS so its reliability term is the rank histogram, formally uniting the scoring-rule thread and the PIT-histogram thread.
The ziggurat method, the modern speed champion for sampling normals, a rejection method and the enduring alternative to inversion.
A Box-Cox-style power family extended to the whole real line, so signed and zero data can be Gaussianized too. It ships as a transform in skaters.
A call to action for the PIT. Testing Density Forecasts gaussianizes the transforms by the inverse normal, z = Φ−1(PIT), so a correct density yields iid N(0,1) and ordinary tests of mean, variance and autocorrelation grade it. Berkowitz handed the field a clean way to score its forecasts and act on the score.
A flat rank histogram is necessary but not sufficient: compensating conditional biases can average to uniform. The caveat that keeps PIT histograms from being over-read.
Correlation and dependence in risk management, the paper that separated linear correlation from copula dependence for practitioners.
The time-rescaling theorem: rescale spike times by the integrated conditional intensity and any point process becomes unit-rate Poisson with uniform interevent times, checked by KS, the same integrated-hazard construction as the credit-default nesting above.
Vines organize pair-copula decompositions into a graphical model, sequential conditioning made into a structure.
The first applied copula goodness-of-fit test built on the Rosenblatt transform as the PIT, tested on high-frequency finance.
The energy distance, in one dimension twice the Cramer squared-L2 gap between cdfs, the population object the CRPS and the energy score estimate. It is a kernel quantity, related to but distinct from the Wasserstein transport distance.
Conformal prediction. Under exchangeability the rank of a fresh nonconformity score is uniform, the PIT without the continuity assumption, giving finite-sample-valid p-values, which the field then narrowed to coverage intervals.
The KS statistic reaches biology. Gene set enrichment analysis, among the most cited methods in the field, runs a weighted Kolmogorov–Smirnov running-maximum-deviation statistic to test whether a gene set clusters at the top of a ranked list.
Validate Bayesian software by checking that posterior quantiles of the true parameters are uniform, a PIT on the computation, the precursor to simulation-based calibration.
Strictly proper scoring rules formalized, the propriety theory under every PIT-based evaluation and the home of the CRPS-as-Cramer-distance identity.
Calibration and sharpness: maximise sharpness subject to calibration, with the PIT histogram as the diagnostic, and the fair dual credit to Dawid 1984 and Diebold 1998.
Copula goodness-of-fit built directly on the Rosenblatt transform: a vector has copula C if and only if its Rosenblatt transform is the independence copula, the multivariate PIT turned into a test.
Pair-copula constructions made practical for inference, among the most cited papers in the field, and Rosenblatt's transform used constructively at scale.
The mid-PIT for count data, the non-randomized fix that keeps discrete-data calibration checks centred, the same correction the sticky lattice path produces for free elsewhere.
Optimal Transport: Old and New gives the Knothe-Rosenblatt rearrangement its textbook home, a named section presenting the triangular coupling to the transport community.
The higher-dimensional bridge. They prove that the Knothe–Rosenblatt map is the limit of Brenier optimal-transport maps as the quadratic cost anisotropically degenerates, so the conditional PIT is optimal transport in a precise limiting sense even where it is not the Brenier map outright.
Esteban Tabak introduces the flow of maps with Eric Vanden-Eijnden (2010) and names it with Cristina Turner (2013): a learned invertible map to a simple base density, trained by exact change of variables. Deep learning then rebuilds Rosenblatt at scale, mostly without knowing it. Rezende and Mohamed popularize the idea for inference, and the autoregressive flows MAF and IAF parameterise conditional cdfs, the Rosenblatt transform with neural networks for the conditionals. None of the founding deep-learning papers cite the 1952 result; the link was drawn retrospectively by the 2021 survey8.
An independent second proof that Knothe’s rearrangement is the degenerate limit of Brenier maps, by a continuation equation for the Kantorovich potential.
Probabilistic Forecasting, the review that fixes the modern vocabulary: proper scores, calibration, sharpness, and the PIT as the working diagnostic.
NICE: the first practical deep invertible density model, additive couplings with a free Jacobian, trained by exact change of variables. The flow era’s engineering begins.
The modern treatment of quantile mapping for climate bias correction, with the warning that naive mapping can corrupt projected trends.
Sampling via measure transport: the program that made the triangular Knothe-Rosenblatt map the practical object for inference, the line normalizing flows grew from.
Real NVP scales the idea: affine coupling layers with cheap Jacobians make exact-likelihood invertible models large enough to matter.
Simulation-based calibration: rank statistics of prior draws within posterior samples are uniform when the computation is correct, the PIT turned into a test of the algorithm rather than the forecast.
Recalibrates a neural network’s predictive cdf at the observation to be uniform by isotonic regression, the deep-learning entry point for PIT recalibration.
Particle physics names the transform outright: Cousins' lectures have a section titled Probability integral transform, recasting any goodness-of-fit test as the claim that the transformed sample is uniform.
Glow adds invertible 1×1 convolutions and takes flows mainstream: large images, exact likelihoods, samples people shared.
FFJORD: the continuous-time limit, invertible dynamics as an ODE with likelihood by an unbiased trace. Transport as literal flow, on a page already full of time changes.
Distribution-free predictive inference for regression, the split-conformal foundation. Coverage rests on the uniformity of conformity-score ranks, the PIT idea in rank clothing.
Conformalized quantile regression adapts interval width to the input, the stepping stone between split conformal and fully distributional conformal.
Visualization in Bayesian workflow makes the LOO-PIT plot a standard check: leave one observation out, transform it by the predictive cdf, and look for uniform.
Neural spline flows sharpen the univariate monotone map itself: rational-quadratic splines as learned cdfs inside the couplings. Of the flow papers, the one closest in spirit to this page: better quantile functions, learned.
The transform becomes a market mechanism in a high-velocity distributional prediction market run by Intech Investments9: data pushed through a community implied cdf and then the inverse normal gives z-streams that are themselves predicted10.
Distributional conformal prediction conformalizes the estimated conditional-cdf rank, stating outright that the conditional rank U = F(Y,X) is uniform and independent of X, the PIT made explicit in conformal prediction.
YOU ARE HERE. The package emphasized online distributional residuals via the "parade" book-keeping, and was the original home for prediction algorithms since moved to skaters.microprediction.org.
States verbatim that autoregressive flows are Knothe-Rosenblatt couplings and proves consistency of estimating them, the precise statement the flow literature had left implicit.
A large-scale study that writes Z = F(Y|X) as the probability integral transform, proves uniform PIT equals probabilistic calibration, and recalibrates 57 neural regression datasets to it.
The representation and learning theory for monotone triangular transport maps, the parametric Knothe-Rosenblatt maps that autoregressive flows approximate.
The height between a conformal band and a full predictive distribution gets a name and a closed form, the residual-information gap (Marginally Useful). A three-line lemma thus kills conformal prediction as a field11.
skaters ships laplace, an online forecaster built around Rosenblatt's transform with its own predictive as the conditional law, tick by tick. The same bijection is offered as a change of coordinates for other models, and in the benchmarks run so far it has improved most of those tried12.
The detection head wald tests each tick's standardized surprise against its own tail, so the advertised false-alarm rate is the measured one13. A century of the same idea, now running live.
Footnotes
1 The combining rule is universally cited as Fisher 1925, but it entered Statistical Methods for Research Workers in the fourth edition of 1932; Karl Pearson's 1933 paper cites that edition contemporaneously. See the tracing.
2 Two Pearsons, routinely conflated: Karl (1933) invented the transform-to-uniform test; his son Egon (1938) coined the phrase "probability integral transformation". There is no Karl Pearson 1938 paper; he died in April 1936.
3 Priority is layered. Eckhardt's 1987 memoir credits the idea to Ulam; the earliest surviving written use is the March 1947 von Neumann letter (report LAMS-551); the first published general statement is his 1951 Various techniques.
4 Precise version: the PIT always pushes the data law onto the uniform, so it is a transport map. It is literally the Brenier optimal map only in one dimension, where for any strictly convex cost the monotone T = Ftarget−1 ∘ Fsource is optimal. In higher dimensions the conditional PIT is the triangular Knothe-Rosenblatt map, optimal only as the degenerate limit of Brenier maps (Carlier, Galichon and Santambrogio 2010).
5 Journal of the Royal Statistical Society Series B 30(2), pages 248-265. The range 248-275 copied through much of the literature is wrong.
6 PIT-based density-forecast evaluation is routinely credited to Diebold, Gunther and Tay 1998 alone. Dawid 1984, Section 5.3, had already stated the program and cited Rosenblatt; Gneiting, Balabdaoui and Raftery 2007 give the correct dual attribution.
7 The inverse PIT had been in finance for decades by 1998; the new part was coupling default times. The mechanism, which is where the two methods nest: exponential thresholds correlated by monotone transform to normal, the copula piece, then a default triggered when the integral of a stochastic hazard rate first exceeds the threshold, the Cox or doubly stochastic intensity piece. That timing half is exactly the point-process time-rescaling of Meyer and Papangelou (1969-1972 above), independently rediscovered for credit. The special case is David X. Li, On Default Correlation: A Copula Function Approach, RiskMetrics working paper with a first draft of September 1999, published in The Journal of Fixed Income 9(4), 2000, and it had been reached independently before: MacKenzie and Spears, documenting the era from 114 interviews, record Vasicek's unpublished correlated-default memoranda of 1987 to 1991 and CreditMetrics arriving at the Gaussian-copula form in 1997 before the copula terminology was attached. Those were the copula piece alone, not the copula-plus-Cox nesting.
8 Papamakarios, Nalisnick, Rezende, Mohamed and Lakshminarayanan, Normalizing Flows for Probabilistic Modeling and Inference (JMLR 2021). A trap: Tabak and Turner 2013 cite Rosenblatt 1956, the kernel-density paper, not the 1952 transform.
9 The mechanism was written down only in 2026: a nearest-the-pin parimutuel that pays forecasters by proximity to the outcome, with a formalized noise addition to the target that makes the game incentive compatible. Paying by proximity rather than for backing a single bucket is what makes a crowd report a full predictive cdf instead of a point, and the calibrated noise is what makes truthful reporting optimal. Formalizing it closes the loop the platform ran by instinct, and connects the PIT to market design.
10 The community implied cdf is the distribution implied by everyone's quarantined predictions at two horizons, roughly a minute and roughly an hour. Feeding the resulting z-streams back as new prediction targets is what lets the crowd sharpen its own coordinate system, across billions of micropredictions. It answered Berkowitz's 2001 call to action at a velocity he never imagined, and unlike conformal prediction it allows sharp predictions crossing the information gap.
11 Tongue in cheek. Nobody here is out to kill conformal prediction, least of all the author, who strengthens it in a separate result, the width of the conformal fan. The lemma is an argument for finishing the job with full distributions, not for abandoning finite-sample validity.
12 On held-out data, forecasters run on the transformed stream and scored back by exact change of variables gain around two nats per point; anomaly detectors gain several times their raw hit rate. It is the forecaster passing through, not the opponent.
13 Alarm on p below alpha and the false-alarm rate is alpha by construction, no threshold to tune. On real anomaly-free streams wald is the best-calibrated of the streaming detectors measured at operating depth.