Distributions
    Distribution

    Nothing to draw: every distribution is hidden or has an error.

    Distributions

    The first draws

    Data

    Fit
    Families
    Method
    Rank by
    Binomial trials n
    Which families

    The fits, best first

    The fits against the data

    Or no family

    the data as a distribution of their own:
    Probabilities
    x between a and b
    Percentiles and intervals
    percentiles
    interval holding %
    Tails and expectations
    threshold t level %
    Two distributions
    X Y
    Monte Carlo of an expression

    Draws the distributions an expression names, by their letters, and makes the results a distribution of their own, which stays up to date with them.

    expression
    draws seed
    scheme

    Probability distributions

    This page defines probability distributions, draws them together, lists their summary statistics, samples them, fits them to data and calculates with them. Everything is computed in the browser; nothing you enter leaves it. The distributions, the data and the settings are kept in this browser's storage, so they are there on your next visit; start again clears them.

    • The distributions
    • Parameterisations
    • Known values
    • Shift
    • Truncation
    • Empirical distributions
    • Tables
    • Monte Carlo
    • The chart
    • Statistics
    • Samples
    • Fitting to data
    • Ranking the fits
    • Calculations
    • The families
    • How the numbers are computed
    • References

    The distributions

    The list on the left holds up to 16 distributions. Each has a letter (A, B, …), a colour and a name you can change. Click a distribution to edit it below the list; its coloured dot shows or hides it in the charts, and × removes it. Duplicate copies the selected distribution, which is the quick way to see what a change of one parameter does. The colours are a palette checked for colour-blind readers; the ninth to sixteenth distributions repeat the eight colours with dashed lines.

    A distribution is one of four kinds, chosen in the family list:

    • a parametric family: 30 continuous ones (normal, lognormal, gamma, Weibull, beta, triangular, the extreme-value families and more) and 8 discrete ones (binomial, Poisson, negative binomial and more), listed under The families;
    • an empirical distribution made from data, with no family assumed;
    • a table typed in: cumulative probabilities, or the probability of each value;
    • a Monte Carlo result: an expression of other distributions, drawn many times.

    Changing the family of a distribution keeps what can be kept: the new family starts at the same mean and standard deviation when it can have them, and at its own default otherwise.

    Parameterisations

    Most families can be given in more than one way. A lognormal, for example, by μ and σ of ln X, by the mean and SD of log₁₀ X, by its geometric mean and geometric SD, by its arithmetic mean and SD or mean and CV, by its median and error factor (the 95th percentile over the median) or by a central interval and the probability it holds. Choosing another parameterisation converts the distribution as it is, so nothing moves until a number is changed. Every parameterisation converts exactly, both ways: a round trip returns the numbers it started from.

    A field turns red when its number is not allowed (a negative SD, a probability above 1), and a message under the fields says why when the numbers together are impossible (a lognormal whose mean is below its median, a triangle whose mode lies outside its ends).

    Known values

    Every continuous family can also be given by known values: as many properties of the distribution as the family has free parameters. A property is the mean, median, mode, standard deviation, variance, coefficient of variation, a percentile, the geometric mean or the geometric SD, in any combination: a gamma with a median of 4 and a 95th percentile of 20, a triangular distribution through three percentiles, a Weibull with a given mean and CV. The page searches for the parameters (Nelder–Mead from several starting points, then Newton's method) until every value is met to about twelve significant figures. When no member of the family has the values, the distribution is not drawn until they are possible, and the message names the nearest member and the values it has: a lognormal cannot have a mean below its median, nor an exponential a CV other than 1.

    Parameters that a family keeps fixed while solving, such as the ends of a beta or the λ of a PERT, stay at the values the distribution had.

    Shift

    Any distribution can be shifted: X + s, every value moved along x by s. The mean, the median, the mode and every percentile move by s; the standard deviation, the skewness, the kurtosis and the entropy stay as they are, and the geometric mean and SD are worked out again. A distribution on whole numbers takes a whole-number shift, so that it stays on whole numbers. The shift comes before a truncation, whose bounds are values of the shifted distribution, and known values describe the shifted distribution: a gamma shifted by 100 with a mean of 105 and an SD of 3 is a gamma of mean 5 and SD 3, moved. A change of family keeps the shift, and the mean and SD. A fit has no shift.

    Truncation

    Any distribution can be truncated to an interval: X conditioned on lower ≤ X ≤ upper, with the density raised so that it again integrates to one. Either bound can be left empty. The probability the interval held before truncation is shown under the bounds. For a discrete distribution the bounds are included, and for one on whole numbers they are rounded inwards: a lower bound of 2.5 starts at 3.

    The bounds can be given as values or as percentiles of the distribution before truncation: 5 and 95 keep the middle 90 % of a continuous distribution, and a percentile of 0 or 100 is no bound. The value a percentile cuts at is shown under the bounds; for a discrete distribution it is the smallest value whose cumulative probability reaches the percentile, and that value is kept. Switching between values and percentiles keeps the bounds where they are.

    A truncated distribution has every statistic recomputed numerically. Truncating a heavy tail gives moments that did not exist before: a Cauchy truncated to [−5, 5] has a variance.

    Empirical distributions

    An empirical distribution is made from data without choosing a family, in one of four ways:

    • Interpolated: the cumulative curve through the sorted values, rising by 1/(n − 1) from each to the next. Its percentiles are the usual sample percentiles (type 7, the default of R, NumPy and Excel's PERCENTILE.INC), and its density is constant between neighbouring values. It never goes below the smallest value or above the largest.
    • Kernel density: a normal curve of width h (the bandwidth) on every value. Silverman's rule, h = 0.9 min(SD, IQR/1.34) n−1/5, or Scott's, h = 1.06 SD n−1/5, or a width you give. In ln x puts the kernels on the logarithms, for positive, skewed data: the result has no probability below zero. Above 2,000 values the kernels are placed on 4,096 bins, which changes nothing visible while the bandwidth spans several bins; the moments are always exact.
    • Histogram: equal-width bins from the smallest value to the largest, by Sturges' (1 + log₂ n bins), Freedman–Diaconis' (width 2 IQR n−1/3), Scott's (width 3.49 SD n−1/3) or the square-root rule, or a count you give; the probability of a bin is spread evenly over it.
    • Empirical CDF: probability 1/n on each value, the sample itself as a discrete distribution.

    The quickest way to make one is the Or no family row of the Fit tab, which takes the data entered there. The data can also be pasted into the distribution itself.

    Tables

    A cumulative table is a list of values with their cumulative probabilities, one pair per line ("x F"), straight lines between them: the way an expert's percentiles are often written down. F must start at 0, end at 1 and never decrease. A probability table lists values with their probabilities ("x p"), a discrete distribution; probabilities that do not add up to one are scaled so that they do, and the page says so.

    Monte Carlo of an expression

    The Monte Carlo kind draws the distributions an expression names, by their letters, and evaluates the expression draw by draw: A * B, exp(A) / (1 + B^2), max(A, B, C). The expression may use numbers, the letters, + − * / ^ (or **), parentheses, pi, e and the functions exp, ln (or log), log10, sqrt, abs, floor, ceil, pow, min and max. The distributions are drawn independently of each other, each from its own stream of a seeded generator (xoshiro128**), so a seed gives the same result every time, and by one of the sampling schemes of the Sample tab: simple random, Latin hypercube (each letter's draws one from each of n equally probable strata, in random order), centred Latin hypercube, or the Sobol or Halton sequence, in which the letters are the dimensions, so that their draws together fill the space of their probabilities evenly and the result's mean and percentiles settle much faster than with random draws. A simple random or Latin hypercube draw of a letter is the Sample tab's draw of it with the same seed. Draws whose result is not a finite number are left out and counted.

    The result is an empirical distribution of the draws, made the way you choose (kernel density by default), and it is recomputed whenever a distribution it uses changes. An expression may use another Monte Carlo result, but not itself, directly or through others.

    The chart

    The chart draws every visible distribution together, one chart above and, if chosen, a second below it: the density (or, for a discrete distribution, the probability of each value), the cumulative probability F(x) = P(X ≤ x), the exceedance probability 1 − F(x), the quantile function, or the hazard rate f(x)/(1 − F(x)). Two charts rather than one with two vertical scales, because two scales on one plot suggest a relation between the curves that is not there.

    A discrete distribution is drawn as stems on the same axis as the densities. Its probabilities and a density are comparable when the values are one apart, as counts are.

    With log x the density is drawn per decade, x f(x) ln 10, so that the area under each curve on the logarithmic axis is still its probability and a lognormal is a symmetric bell. Log y is the way to read tails, in the exceedance chart above all.

    The x-axis covers the 0.1st to the 99.9th percentile of every visible distribution unless other percentiles or values are given. Data draws the data a distribution carries (a fitted distribution carries its data) as a histogram behind its density and as steps beside its cumulative curve; samples draws the draws of the Sample tab the same way, with dotted outlines. Pointing at the chart shows the value of every curve at that x; clicking a name in the legend hides or shows that distribution, as its dot in the list does. bins sets the histograms: a rule (Freedman–Diaconis, the default, Scott, Sturges, square root or Rice), a number of bins, or a bin width, in decades on a logarithmic x-axis, with the edges at multiples of the width so that histograms side by side share them; the Fit tab's density plot has a bins of its own. Under the chart are the mean, SD, median, mode and 5th and 95th percentiles of every distribution it shows, the selected one in bold, and with samples ticked those of each one's draws below it. Save the curves as CSV writes the values drawn.

    Statistics

    The Statistics tab lists every distribution side by side: its parameters, support, mean, median, mode, standard deviation, variance, coefficient of variation, skewness, excess kurtosis (zero for a normal), entropy (differential entropy, in nats, for a continuous distribution), geometric mean and geometric SD (for a distribution on positive values), interquartile range, median absolute deviation and the percentiles you list. A distribution carrying data also gets its sample size, log-likelihood and Kolmogorov–Smirnov D against its data.

    A moment that does not exist is shown as ∞ when one tail makes it infinite and the moments below it are finite (a Pareto with α = 3 has an infinite skewness), and as undefined when both tails do, or when a lower moment is already infinite (a Cauchy has no mean; a Pareto with α = 1 has an infinite mean and so no variance). A distribution of a single value has no skewness or kurtosis. Closed forms are used where a family has them; otherwise, and always after truncation, the moments, entropy and logarithmic moments are integrals over the quantile function, computed numerically.

    Samples

    The Sample tab draws n values of each distribution you choose. Every draw is an inverse transform: a number u between 0 and 1 becomes the value at probability u, read from the upper tail when u is above one half, so that the tail keeps its digits. The scheme decides the u's:

    • Simple random: independent draws from the generator (xoshiro128**), as a Monte Carlo simulation takes them. The mean of n draws is off the distribution's by about SD/√n.
    • Latin hypercube: the probabilities cut into n equal strata, one draw at a random place in each, in random order (McKay, Beckman and Conover 1979). Every part of the distribution is drawn from in proportion, the tails included.
    • Latin hypercube, centred: the same at the middle of each stratum, so that the draws are the quantiles at (i − ½)/n, in random order.
    • Sobol sequence: Sobol's low-discrepancy sequence (1967) with the direction numbers of Joe and Kuo (2008), scrambled by a random lower triangular matrix and a random digital shift (Matoušek 1998), which keep its structure: when n is a power of two, each of n equal strata of probability holds one draw of every distribution.
    • Halton sequence: the radical inverse of 0, 1, 2, … in a prime base for each distribution (Halton 1960), the digits at each position permuted at random.

    Each distribution draws from a stream of its own, made from the seed and its letter, so its draws do not change when another distribution is added, changed or left out, and a seed always gives the same draws. A simple random or Latin hypercube sample is exactly what a Monte Carlo expression with the same seed draws from that distribution. In the two sequences the letter is also the dimension, A the first and P the sixteenth, so that the samples of several distributions together fill the space of their probabilities evenly. No correlation is imposed between distributions.

    The table sets each statistic of the draws (the mean, the SD with n − 1, the ends and the percentiles, type 7) over the distribution's own, and gives Kolmogorov–Smirnov's D between the draws' cumulative curve and the distribution's, with its p-value. A Latin hypercube or low-discrepancy sample is more even than random draws are, so its D is smaller and its p-value close to 1. With show in the chart (the same setting as the chart's samples box) the draws appear in the chart, with dotted outlines: a histogram behind each density, steps beside each cumulative curve, points along each quantile function. They can be saved as CSV or as Excel, copied for a spreadsheet, or handed to the Fit tab.

    Up to 200,000 draws in all are taken by themselves whenever a distribution or a setting changes, while the tab is open or the chart shows them; above that, Draw takes them (up to 1,000,000 of each distribution and 4,000,000 in all). They are not stored with the page: the seed brings them back.

    Fitting to data

    Paste data into the Fit tab, open a CSV or text file, or drop one on the box. Numbers may be separated by spaces, tabs, commas, semicolons or new lines; with decimal comma a comma is part of a number and only the other separators count. A file with several columns, or pasted columns, offers a choice of column; text that is not a number is skipped and counted.

    Families chooses continuous or discrete families; by the data picks discrete when every value is a whole number of zero or more, and continuous otherwise. A continuous and a discrete fit cannot be ranked against each other: one's likelihood is a density, the other's a probability.

    Maximum likelihood and the method of moments

    Maximum likelihood finds the parameters under which the data are the most probable. Closed forms where they exist (normal, lognormal, exponential, Pareto, inverse Gaussian and others), Newton's method on the likelihood equations of the gamma and the beta, Brent's method on the one equation left for the Weibull and the Gumbel once the other parameter is written in terms of it, and Nelder–Mead elsewhere; the triangles have a mode that is always one of the values, and are searched over the values for the mode and over the ends. Where SciPy fits the same family, the page reaches at least SciPy's maximum (its tests check it). A fit that is regular gets standard errors from the curvature of the log-likelihood at its maximum.

    The method of moments gives the member whose mean and variance (and skewness, for three parameters; the kurtosis for the Student t, which is symmetric) are the sample's. It is the method that matches what is known from a mean and an SD alone, and it is often worse in the tails. It matches the moments of the values themselves, not of their logarithms, so on a log shape it differs from maximum likelihood. A family that cannot be as skewed as the sample gets the nearest member, and a note says so.

    Some parameters are fixed, not fitted: the beta's ends at 0 and 1, the PERT's λ at 4, the generalized Pareto's threshold at 0, and for the binomial and beta-binomial the number of trials when you give it. Without a number of trials the binomial's n is estimated by profile likelihood.

    Ranking the fits

    Every fit is scored on the data and the table is sorted by the criterion you choose:

    • AIC = 2k − 2 ln L, where L is the likelihood and k the number of parameters fitted; AICc adds 2k(k + 1)/(n − k − 1), the correction for a small sample; BIC = k ln n − 2 ln L charges more for parameters. Smaller is better. The table shows each fit's difference from the best (Δ) and its Akaike weight, exp(−Δ/2) normalised over the fits of the same method: roughly, the probability that the fit is the best of those tried.
    • K–S D, the Kolmogorov–Smirnov statistic: the largest vertical gap between the empirical and the fitted cumulative curves.
    • A², Anderson–Darling: the squared gaps weighted towards the tails, where a limit is usually read.
    • W², Cramér–von Mises: the squared gaps, unweighted.
    • χ² p, for a discrete fit: Pearson's chi-square over cells with an expected count of at least five, with its degrees of freedom reduced by the parameters fitted. A small p says the fit is poor.

    The p-values next to D and A² are those for a distribution chosen before the data were seen. For a distribution fitted to the same data they are too large, often much too large (Lilliefors), so they rank and do not test. A fit whose support ends at the extremes of the data (the uniforms and the Pareto by maximum likelihood) puts those values at probability 0 or 1, where A² is infinite; its tests are then taken on the values strictly inside.

    Add puts a fit into the list of distributions, carrying the data with it. The P–P and Q–Q plots show the fits against the data: on the diagonal is a perfect fit, and a curve away from it at the ends is a tail the fit gets wrong.

    Calculations

    • Probabilities: P(X ≤ x), P(X > x), P(a < X ≤ b), the density (or probability) at x and the hazard rate at x, for every distribution.
    • Percentiles and intervals: the values at the percentiles you list, the central interval holding a given probability (equal tails), and the shortest interval holding it, which differs from the central one for a skewed distribution.
    • Tails and expectations: given a threshold t, the conditional means E[X | X > t] and E[X | X ≤ t] and the expected excess E[(X − t)⁺]; given a level, the value at risk (the percentile) and the mean beyond it (expected shortfall, or tail value at risk).
    • Two distributions: the probability that X exceeds Y when they are independent, the overlap of their densities, the largest gap between their cumulative curves (K–S), the Wasserstein distance (the area between their quantile functions), the Hellinger distance and the Kullback–Leibler divergence each way.
    • Monte Carlo of an expression: see above.

    Samples of the distributions have a tab of their own: see Samples.

    The families

    Each family below with its parameters, its density (or probability) and cumulative distribution, its mean and variance, and the same distribution in SciPy (scipy.stats), for checking a number in Python. Γ is the gamma function, B the beta function, Φ the standard normal CDF, P and Q the regularized incomplete gamma functions and I the regularized incomplete beta function.

    How the numbers are computed

    The special functions are written for this page: Cody's rational approximations for erf and erfc, Wichura's algorithm AS 241 for the normal quantile (with one Halley step), Stirling's series for log Γ, Loader's saddle-point form for binomial and Poisson probabilities, which stays accurate for millions of trials, and continued fractions for the incomplete gamma and beta functions, with their inverses by safeguarded Halley iteration. Each of a pair of complementary probabilities is computed directly where it is small, so that tail probabilities of 10⁻¹⁰⁰ keep their precision, and every quantile function has an upper-tail form for the same reason.

    Where a family has no closed form, the moments, entropy and logarithmic moments are integrals over the quantile function by tanh-sinh quadrature, split where the density has a corner, and quantiles are found by Brent's method. The page is tested against SciPy for every family at ordinary and awkward parameters, in both tails, and where SciPy itself loses precision (some far tails, binomials of millions of trials, the GEV near ξ = 0) against mpmath at 40 digits: densities, probabilities and quantiles agree to about ten significant figures.

    References

    Cody 1969Rational Chebyshev approximations for the error function. Mathematics of Computation 23, 631–637.
    Wichura 1988Algorithm AS 241: the percentage points of the normal distribution. Applied Statistics 37, 477–484.
    Loader 2000Fast and accurate computation of binomial probabilities. Technical note, Bell Laboratories.
    DLMFNIST Digital Library of Mathematical Functions, chapters 5 (gamma), 7 (error functions) and 8 (incomplete gamma and beta).
    Johnson, Kotz and Balakrishnan 1994–95Continuous Univariate Distributions, volumes 1 and 2. Wiley.
    Johnson, Kemp and Kotz 2005Univariate Discrete Distributions, 3rd ed. Wiley.
    Coles 2001An Introduction to Statistical Modeling of Extreme Values. Springer.
    Akaike 1974A new look at the statistical model identification. IEEE Transactions on Automatic Control 19, 716–723.
    Burnham and Anderson 2002Model Selection and Multimodel Inference, 2nd ed. Springer.
    Stephens 1974EDF statistics for goodness of fit and some comparisons. Journal of the American Statistical Association 69, 730–737.
    Marsaglia and Marsaglia 2004Evaluating the Anderson–Darling distribution. Journal of Statistical Software 9(2).
    Silverman 1986Density Estimation for Statistics and Data Analysis. Chapman and Hall.
    Hyndman and Fan 1996Sample quantiles in statistical packages. The American Statistician 50, 361–365.
    McKay, Beckman and Conover 1979A comparison of three methods for selecting values of input variables in the analysis of output from a computer code. Technometrics 21, 239–245.
    Sobol 1967On the distribution of points in a cube and the approximate evaluation of integrals. USSR Computational Mathematics and Mathematical Physics 7(4), 86–112.
    Bratley and Fox 1988Algorithm 659: implementing Sobol's quasirandom sequence generator. ACM Transactions on Mathematical Software 14, 88–100.
    Joe and Kuo 2008Constructing Sobol sequences with better two-dimensional projections. SIAM Journal on Scientific Computing 30, 2635–2654.
    Matoušek 1998On the L2-discrepancy for anchored boxes. Journal of Complexity 14, 527–556.
    Halton 1960On the efficiency of certain quasi-random sequences of points in evaluating multi-dimensional integrals. Numerische Mathematik 2, 84–90.
    Blackman and Vigna 2021Scrambled linear pseudorandom number generators. ACM Transactions on Mathematical Software 47(4).
    Takahasi and Mori 1974Double exponential formulas for numerical integration. Publications of RIMS 9, 721–741.