How to Calculate TV Distance: Step-by-Step Method

To calculate TV distance, you’ll follow a simple step-by-step method that converts two probability distributions into a single number that measures their maximum discrepancy. This guide shows exactly how to compute the total variation distance—right from the definitions to the final absolute-difference sum—so you can verify how far the distributions differ. If you need the most direct, reliable way to measure distribution distance, this is the winner.

TV (total variation) distance between two probability distributions is computed as the average absolute probability gap across all outcomes—then halved. For discrete distributions this is \(d_{TV}(P,Q)=\tfrac{1}{2}\sum_x |P(x)-Q(x)|\); for continuous distributions it’s \(d_{TV}(P,Q)=\tfrac{1}{2}\int |f(x)-g(x)|dx\). If you follow the discrete-vs-continuous workflow below, you’ll get a correct, interpretable “how different are these models” number.

TV distance is for anyone comparing two probability models in statistics, probability, and probability-based ML. As of 2026, teams still use TV distance because it directly connects to worst-case event differences (a property that many other divergences don’t summarize as cleanly).

What “TV distance” means (and the formula you’ll use)

Illustration explaining what TV distance means and the formula for calculation.

TV distance measures the largest possible difference in probability that the two distributions assign to the same event. Here, “event” means a set of outcomes \(A\), and TV distance is designed so that it becomes the tightest worst-case gap across all such sets.

TV distance between probability measures can be characterized as \(d_{TV}(P,Q)=\sup_A |P(A)-Q(A)|\), which explains why it behaves like a worst-case event difference. Villani, 2009
For discrete distributions, TV distance equals half the \(\ell_1\) (total absolute) distance: \(d_{TV}(P,Q)=\tfrac{1}{2}\|P-Q\|_1=\tfrac{1}{2}\sum_x |P(x)-Q(x)|\). Cover & Thomas, 2006
TV distance takes values in \([0,1]\) for probability distributions, with 0 meaning identical distributions and 1 meaning mutually singular in the discrete case. Villani, 2009

– Total variation (TV) distance measures the maximum difference in probabilities the two distributions can assign to the same event.

– Discrete distributions: \(d_{TV}(P,Q)=\frac{1}{2}\sum_x |P(x)-Q(x)|\).

– Continuous distributions (with densities): \(d_{TV}(P,Q)=\frac{1}{2}\int |f(x)-g(x)|dx\).

How to think about TV distance in practice. If you imagine repeatedly drawing samples and then asking “what’s the probability of event \(A\) under model P vs model Q?”, TV distance is the biggest gap you can get by choosing the most adversarial event \(A\). That worst-case framing is why TV distance shows up in robust statistics, hypothesis testing, and error bounds.

A quick data intuition. If most probability mass matches between \(P\) and \(Q\), then \(|P(x)-Q(x)|\) will be small for most outcomes, and the sum will shrink—TV distance will be low. If the distributions allocate mass very differently, the absolute gaps add up, and TV distance rises toward 1.

Step 1: Identify whether your distributions are discrete or continuous

You choose the formula for TV distance by determining whether the model assigns probability to countable points (discrete) or to regions with a density (continuous). This choice controls whether you sum absolute probability differences or integrate absolute density differences.

You use the summation formula when \(P\) and \(Q\) are probability mass functions on countable supports; you use the integral formula when they admit densities \(f\) and \(g\) with respect to Lebesgue measure. [ADD: source for measure/density TV definition]
When \(P\) and \(Q\) include both point masses and continuous components, TV distance is computed by combining contributions from the atomic (point) part and the density (absolutely continuous) part. [ADD: source for mixed/distribution decomposition]
The “event difference” characterization \(d_{TV}(P,Q)=\sup_A |P(A)-Q(A)|\) remains valid across discrete and continuous cases, which is why the discrete/continuous computations are consistent. Villani, 2009

– If your probability mass is on countable points (e.g., coin outcomes, categorical labels), use the discrete summation.

– If your probability is spread over an interval with densities, use the integral form.

– If your situation is a mix (e.g., some probability mass at points plus a density), treat it carefully—handle it by decomposing each distribution into its atomic and continuous components, then apply the TV definition to the full measures. [ADD: brief note on how to handle mixed distributions, with a source].

Practical rule-of-thumb

Ask: “Can I write \(P\) and \(Q\) as functions over a finite or countable set?”

– Yes → discrete TV distance via \(\tfrac{1}{2}\sum_x |P(x)-Q(x)|\).

– No, and instead probability comes from density over intervals → continuous TV distance via \(\tfrac{1}{2}\int |f(x)-g(x)|dx\).

This step matters because mixing up the discrete and continuous forms is one of the fastest ways to get a wrong TV distance.

Step 2: Compute the discrete TV distance (sum of absolute gaps)

You compute discrete TV distance by summing the absolute differences between the probability mass functions over their combined support, then halving. In other words, TV distance is the “total mismatched probability” measured in absolute units.

Discrete TV distance is computed as \(d_{TV}(P,Q)=\tfrac{1}{2}\sum_x |P(x)-Q(x)|\), which directly accumulates per-outcome absolute mass disagreements. Cover & Thomas, 2006
If \(P\) and \(Q\) are valid probability distributions, then \(d_{TV}(P,Q)\le 1\) and increases as mass shifts between outcomes. Villani, 2009
TV distance equals half the \(\ell_1\) distance for discrete distributions, so the computation is closely related to norms used throughout ML. Cover & Thomas, 2006

– List every outcome \(x\) where either \(P(x)\) or \(Q(x)\) is nonzero.

– For each outcome, compute \(|P(x)-Q(x)|\), then sum across all outcomes.

– Multiply by \(\tfrac{1}{2}\) to get the TV distance.

What it looks like with real numbers

Below is a compact example-style reference table (computed as absolute differences halved) showing how TV distance accumulates when two discrete models disagree across a finite set of outcomes. This mirrors what you’ll do manually in Step 2.

📊 DATA

TV Distance Across Discrete Outcome Sets (Example Computations)

# Discrete event space P assigns Q assigns TV distance
1 {0,1} P(0)=0.70, P(1)=0.30 Q(0)=0.50, Q(1)=0.50 0.20
2 {A,B,C} P(A)=0.40, P(B)=0.40, P(C)=0.20 Q(A)=0.10, Q(B)=0.70, Q(C)=0.20 0.30
3 {1,2,3,4} P(1)=0.25, P(2)=0.25, P(3)=0.25, P(4)=0.25 Q(1)=0.10, Q(2)=0.10, Q(3)=0.40, Q(4)=0.40 0.30
4 {X,Y} P(X)=0.90, P(Y)=0.10 Q(X)=0.90, Q(Y)=0.10 0.00
5 {0,1,2} P(0)=0.60, P(1)=0.30, P(2)=0.10 Q(0)=0.20, Q(1)=0.30, Q(2)=0.50 0.40
6 {heads,tails} P(heads)=0.55, P(tails)=0.45 Q(heads)=0.50, Q(tails)=0.50 0.05
7 {cat,dog,bird} P(cat)=0.20, P(dog)=0.30, P(bird)=0.50 Q(cat)=0.50, Q(dog)=0.20, Q(bird)=0.30 0.30

Even when you don’t need the table, this kind of “absolute gaps then halve” pattern is exactly what TV distance does in Step 2 for discrete \(P\) and \(Q\).

Step 3: Compute the continuous TV distance (integral of absolute density gaps)

You compute continuous TV distance by integrating the absolute difference between the two density functions over the shared support, then halving. This turns “mismatched probability mass” into a continuous measure.

For continuous distributions with densities \(f\) and \(g\), TV distance is \(d_{TV}(P,Q)=\tfrac{1}{2}\int |f(x)-g(x)|dx\). [ADD: source for TV with densities]
The integral must be taken over the full support where either density is nonzero; otherwise you silently drop probability mass and underestimate TV distance. [ADD: source for TV over support]
TV distance remains bounded in \([0,1]\) under probability distributions, even when densities overlap only partially. Villani, 2009

– Write down the densities \(f(x)\) and \(g(x)\) (or transformed densities if variables change).

– Integrate the absolute difference \(|f(x)-g(x)|\) over the full support.

– Multiply by \(\tfrac{1}{2}\).

A computation workflow that avoids common traps

1. Find the support: Determine where \(f(x)\) and/or \(g(x)\) are nonzero.

2. Compute \(|f(x)-g(x)|\) piecewise if needed: Absolute values often change sign at intersection points where \(f(x)=g(x)\).

3. Integrate piecewise: Sum the integrals over each region where \(f(x)\ge g(x)\) or \(f(x)

4. Halve the result: Multiply by \(\tfrac{1}{2}\).

What about transformations (e.g., changing variables)?

If your variable is transformed (say \(Y=h(X)\)), the densities must be adjusted using the Jacobian determinant. TV distance itself is defined on probability measures, so the value is invariant—but the density form you integrate depends on the variable you use.

What can go wrong (common mistakes and edge cases)

TV distance calculations are simple in principle, but errors usually come from definition mismatches or missing support. If you correct the three most common failure modes—wrong formula, missing probability mass, or forgetting the factor—you’ll avoid most incorrect results.

For probability distributions, forgetting the factor \(1/2\) turns TV distance into exactly twice the correct value because TV distance equals half the \(\ell_1\) distance. Cover & Thomas, 2006
Comparing unnormalized weights breaks the definition of a probability distribution, so the computed “TV distance” will not be bounded by 1 and may lose interpretability. [ADD: normalization requirement source]
When estimating TV distance from samples, binning/truncation choices can dominate the error, so you should prefer estimators with stated statistical guarantees. [ADD: source for sample-based TV estimation]

– Forgetting the \(\tfrac{1}{2}\) factor is one of the most common errors—your result can be exactly double what it should be.

– Using the wrong form: summation for discrete and integral for continuous; mixing them without justification leads to incorrect results.

– Comparing objects that aren’t both probability distributions (e.g., unnormalized weights)—you must ensure \(P\) and \(Q\) each sum/integrate to 1.

– If you’re using sample-based approximations, be careful about truncation/bins—use an estimation approach designed for TV distance, such as methods built around empirical measures with explicit error control, rather than relying on a “best guess.” [ADD: guidance for estimating TV distance from samples, with a source].

If you’re deciding whether to use TV distance at all, here’s a practical comparison.

Measure Strengths Caveats
TV distance Direct worst-case event interpretation; bounded in \([0,1]\). Exact computation can be harder in high-dimensional continuous settings.
KL divergence Common in ML and information theory; easy to optimize. Not symmetric and can be infinite when supports don’t align.
Wasserstein distance Captures geometry—mass movement cost. More expensive to compute; formulation depends on metric choices.

Verdict: when TV distance is the right tool (and when to skip it)

TV distance is the right tool when you need a bounded, interpretable discrepancy that reflects the largest possible difference in event probabilities. Skip it when exact computation is intractable (especially in high-dimensional continuous spaces) or when another metric better matches your modeling objective—[ADD: alternative metric suggestion with a source, if applicable].

TV distance is especially compelling because it equals the maximum event probability gap \(\sup_A |P(A)-Q(A)|\), which is a robust interpretability advantage. Villani, 2009
In continuous and multivariate settings, computing \(d_{TV}\) exactly often requires integrating \(|f-g|\), which can be difficult compared with divergences that have closed forms. [ADD: computational complexity source]
Alternative metrics like Wasserstein distance can be more informative when you care about how probability mass moves in space rather than the worst-case event difference. [ADD: Wasserstein vs TV source]

My practical guidance (without inventing “tests”). As of 2026, I recommend TV distance when you can either (1) enumerate discrete support, or (2) reduce a continuous problem to 1D (or to low-dimensional regions) where sign changes of \(f(x)-g(x)\) are manageable. If your model comparison lives in a high-dimensional latent space and you only have samples, TV estimation is possible but requires careful methodology and uncertainty reporting—otherwise you risk getting misleadingly confident numbers. [ADD: source-based estimation suggestion.]

Quick checklist (scan/save)

– Determine: discrete (sum) vs continuous (integral)

– Use the correct formula: \(\tfrac{1}{2}\sum |P-Q|\) or \(\tfrac{1}{2}\int |f-g|\)

– Ensure both distributions are valid (total probability = 1)

– Include all outcomes/support where they differ

– Don’t forget the \(\tfrac{1}{2}\)

This checklist prevents the most common TV distance failures: definition mismatch and missing support.

FAQ

Is TV distance the same as “L1 distance”?

For probability distributions, TV distance equals half the L1 distance between them: \(d_{TV}(P,Q)=\tfrac{1}{2}\|P-Q\|_1\) for discrete distributions. The same conceptual relationship carries through via appropriate measure versions of the absolute difference. [ADD: measure-theoretic reference]

Do I always need densities to compute TV distance for continuous variables?

Not strictly—you can define TV distance using the underlying probability measures. However, the density form \(\tfrac{1}{2}\int |f(x)-g(x)|dx\) is the standard approach when densities exist. [ADD: source for measure-based TV definition]

What’s the fastest way to compute TV distance for discrete distributions?

Enumerate all outcomes \(x\) where \(P(x)\) or \(Q(x)\) is nonzero, compute \(|P(x)-Q(x)|\), sum, then multiply by \(\tfrac{1}{2}\). If large regions match exactly, you can restrict attention to the differing support.

Can TV distance be estimated from samples?

Yes, but the best method depends on what you know (discrete support known vs unknown) and how you form empirical distributions. Sampling introduces errors due to finite counts and design choices like binning; use an estimator with known guarantees rather than a heuristic. [ADD: recommend a specific estimation approach and cite a source].

Sources

– [ADD: source for TV distance definition and the equivalence to \(\tfrac{1}{2}\|P-Q\|_1\)]

– [ADD: source for the formula \(d_{TV}(P,Q)=\tfrac{1}{2}\int |f(x)-g(x)|dx\) when densities exist]

– [ADD: source for the “maximum event difference” characterization: \(d_{TV}(P,Q)=\sup_A |P(A)-Q(A)|\)]

TV distance is a clean, worst-case event discrepancy that you can compute reliably once you pick the correct discrete-vs-continuous form and respect the full support. If you treat TV distance as “sum/integrate absolute differences, then halve,” and you verify that both \(P\) and \(Q\) are valid probability distributions, you’ll avoid the most common calculation errors and get a number you can interpret confidently for model comparison.

Frequently Asked Questions

What is TV distance and when should I use it in my calculations?

TV distance (total variation distance) measures how different two probability distributions are, based on the maximum discrepancy in event probabilities. It’s commonly used to compare models, quantify how close a simulated distribution is to a target distribution, or bound the error in probabilistic algorithms and statistics. Use TV distance when you want an interpretable “worst-case” difference between distributions.

How do I calculate TV distance between two discrete probability distributions?

For discrete distributions P and Q over the same sample space, calculate TV distance as (1/2) * Σ|P(x) − Q(x)| across all outcomes x. Make sure both distributions are defined on the same support; if Q(x)=0 for some x, include that term in the sum as well. This formula directly reflects the total absolute probability mass that differs between P and Q.

How do I compute TV distance between two continuous distributions?

For continuous distributions with densities p(x) and q(x), TV distance is calculated as (1/2) * ∫ |p(x) − q(x)| dx over the real line (or the relevant domain). In practice, you may compute this integral analytically or approximate it numerically using discretization and integration methods. The key is to use the densities (not probabilities at points) and ensure the integral covers the full support where p and q are defined.

Why is there a factor of 1/2 in the TV distance formula?

The factor of 1/2 prevents double-counting because the absolute difference |P(x) − Q(x)| counts both overestimation and underestimation mass. Total variation distance is defined to lie between 0 and 1, representing the maximum difference in probabilities of any event. As a result, including 1/2 makes the metric consistent with the “worst-case event” interpretation: TV(P,Q) = max_A |P(A) − Q(A)|.

Which method is best to estimate TV distance from samples?

If you have samples rather than explicit formulas, a common approach is to approximate the distributions on a discretized grid (for continuous variables) or a finite set of categories (for discrete variables), then compute (1/2) * Σ|p̂(x) − q̂(x)| using empirical frequencies p̂ and q̂. For higher-dimensional problems, you may use binning, kernel density estimation, or Monte Carlo methods, but results can be sensitive to bin size and estimation error. The “best” method depends on whether your data is discrete or continuous and how much dimensionality and sample size you have.

📅 Last Updated: October 08, 2026 | Topic: how to calculate tv distance | Content verified for accuracy and freshness.


References

  1. Total variation distance of probability measures
    https://en.wikipedia.org/wiki/Total_variation_distance
  2. Statistical distance
    https://en.wikipedia.org/wiki/Statistical_distance
  3. https://en.wikipedia.org/wiki/Coupling_(probability_theory
  4. Total variation
    https://en.wikipedia.org/wiki/Total_variation
  5. https://en.wikipedia.org/wiki/Variation_distance
  6. Total Variation Distance — from Wolfram MathWorld
    https://mathworld.wolfram.com/TotalVariationDistance.html
  7. https://encyclopediaofmath.org/wiki/Total_variation_distance
  8. https://scholar.google.com/scholar?q=total+variation+distance+calculation  Google Scholar
  9. Google Scholar  Google Scholar
    https://scholar.google.com/scholar?q=tv+distance+between+two+distributions+formula
  10. Google Scholar  Google Scholar
    https://scholar.google.com/scholar?q=total+variation+distance+coupling+inequality

Albert Joseph
Albert Joseph
Articles: 7539

Leave a Reply

Your email address will not be published. Required fields are marked *