Blog banner

Conservation of Distance

A working principle for model building and diagnosis

machine learning
mathematics
model evaluation
Conservation of Distance as a working principle for building and debugging supervised learning models: task-relevant similarity, stability, and independent evaluation.
Author

Soma S Dhavala

Published

October 7, 2026

Conservation laws provide a common basis for modelling physical systems. For a closed material system in classical mechanics, mass, momentum and energy balances take the form

\[ \frac{dM}{dt}=0, \qquad \frac{d\mathbf P}{dt}=\mathbf F_{\mathrm{ext}}, \qquad \frac{dE}{dt}=\dot Q-\dot W. \]

Here \(M\) is mass, \(\mathbf P\) is momentum, \(\mathbf F_{\mathrm{ext}}\) is net external force, \(E\) is total energy, \(\dot Q\) is heat supplied, and \(\dot W\) is work done by the system. These balances underlie calculations of fluid flow, motion and heat transfer. Constitutive laws and initial and boundary conditions complete the specification of a particular problem. NASA: Navier–Stokes Equations

The same balance structure applies to probability mass. For a probability density \(p_t\) and probability flux \(J_t\), the continuity equation is

\[ \frac{\partial p_t}{\partial t}+\nabla\cdot J_t=0. \]

With zero net boundary flux, or sufficient decay at infinity, it preserves total probability. In continuous-time diffusion models, the Fokker–Planck equation has this form; the corresponding probability-flow ODE provides a deterministic evolution with the same time-marginal distributions under suitable conditions. Song et al.: Score-Based Generative Modeling through Stochastic Differential Equations

Can a comparable principle help us build and assess supervised learning models?

In an earlier LinkedIn post, I proposed the following informal statement and asked whether it had exceptions:

Any two examples closer in the input space must also be closer in the output space.

The statement leaves the meaning of closeness and the choice of input and output spaces unspecified. This article develops those choices and examines when the resulting condition is useful for building and diagnosing models.

In supervised learning, we use observed input–output pairs to predict outcomes for new inputs. This requires assumptions about which relationships extend beyond the training data. One assumption is that inputs which are similar in ways relevant to the task should have related expected outputs or output distributions. Making this assumption explicit provides both a constraint for model construction and a criterion for checking predictions.

The proposed connection is through this role as a constraint. Physical conservation equations constrain changes in quantities through time and space. For supervised learning, the condition considered here constrains differences between predictions relative to differences between inputs. This motivates Conservation of Distance as a working principle.

Conservation of Distance: A working principle for supervised learning

Conservation of Distance directs attention to whether differences in predictions are consistent with differences that matter to the task. This applies across nearest neighbours, regression, decision trees, support vector machines and neural networks, although each method represents those relationships differently.

NoteConservation of Distance

Inputs which are close under a task-relevant distance should have correspondingly close expected outputs or output distributions.

Let \(D_T\) be a justified task-reference distance and \(m(x)\) the target conditional mean or probability vector. For a finite nonnegative constant \(L_m\), assume

\[ D_Y(m(x),m(x'))\leq L_m D_T(x,x'). \]

If \(D_Y(f(x),m(x))\leq\varepsilon\) at both inputs being compared, then

\[ D_Y(f(x),f(x'))\leq L_m D_T(x,x')+2\varepsilon. \]

Here \(D_Y\) is a metric or pseudometric on outputs. The first inequality is an assumption about the target; the second is a necessary consequence of that assumption and the stated accuracy. Satisfaction does not establish predictive usefulness. The distance choices and the derivation are developed below.

The qualification “task-relevant” is essential. Learning involves identifying relevant differences, expressing them in a representation, and testing whether the resulting relationships help predict unseen outcomes.

As a working principle, Conservation of Distance requires the modeller to specify which input changes should leave predictions approximately stable. As a diagnostic, it tests whether a model meets that requirement. The necessary-condition argument follows from assumed target stability and prediction accuracy: together they bound the difference between predictions. A violation calls at least one of those assumptions into question. The distance, domain and acceptable output variation must therefore be justified independently. Unlike physical conservation, the condition does not assert an invariant total.

Task-dependent similarity

In the original discussion, Jigar Doshi used the word “bank” to illustrate dependence on context. Mani Srinivasan raised negation in sentiment analysis. Both comments identify information that a suitable representation must retain.

Consider three messages sent to a subscription service:

Message Requested action
A: “Please renew my subscription.” Renew
B: “Please do not renew my subscription.” Do not renew
C: “Keep my membership active for another year.” Renew

By word overlap, A and B are close. For predicting the requested action, A and C belong together. Word overlap does not capture the action requested in these examples.

Even this grouping depends on the task. If we are routing messages to departments, all three may belong to subscriptions. If we are deciding what action to take, the negation is essential. If we are identifying the language, all three again have the same target.

Similarity must therefore be defined with respect to a task. There is no reason to expect a single distance to serve all three predictions equally well.

The same issue appears in numerical data. Adjacent customer identifiers do not imply similar customers. Two houses on the same street may differ greatly in size and condition. Standardising a feature changes its scale; it cannot establish its relevance.

Representation and prediction

Write a predictor as

\[ f(x)=g(\phi(x)). \]

The map \(\phi\) constructs a representation. The map \(g\) makes a prediction from it. These components need not be trained separately: the decomposition is a way to examine their roles.

For the subscription messages, a useful representation for action prediction would suppress some differences in wording while retaining the distinction introduced by “not.” In a house-price model, it might express size, condition and location in a form from which price is easier to estimate.

One possible distance is then

\[ D_\phi(x,x')=\lVert\phi(x)-\phi(x')\rVert_2. \]

This is always a pseudometric on the original input space; it is a metric when \(\phi\) is injective. Distinct observations may have zero distance if their representations coincide.

Learning a useful representation often requires changing distances. We may want paraphrases to move together and opposite requests to move apart. The intended consistency is between task-relevant proximity and predictive behaviour. Preservation of the original geometry is not required.

Representation learning develops this idea systematically. A representation is useful insofar as it makes the subsequent task easier, and its properties depend on that task and the predictor that follows it. Ordinary supervised training can produce such representations without explicitly optimising a distance objective. Goodfellow, Bengio and Courville, Chapter 15

The model-internal distance \(D_\phi\) and task-reference distance \(D_T\) have different roles. The former describes relationships within the fitted model; the latter specifies the relationships against which its behaviour is assessed. For example, \(D_T\) may come from scaled physical measurements or independently judged equivalence classes of requests. It must be fixed before final evaluation and justified through task evidence rather than defined from the predictions being tested.

A learned distance can serve as \(D_T\) if it is frozen and independently validated. The distinction concerns its evidential role, not whether it was learned. The two distances may coincide, but that agreement is a modelling choice. Defining \(D_T\) from the predictor’s outputs would make the strict predictor bound true by construction; defining it from the target \(m\) would do the same for the target-stability assumption.

Separating representation, distance and prediction

Supervised methods specify how observations share predictive information. The representation, comparison rule and predictor may be fixed, learned separately or learned jointly. It would be too strong to say that every supervised method learns a representation or a metric: nearest-neighbour prediction can use fixed features and a prescribed distance while retaining labelled examples.

Separating these components helps identify where a stability assumption enters the model. A learned tree partition defines which observations share a leaf. A neural representation determines the inputs to its output layer. Neither construction alone establishes that the resulting proximity is appropriate for the target.

The separation is also not unique. For an invertible linear transformation \(A\) of the representation,

\[ f=g\circ\phi =(g\circ A^{-1})\circ(A\circ\phi). \]

Predictions remain unchanged, while Euclidean distances in the transformed representation can change substantially. Representation constraints, distance scales and the prediction rule must therefore be examined together. Separating them makes their roles inspectable; it does not automatically establish conservation.

A mathematical formulation

Physical conservation equations describe changes through time or space. The proposed stability condition instead compares pairs of inputs. In the LinkedIn discussion, Ajay Shenoy identified Lipschitz continuity as the condition that supplies such a bound.

A strict design requirement on the predictor can be written as

\[ D_Y(f(x),f(x'))\leq L_f\,D_T(x,x'). \]

Here \(L_f\) is a chosen predictor sensitivity bound. This strict design requirement differs from the error-aware bound derived from target stability with constant \(L_m\). With metric or pseudometric input and output distances, the strict condition is Lipschitz continuity. The necessity proof below uses metric properties only on the output side; its input comparison can be any specified nonnegative dissimilarity.

For regression, \(D_Y\) might be absolute difference between predicted means. For classification, it might compare probability vectors. Numeric differences between arbitrary class identifiers have no corresponding meaning.

The inequality allows contraction. Very different inputs may have the same prediction. It also allows expansion up to the factor \(L_f\), and says nothing about a converse: similar outputs need not imply similar inputs. Input and output distances can have different units. In my reply to Ajay, I suggested bi-Lipschitz behaviour as desirable. That stronger requirement also bounds contraction and needs a narrower scope: classification and other many-to-one tasks must allow distinct inputs to share an output.

Exact preservation up to scale would require \(D_Y(f(x),f(x'))=c D_T(x,x')\) for a fixed positive scale \(c\). The upper bound does not require this equality.

The conventional connection is to the smoothness or local constancy prior: nearby inputs are expected to have related outputs. This is an established and useful inductive bias, with limitations, especially when learning depends only on local neighbourhoods in high-dimensional spaces. Goodfellow, Bengio and Courville, Chapter 5

Observed targets need additional care. Two customers with identical recorded features can behave differently. The relevant regularity may concern \(\mathbb E[Y\mid X=x]\) or \(P(Y\mid X=x)\) rather than each realised outcome. Even these quantities need not be smooth in the chosen coordinates. A real eligibility threshold can introduce a discontinuity, and thresholding a smoothly varying class probability can create a hard decision boundary.

Design choices left unspecified

The principle does not specify a complete model. It leaves the representation \(\phi\), prediction rule \(g\), input distance, output discrepancy, domain of application and acceptable sensitivity bounds open for design. Some choices are fixed using domain knowledge; others are estimated from training data and selected through validation. Representations and distances need not all be learned, but their suitability must be assessed.

Component Examples of design choices
Input representation Standardised measurements, selected physical variables, text embeddings, learned neural features
Input distance or dissimilarity Euclidean distance, a weighted distance, a learned Mahalanobis distance, cosine dissimilarity
Prediction rule Linear regression, neighbour averaging, a tree, a neural output layer
Output discrepancy Absolute error for scalar predictions, a weighted norm for multiple outputs, a divergence between predicted class distributions
Scope and sensitivity A specified neighbourhood, operating range or family of transformations; an acceptable bound on output variation

Rajesh Talluri emphasised that input and output spaces can have different notions of proximity and suggested considering transformed spaces. These choices remain part of the model specification.

For example, a house-price model may weight differences in floor area and location differently. A multi-output model may require separate scales for temperature and pressure. A classifier requires a comparison between probability distributions, rather than subtraction of category identifiers. Each choice changes what the stability requirement means.

Metrics, pseudometrics and divergences

Here, “distance” is used broadly for a task-specific measure of discrepancy. Its precise mathematical properties must be stated when specifying a model or diagnostic.

A metric is nonnegative, symmetric, satisfies the triangle inequality, and is zero exactly when its arguments are identical. A pseudometric retains the other properties but permits distinct observations to have zero distance. The representation distance \(\lVert\phi(x)-\phi(x')\rVert_2\) is an example when \(\phi\) maps distinct inputs to the same representation. This can be appropriate when their differences are irrelevant to the task.

A divergence need not be symmetric or satisfy the triangle inequality. For class-probability vectors \(p\) and \(q\), one option is

\[ D_{\mathrm{KL}}(p\Vert q)=\sum_i p_i\log\frac{p_i}{q_i}. \]

Its direction matters, and it can be infinite when \(q_i=0\) for a class with \(p_i>0\). A bound using such a divergence is a discrepancy bound; it is not a standard metric Lipschitz condition. Properties that depend on symmetry or the triangle inequality cannot be assumed.

Cross-entropy as an output discrepancy

Cross-entropy is another common choice for assessing outputs:

\[ H(p,q)=-\sum_i p_i\log q_i =H(p)+D_{\mathrm{KL}}(p\Vert q). \]

When \(p\) is a fixed target distribution and \(q\) is the prediction, minimising cross-entropy is equivalent to minimising KL divergence. For a one-hot target with class \(y\), it reduces to \(-\log q_y\). This makes cross-entropy a suitable prediction loss, although it is not a metric or a pseudometric. Relation between cross-entropy, entropy and KL divergence

Its use in the pairwise conservation condition needs an adjustment. If \(p=f(x)\) and \(q=f(x')\), then even identical distributions have \(H(p,p)=H(p)\), which is generally positive. Directly substituting cross-entropy into the zero-baseline bound would therefore fail at \(x=x'\) for a non-degenerate probability vector.

The corresponding excess cross-entropy condition is

\[ H(f(x),f(x'))-H(f(x)) =D_{\mathrm{KL}}(f(x)\Vert f(x')) \leq C_{\mathrm{KL}}\,D_T(x,x'). \]

Equivalently, cross-entropy can be bounded by \(H(f(x))+C_{\mathrm{KL}}D_T(x,x')\). The entropy term accounts for the reference distribution’s uncertainty. Thus cross-entropy can be used in the output comparison, but its baseline must be included. This is a separate design condition, not a consequence of the metric necessity result.

Pairs are ordered for KL; imposing the bound in both directions requires checking both orders. For classification, total variation \(D_Y(p,q)=\tfrac12\sum_i|p_i-q_i|\) provides a metric alternative, while cross-entropy can remain the training loss. Low average cross-entropy does not itself establish a pointwise accuracy allowance.

Comparison across algorithms

Conservation of Distance offers a common set of questions without erasing the differences between algorithms.

Nearest neighbours makes the choice explicit. Select a representation and distance, find nearby training examples, and combine their targets. The method depends directly on whether those neighbours are informative for the query. Feature selection and scaling affect which examples are retrieved. Hard neighbour selection can produce jumps when the neighbour set changes, so it does not generally guarantee a Lipschitz predictor in the original coordinates.

Linear regression describes a global relationship. For \(f(x)=w^\top x+b\),

\[ |f(x)-f(x')|\leq\lVert w\rVert_2\,\lVert x-x'\rVert_2. \]

The coefficient norm bounds sensitivity in these coordinates. Movement perpendicular to \(w\) leaves the prediction unchanged; movement along \(w\) changes it. Regression can extrapolate through this global structure without retrieving nearby observations.

Decision trees learn a partition. A regression tree gives observations in the same leaf the same fitted mean. Its splits determine which distinctions affect the prediction. Across a leaf boundary, however, an arbitrarily small input change can produce a jump. Tree predictions therefore need not be smooth in the original coordinates.

An RBF-kernel support vector machine uses a similarity that decays exponentially with squared distance. Its decision score combines kernel evaluations against support vectors with fitted signed coefficients. This connects prediction to geometry, but the computation is not an average of neighbouring labels: margins, coefficients and the intercept matter.

Neural networks can learn the representation and prediction rule together. Hidden layers change the coordinates available to later layers. Whether ordinary Euclidean distance in one of those layers is useful for a particular task remains something to assess; a trained network does not automatically make every hidden-space distance meaningful.

These methods give us several ways to specify which observations share predictive information: explicit neighbourhoods, global directions, learned regions, kernel similarities and learned features. Their assumptions and mechanisms remain distinct.

Necessity, sufficiency and testability

The necessity of Conservation of Distance depends on the target relationship and the independently justified task-reference comparison \(D_T\). Let \(m(x)\) denote the target conditional mean or conditional probability vector. Suppose that, on the domain under consideration,

\[ D_Y(m(x),m(x'))\leq L_m D_T(x,x'). \]

If a predictor approximates this target within \(\varepsilon\) at both inputs, then

\[ D_Y(f(x),m(x))\leq\varepsilon, \qquad D_Y(f(x'),m(x'))\leq\varepsilon. \]

For a metric or pseudometric output distance, the triangle inequality gives

\[ \begin{aligned} D_Y(f(x),f(x')) &\leq D_Y(f(x),m(x)) +D_Y(m(x),m(x')) +D_Y(m(x'),f(x'))\\ &\leq L_m D_T(x,x')+2\varepsilon. \end{aligned} \]

Thus, stability of the target and accuracy at both inputs imply a corresponding stability bound for the predictions, with an allowance for approximation error. A violation means that the assumed target stability or the claimed accuracy fails. It identifies an inconsistency without determining its cause. Low average test error does not establish the required accuracy at every input. This argument also does not transfer directly to divergences that lack the triangle inequality.

Why satisfaction is insufficient

A constant predictor satisfies the strict design inequality with \(L_f=0\), even when it fails to represent variation in the target. Stability therefore cannot establish predictive usefulness. Accuracy and appropriate responses to meaningful changes remain separate requirements.

This does not remove the condition’s diagnostic value. If independently checked paraphrases receive substantially different renewal probabilities, their predictions may be inconsistent with the assumed target stability and accuracy. Aggregate metrics can conceal such behaviour. A condition can be necessary under stated assumptions without being sufficient for the full task.

Why necessity is conditional

Gabriel Chua noted that neighbouring inputs can receive different class labels across a decision boundary. A deterministic threshold gives a direct example. Consider the target

\[ m(x)=\mathbf 1\{x\geq0\}. \]

An exact predictor has zero error. With ordinary absolute distance on inputs and outputs, however, it cannot satisfy any finite global Lipschitz bound. Inputs on opposite sides of zero can be arbitrarily close while their outputs differ by one. Successful supervised prediction therefore does not universally require this bound in the original coordinates. Here the target-stability premise fails for every finite \(L_m\), so the necessity result does not apply.

Lalit Verma also raised changes in output behaviour, using the term “inflection points.” For the bound, the relevant cases are discontinuities or unbounded local sensitivity; a change in curvature alone need not violate it.

A representation that separates the two regions can restore a finite bound. This makes the representation important, but does not establish a universal law: if unrestricted representations are allowed, choose \(\phi(x)=f(x)\) and define the induced distance

\[ D_f(x,x')=D_Y(f(x),f(x')). \]

Every predictor then satisfies a bound in its own induced distance with constant one, including an inaccurate predictor. The existence of some representation satisfying the condition cannot distinguish successful learning from failed learning.

Sensitivity to initial conditions

Responding to the earlier post, Nithin Nagaraj raised chaotic nonlinear systems as an objection: small differences in initial conditions can grow substantially, and nearby states on opposite sides of a separatrix can lead to different outcomes. This identifies two distinct limitations of the informal claim.

First, sensitivity depends on the prediction horizon. For an ODE \(\dot z=b(z)\) whose vector field is Lipschitz with constant \(K\) on a region containing both trajectories over \([0,t]\), the flow \(\Phi_t\) satisfies, for \(t\geq0\),

\[ \lVert\Phi_t(x)-\Phi_t(x')\rVert \leq e^{Kt}\lVert x-x'\rVert. \]

A finite-time bound can therefore coexist with rapid separation. In chaotic dynamics, positive Lyapunov exponents describe exponential growth of infinitesimal perturbations along unstable directions; they are not themselves uniform bounds for all pairs. Politi: Lyapunov Exponent

For forecasting, a bound that grows rapidly with time may permit output differences too large to be useful. The existence of a finite Lipschitz constant alone therefore says little about practical predictability. The horizon, measurement precision and acceptable error must be specified.

Second, predicting an eventual outcome is different from predicting the state at a finite time. If the target identifies which basin of attraction contains an initial condition, arbitrarily close inputs across a separatrix forming a basin boundary can receive different labels. The target can be discontinuous in the original coordinates, even when the underlying flow is smooth. Such boundaries can occur without chaos.

The objection thus limits a universal reading of the original statement. A model should reproduce genuine sensitivity or a genuine outcome boundary when the task requires it. A violation of an assumed stability bound can reveal that the bound is unsuitable for the target, rather than that the model is defective.

A testable formulation

A specified instance of the principle is falsifiable. Rishabh Bhardwaj distinguished a guiding principle from a universal law in the original discussion. The conditional formulation here makes the assumptions that can be tested explicit. Freeze the task-reference distance \(D_T\), the output distance, the domain, the bound and any approximation-error allowance before evaluation. A pair that violates the resulting inequality is a counterexample to that instance. Independently assessed targets help determine whether the failure concerns prediction accuracy or the assumed target stability.

For subscription messages, humans can identify paraphrases and action reversals without consulting the model’s embeddings. Training labels may legitimately shape the representation, but evaluation must assess its relevance on unseen cases. Choosing a new reference distance or increasing the bound after every failure would remove the test’s ability to reject the specification.

For noisy outcomes, a single observed label does not reveal the conditional mean or probability vector \(m(x)\). The error allowance therefore needs separate evidence, such as a known synthetic target, repeated observations with uncertainty estimates, or a justified accuracy guarantee. Without that evidence, pairwise violations remain diagnostics of a proposed specification; they do not by themselves establish that the target-stability assumption is false.

Building models with the principle

The principle can guide feature selection, representation learning and the training objective. If units or irrelevant identifiers dominate the distance, the representation needs revision. If valid paraphrases should receive similar predictions, independently checked paraphrase pairs can provide training examples or consistency constraints. Meaningful changes, such as negation, must retain the distinctions needed for prediction.

For a specified collection of training pairs \(\mathcal P\), one possible consistency penalty is

\[ \mathcal L_{\mathrm{cons}} =\frac{1}{|\mathcal P|}\sum_{(x,x')\in\mathcal P} \left[\max\left\{0, D_Y(f(x),f(x'))-L_f D_T(x,x') \right\}\right]^2. \]

It can be combined with the supervised prediction loss as \(\mathcal L_{\mathrm{pred}}+\lambda\mathcal L_{\mathrm{cons}}\), where \(\lambda\geq0\) controls the penalty’s contribution. This is one possible implementation of the principle, not a requirement for every model. Enforcing it on sampled pairs does not establish a global Lipschitz bound.

This penalty uses a fixed task-reference distance and chosen design bound \(L_f\). If an internal learned distance is used instead, its scale must be constrained: enlarging it could reduce the penalty without improving predictions. Pair construction and the prediction loss remain essential. The penalty controls excessive output variation; it supplies no requirement that distinct targets receive distinct predictions.

Debugging and empirical evaluation

Pramod Kompalli connected the original statement to robustness under small input perturbations. The subscription example makes the required checks concrete. An independently verified paraphrase should produce similar class-probability vectors under the specified bound. A negation that changes the target calls for an appropriate probability change; that is an additional accuracy test, not a consequence of the upper stability bound. Near a decision threshold, a small permissible probability change can flip the selected action, so an action flip alone is not a violation.

This pairing is useful beyond text. We can specify changes that should preserve a target and changes that should alter it, then examine both prediction quality and sensitivity. Those expectations must come from the application; even a seemingly harmless image transformation can change the label in some tasks.

A practical study can proceed through four steps:

  1. Specify the target. Department routing and requested-action prediction require different notions of similarity.
  2. Specify the relevant relationships. Fix the reference distance, output distance, bound and error allowance, with examples of irrelevant variation and meaningful change, before inspecting test results.
  3. Fit and select the representation. Fit preprocessing, features and model parameters using the permitted training data; select choices using validation data.
  4. Evaluate generalisation. Evaluate accuracy and the specified consistency tests on untouched cases, with splits that reflect deployment. Closely related templates or customers may need to stay in the same split.

A failure can reveal an omitted variable, a poor representation, an unsuitable model, or a change in the data-generating process. Conservation of Distance helps formulate the diagnosis; it does not determine the cause by itself.

For each evaluation pair, a direct diagnostic is the excess output variation

\[ e(x,x')=\max\left\{0, D_Y(f(x),f(x'))-L_m D_T(x,x')-2\varepsilon \right\}. \]

This checks the error-aware necessary bound. To test a strict predictor design requirement instead, use \(L_f\) and omit the error allowance. Positive values can be grouped by transformation, data source or operating condition. The reference distance, bound and allowance must be fixed before examining these results.

Inspection can then follow the prediction pipeline. Equivalent raw inputs that become distant after preprocessing suggest a problem with parsing, units, scaling or feature construction. Inputs that remain close in the selected representation but produce excessive output differences point towards the prediction rule’s sensitivity. Inputs that are close only because a relevant variable was omitted indicate that the distance is unsuitable. A genuine target discontinuity calls for revising the assumed neighbourhood or bound.

Passing a finite set of these checks provides evidence about the cases examined. A violation is evidence of a failed stability requirement when that requirement is justified; it is not automatically evidence that every discontinuity or sharp prediction change is an error.

Scope of the principle

Conservation of Distance connects model specification to model assessment. During construction, it makes assumptions about relevant similarity explicit. During debugging, it identifies prediction changes that exceed the variation allowed by those assumptions.

Its reach has limits. Successful learning can exploit global algebraic or compositional structure that raw local distance obscures. For example, parity changes whenever one bit flips. With unnormalised Hamming distance on bit strings and absolute distance on binary outputs, every binary-valued function is nevertheless 1-Lipschitz: distinct strings are at distance at least one. The bound holds, but supplies no useful preference for parity over other functions. A finite bound must therefore be assessed for its practical restrictiveness, not merely its existence.

The principle is grounded in smoothness and representation, with a scope determined by the task. If the target relationship is stable and the predictor is accurate, the predictions must satisfy the corresponding bound with an approximation-error allowance. A violation shows that at least one of these assumptions fails. Satisfaction does not establish accuracy or generalisation. Finite evaluation pairs provide a partial check, while a counterexample can reject the specified bound.

Applied to a new problem, this viewpoint requires specifying which input differences should affect the prediction, how the representation expresses those differences, and how the resulting behaviour will be evaluated on unseen data.


A companion post will examine query–key–value formulations as a way to compare prediction and retrieval computations.