What is Delay Flow Matching doing?
Ordinary flow matching fits a marginal vector field under a sampled coupling of source and target endpoints. Delay Flow Matching (DFM) also conditions the vector field on an earlier point from the same constructed path. When both points come from an invertible deterministic interpolation of one endpoint pair, they identify that pair. The DFM regression target is then the target vector for that pair, rather than an average over all pairs passing through the current state.
The two regression targets
Let a coupling pair source \(x_0\) with target \(x_1\), and define the constructed path and its target vector by
\[ x_t=\alpha _tx_0+\beta _tx_1, \qquad u_t=\dot \alpha _tx_0+\dot \beta _tx_1. \]
Under squared loss, ordinary flow matching has the population solution
\[ v^*_{\mathrm {FM}}(t,x)=\mathbb E[u_t\mid x_t=x]. \]
If target vectors from several endpoint pairs meet at \(x\), this conditional mean averages them. DFM instead uses \(v_\theta (t,x_t,x_{t-\tau })\), whose population solution is
\[ v^*_{\mathrm {DFM}}(t,x,y) =\mathbb E[u_t\mid x_t=x,\ x_{t-\tau }=y]. \]
The two objectives therefore estimate different conditional expectations.
When two points identify the endpoints
Set \(s=t-\tau \), and assume that both \(x_t\) and \(x_s\) follow the endpoint interpolation above. This assumption excludes initial-history intervals in which a DFM construction defines \(x_s\) by a separate history function. If
\[ \Delta _{t,s}=\alpha _t\beta _s-\alpha _s\beta _t\neq 0, \]
then the coordinatewise \(2\times 2\) system is invertible:
\[ x_0=\frac {\beta _sx_t-\beta _tx_s}{\Delta _{t,s}}, \qquad x_1=\frac {-\alpha _sx_t+\alpha _tx_s}{\Delta _{t,s}}. \]
Substitution gives the target vector directly from the two DFM inputs,
\[ u_t(x_t,x_s)= \frac {(\dot \alpha _t\beta _s-\dot \beta _t\alpha _s)x_t +(\alpha _t\dot \beta _t-\beta _t\dot \alpha _t)x_s} {\Delta _{t,s}}. \]
For the straight path \(x_r=(1-r)x_0+rx_1\), this reduces to
\[ u_t=x_1-x_0=\frac {x_t-x_{t-\tau }}{\tau } \qquad (t>\tau ). \]
Under the determinant condition, the DFM target is deterministic given its two state inputs, whereas ordinary flow matching still estimates \(\mathbb E[u_t\mid x_t]\) ( Lipman et al. 2022; Tong et al. 2023). Conditioning on information that identifies the endpoint pair makes DFM analogous to Augmented Bridge Matching, which conditions on the initial sample to preserve the training coupling ( De Bortoli et al. 2023).
This inversion concerns samples on the constructed training path. During inference, \(x_{t-\tau }\) is generated by the fitted delay equation, so endpoint recovery applies only when the resulting state pair still satisfies the assumed interpolation relation.
What an endpoint readout requires
For a current state \(x_t\) and a fitted vector prediction \(\widehat u_t\), the linear-in-endpoints schedule implies
\[ \widehat x_1= \frac {\alpha _t\widehat u_t-\dot \alpha _tx_t} {\alpha _t\dot \beta _t-\dot \alpha _t\beta _t}, \]
provided the denominator is nonzero. This identity gives the endpoint implied by a vector prediction; it does not by itself make the first model call a pair-specific endpoint predictor.
At time zero with constant history and a product coupling, the input contains \(x_0\) but no information about the independently sampled \(x_1\). The squared-loss solution is then a conditional mean vector, and the formula above returns a conditional mean endpoint. A pair-specific endpoint estimate therefore requires an informative history function or a delayed state produced later in the rollout.
Why the coupling matters
Once the two-state input identifies an endpoint pair, the training target retains the coupling used to sample that pair. A product coupling therefore supplies arbitrary independent pairings, while paired observations, optimal transport, minibatch matching, or a bridge construction supply pairings with additional structure ( Pooladian et al. 2023; Shi et al. 2023). This choice determines which pair-specific target vectors DFM fits.
Scope: delay differential equations
For a delay differential equation,
\[ \dot x(t)=f(x(t),x(t-\tau )), \]
the earlier state is part of the physical state and need not encode two interpolation endpoints. The endpoint-inversion argument applies only to paths constructed from endpoint pairs, not to general delayed dynamics ( Zhu et al. 2021; Zhao et al. 2026).