← All essays

Physics & literature

Does Light Know Where It Is Going?

Story of Your Life and the calculus of variations

We usually understand a life in the order it is lived. We make a choice, live with its consequences, and revise our sense of who we are. Only afterward do we look back and ask whether we would take the same path again. In Existentialism Is a Humanism, Jean-Paul Sartre places choice at the center of human existence: we become who we are through what we do. [1]

But reading Ted Chiang’s Story of Your Life left me wondering how much this view depends on the future remaining unknown. If we could see our lives all at once, including the joy and the grief still ahead, what would it mean to choose?

Chiang approaches this question through language. His narrator, the linguist Louise Banks, learns to read the writing of the heptapods, whose sentences take the form of interconnected designs. Each part is arranged in relation to the whole; the very first stroke depends on how the entire sentence will fit together. To write this way, one seems to need the ending in mind before beginning.

Their physics reflects this way of seeing. Variational principles, which describe a complete path through a condition on the whole, come naturally to the heptapods. To a human reader, this can sound almost teleological, as though events unfold in order to fulfill an ending. Where we tend to ask what earlier event caused the next one, they understand how each event belongs to the complete pattern. [2]

Chiang brings the difference into focus with Fermat’s principle. When light passes from air into water, it bends. We can describe that change with a rule relating the angles at the boundary, or consider the entire journey between two points and ask which path makes the travel time stationary. The second description seems to demand a strange kind of foresight: how can light set off in the right direction without knowing where it will arrive?

The connection stayed with me after I finished the story. I wanted to understand how a condition on an entire journey could express the same law as a rule that holds at each point. Following that question leads into the calculus of variations, and eventually back to the question of what it means to see a life as a whole.

To see how this works, let’s begin with a simpler question: why might a longer route take less time?

A longer path can take less time

Imagine moving through two regions at different speeds. Perhaps one has a smooth surface, while the other is difficult to cross. The shortest route in distance need not be the quickest route. A small detour through the faster region might save enough time to make it worthwhile.

Light presents a similar problem. In the usual geometrical-optics setting, light travels at different speeds in different media. Fermat’s principle says that the actual ray between fixed endpoints has a stationary travel time: a small change in the path produces no first-order change in that time. In the simple refraction problem we are about to consider, the stationary path is also the fastest one. [3]

Let us make this concrete. Suppose a horizontal boundary separates two media. The starting point is \(A=(0,a)\), above the boundary, and the destination is \(B=(d,-b)\), below it, where \(a,b,d>0\). The light crosses the boundary at \(P=(u,0)\).

A ray crossing between two media Light travels from A above a horizontal boundary through P on the boundary to B below it, bending toward the dashed normal in the slower medium. A (0, a)P (u, 0)B (d, −b) θ₁θ₂ MEDIUM 1 · v₁ MEDIUM 2 · v₂ < v₁
A change in the crossing point P changes both parts of the journey. The angles are measured from the normal.

Assume it travels in a straight line within each medium, at speeds \(v_1\) and \(v_2\). Its total travel time is

\[ T(u) = \frac{\sqrt{u^2+a^2}}{v_1} + \frac{\sqrt{(d-u)^2+b^2}}{v_2}. \]

The first term is the time from \(A\) to \(P\); the second is the time from \(P\) to \(B\). Changing \(u\) changes the route.

At the fastest route, moving the crossing point a little to the left or right cannot improve the travel time to first order. We therefore set

\[ \frac{dT}{du} = \frac{u}{v_1\sqrt{u^2+a^2}} - \frac{d-u}{v_2\sqrt{(d-u)^2+b^2}} = 0. \]

The fractions involving distances are the sines of the angles that the two segments make with the normal to the boundary. Calling these angles \(\theta_1\) and \(\theta_2\), we obtain

\[ \frac{\sin\theta_1}{v_1} = \frac{\sin\theta_2}{v_2}. \]

Since the refractive index is \(n=c/v\), this becomes

\[ \boxed{n_1\sin\theta_1=n_2\sin\theta_2.} \]

This is Snell’s law. The calculation also gives the equation a geometric meaning. Moving the crossing point changes both parts of the journey. At the correct crossing point, the time gained on one part exactly balances the time lost on the other, to first order.

We began with a statement about the complete route and obtained a rule about the angles at a single boundary.

So far, however, we have only varied one number, \(u\). What happens when the whole curve is unknown?

From choosing a number to choosing a curve

Suppose we want to find the shortest curve joining two fixed points. Write the curve as \(y(x)\), with

\[ y(x_0)=y_0, \qquad y(x_1)=y_1. \]

A small segment of the curve has length \(ds=\sqrt{dx^2+dy^2}\), so the total length is

\[ \mathcal J[y] = \int_{x_0}^{x_1}\sqrt{1+y'(x)^2}\,dx. \]

This looks like an ordinary integral, but notice what we are allowed to change. We are not choosing a value of \(x\), nor asking where to stand on an existing curve. We are choosing the function \(y(x)\) itself.

The expression \(\mathcal J\) takes an entire function as input and returns a number. Such an object is called a functional. The square brackets in \(\mathcal J[y]\) emphasize that its input is a function rather than a single numerical value.

There is a useful way to make this less abstract. Imagine drawing the curve through a large number of adjustable points. Its shape is then specified by their heights,

\[ (y_1,y_2,\ldots,y_N). \]

Finding the shortest curve becomes an optimization problem over these \(N\) numbers. As the points become more closely spaced, the finite list approaches a continuous function.

In this sense, the calculus of variations extends multivariable calculus from a finite set of adjustable numbers to an entire adjustable shape.

The central question remains familiar: how does the quantity we care about change when we slightly change its input?

What it means to vary a path

For an ordinary differentiable function \(f(z)\), a small change in its argument gives

\[ f(z+\varepsilon) = f(z)+\varepsilon f'(z)+O(\varepsilon^2). \]

At an interior minimum, \(f'(z)=0\). Otherwise, choosing the appropriate sign of \(\varepsilon\) would lower the function.

To do the same thing with a curve, we need a way to describe a small deformation. We write

\[ y_\varepsilon(x)=y(x)+\varepsilon\eta(x). \]

Here, \(\eta(x)\) specifies the shape of the deformation, while \(\varepsilon\) controls its size. The deformation might lift the middle of the curve, lower a short section, or bend different regions in different directions.

Because we want to keep the endpoints fixed, we require

\[ \eta(x_0)=\eta(x_1)=0. \]

It is worth pausing over the two different changes involved. Changing \(x\) moves us along a particular curve. Changing \(\varepsilon\) moves us from one curve to a neighboring curve. The derivative \(y'(x)\) describes the first kind of change; a variation describes the second.

Once we choose \(\eta\), the expression \(\mathcal J[y+\varepsilon\eta]\) is just an ordinary function of the single number \(\varepsilon\). We can differentiate it:

\[ \boxed{ \delta\mathcal J[y;\eta] = \left. \frac{d}{d\varepsilon} \mathcal J[y+\varepsilon\eta] \right|_{\varepsilon=0}. } \]

This is the first variation of \(\mathcal J\) in the direction \(\eta\).

A curve \(y_*\) is stationary when

\[ \boxed{ \delta\mathcal J[y_*;\eta]=0 \quad \text{for every allowed deformation }\eta. } \]

The words “for every” matter. A curve is not stationary merely because one particular deformation leaves its value unchanged to first order. Every allowed direction must do so.

Geometrically, this is the counterpart of a zero gradient. But the geometry now belongs to a space of curves: each point in this space represents a complete curve, and each direction represents a possible deformation.

Stationarity does not mean that the curve itself is horizontal. It means that the functional has no first-order slope in any allowed direction.

Nor does stationarity automatically mean a minimum. A zero gradient can also occur at a maximum or a saddle point. This is why “stationary action” is more precise than the familiar phrase “least action.”

How an equation for the whole becomes an equation at each point

We can now derive the main result. Consider a functional of the form

\[ \mathcal J[y] = \int_{x_0}^{x_1}F(x,y(x),y'(x))\,dx. \]

The function \(F\) assigns a local contribution based on the position \(x\), the value \(y\), and the slope \(y'\). The integral adds these contributions into a number associated with the whole curve.

Assume that the functions are sufficiently smooth for the differentiations and integrations below. Substituting the deformed curve gives

\[ \mathcal J[y+\varepsilon\eta] = \int_{x_0}^{x_1} F(x,y+\varepsilon\eta,y'+\varepsilon\eta')\,dx. \]

Differentiating with respect to \(\varepsilon\), using the chain rule, yields

\[ \delta\mathcal J[y;\eta] = \int_{x_0}^{x_1} \left( \frac{\partial F}{\partial y}\eta + \frac{\partial F}{\partial y'}\eta' \right)dx. \]

There are two terms because deforming the curve changes both its height and its slope.

We would like to express everything in terms of \(\eta\). We cannot simply treat \(\eta\) and \(\eta'\) as independent choices: one is the derivative of the other. Integration by parts resolves this:

\[ \int_{x_0}^{x_1} \frac{\partial F}{\partial y'}\eta'\,dx = \left[ \frac{\partial F}{\partial y'}\eta \right]_{x_0}^{x_1} - \int_{x_0}^{x_1} \frac{d}{dx} \left(\frac{\partial F}{\partial y'}\right) \eta\,dx. \]

The boundary term vanishes because \(\eta\) is zero at both endpoints. We are left with

\[ \delta\mathcal J[y;\eta] = \int_{x_0}^{x_1} \left[ \frac{\partial F}{\partial y} - \frac{d}{dx} \left(\frac{\partial F}{\partial y'}\right) \right]\eta(x)\,dx. \]

Now comes the decisive step. Suppose the expression in brackets were positive around some point. We could choose a small, positive deformation supported only in that region. The integral would then be positive, contradicting stationarity. The same reasoning rules out a negative value.

Since we can place such a deformation anywhere in the interior, the bracketed expression must vanish everywhere there:

\[ \boxed{ \frac{d}{dx} \left(\frac{\partial F}{\partial y'}\right) - \frac{\partial F}{\partial y} = 0. } \]

This is the Euler–Lagrange equation, the local condition associated with stationarity of this integral functional. [4]

The derivation explains why a global statement can have local consequences. Requiring the integral to be stationary under every small deformation includes deformations concentrated in arbitrarily small regions. The whole cannot be stationary if one of its parts still admits a first-order change.

For a quick check, return to the length functional, where \(F=\sqrt{1+y'^2}\). Since \(F\) does not depend explicitly on \(y\), the equation becomes

\[ \frac{d}{dx} \left(\frac{y'}{\sqrt{1+y'^2}}\right)=0. \]

The expression inside the derivative is constant, which means \(y'\) is constant. The curve is therefore a straight line, as expected.

From the shape of a curve to the motion of a particle

So far, we have developed a mathematical method. To use it in physics, we must specify the functional that describes the system.

For a particle of constant mass \(m\), moving in one dimension under a potential \(V(q)\), define the Lagrangian as

\[ L(q,\dot q) = \frac12m\dot q^2-V(q). \]

It is the kinetic energy minus the potential energy. The corresponding functional is the action:

\[ \boxed{ S[q] = \int_{t_0}^{t_1} \left(\frac12m\dot q^2-V(q)\right)dt. } \]

The input is now a complete motion \(q(t)\), including not only where the particle goes but when it reaches each position. Hamilton’s principle states that the physical trajectory makes this action stationary under variations that keep the positions at the two endpoint times fixed. [5]

Action is not energy. It has units of energy multiplied by time, and stationarity of action is not a claim that a particle uses the least energy or travels in the shortest time. In this problem, the starting and finishing times are already fixed.

We can apply the Euler–Lagrange equation by replacing the spatial variable \(x\) with time \(t\):

\[ \frac{d}{dt} \left(\frac{\partial L}{\partial\dot q}\right) - \frac{\partial L}{\partial q} = 0. \]

For our Lagrangian,

\[ \frac{\partial L}{\partial\dot q}=m\dot q, \qquad \frac{\partial L}{\partial q}=-V'(q). \]

The result is

\[ \boxed{m\ddot q=-V'(q).} \]

The negative derivative of the potential is the force. We have recovered Newton’s second law.

For this conservative mechanical system, the local equation of motion and the stationary-action description are two ways of expressing the same dynamics.

There is an important distinction here. Calculus alone did not tell us to choose kinetic energy minus potential energy. That choice contains physical information. What we have derived is the relationship between that choice and the resulting motion.

A variational principle is not the assertion that any attractive-looking objective will reproduce nature.

A simple example: motion without a force

When \(V=0\), the action is

\[ S[q]=\frac m2\int_{t_0}^{t_1}\dot q^2\,dt. \]

The Euler–Lagrange equation says \(\ddot q=0\), so the stationary trajectory has constant velocity:

\[ q_*(t) = q_0+ \frac{q_1-q_0}{t_1-t_0}(t-t_0). \]

We can also check directly what happens when we deform it. Let \(q=q_*+\varepsilon\eta\), with \(\eta\) zero at the endpoints. Expanding the square gives a cross term proportional to

\[ \int_{t_0}^{t_1}\dot\eta\,dt = \eta(t_1)-\eta(t_0) = 0. \]

Consequently,

\[ S[q_*+\varepsilon\eta]-S[q_*] = \frac{m\varepsilon^2}{2} \int_{t_0}^{t_1}\dot\eta^2\,dt \geq 0. \]

In this example, the stationary trajectory really is a minimum. Any nonzero deformation increases the action, but only at second order in its size.

This calculation also clarifies what the alternatives are. They are candidate motions satisfying the endpoint conditions, not necessarily motions that obey Newton’s law. We compare these possibilities to identify which ones satisfy the physical law; we do not assume that law before making the comparison.

Two ways of describing the same motion

We can now see why the physics that feels natural to the heptapods can seem so strange to us. The difference begins with how we pose the problem.

In the local description of mechanics, we specify an initial position and velocity, then solve the equation of motion. In the variational description, we specify positions at two times and ask which connecting trajectories make the action stationary. These are different ways of posing a problem, but the equations we derived show how they are related.

Imagine first solving the initial-value problem. The particle follows a trajectory and reaches some position at \(t_1\). Now take the two endpoints of that trajectory and compare it with neighboring paths between them. The trajectory you already found will satisfy the stationary-action condition. No information has traveled backward. We have described the same motion from a different perspective.

Conversely, specifying two endpoint positions does not always guarantee a unique trajectory. Depending on the system and the boundary conditions, there may be several solutions or none. The variational formulation is not a general promise that a beginning and an ending determine one inevitable story.

The key distinction is between a law and an account of how a system computes its behavior. An equation involving an entire trajectory does not imply that the particle evaluates every candidate trajectory before moving. Nor does \(\delta S=0\) mean that the particle performs gradient descent on its action. Stationarity characterizes the motion; it does not describe a search procedure. What looked like foresight arose because we confused a condition used in our description with information the physical object must possess.

Knowing the ending

We can now return to the question we began with.

Does light know where it is going?

Nothing in the derivation requires it to. We have found two ways of describing a path. One follows the motion locally, relating what happens at a point to what happens nearby. The other considers the complete trajectory and asks what happens when we deform it.

Under the assumptions we have used, the stationary-action condition and the local equation of motion express the same dynamics:

\[ \delta S[q]=0 \qquad\Longleftrightarrow\qquad \frac{d}{dt}\left(\frac{\partial L}{\partial\dot q}\right) - \frac{\partial L}{\partial q} =0. \]

The endpoint has not sent a message backward in time. We have changed the way we describe the motion, not the direction in which the particle moves.

There is also something important in what this equation does not say.

The familiar phrase “least action” encourages us to imagine nature finding the best possible route. But the condition we derived was stationarity: a small, allowed deformation produces no first-order change in the action. Even when the action is a minimum, the comparison concerns a particular quantity, defined by the system we are studying. Calling a path “best” already assumes that we know what to measure.

That question becomes less straightforward when we return to Chiang’s story. We often look back on our own lives by imagining alternatives: a different choice, another direction, a different ending. In the calculation, we could compare paths because we had a quantity to evaluate. What would it mean to make that comparison for a life?

For Louise, knowing the shape of her life does not supply a calculation that tells her whether it is worth living. There is no given functional that assigns the right weights to everything she will love and everything she will lose.

Even where the action is minimized, it is not a measure of happiness, suffering, or the worth of a journey. The mathematics gives us no reason to call a minimum fortunate, or a maximum tragic.

Yet the distinction between seeing a whole and moving through its parts remains.

From outside, a trajectory can be represented by a single curve. From within, there is still the passage from one moment to the next. Seeing the destination does not make the intervening points disappear. Knowing that something will end does not mean it has already been experienced.

The story holds these two perspectives together without making either one comfortable. Louise knows, and she chooses. The choice does not remove the pain she foresees; the pain does not erase the joy.

Perhaps that is why the language of a variational principle feels so at home here. It lets us ask about the shape of an entire life without pretending that the answer must be a better outcome, a shorter route, or a smaller loss.

The mathematical question has become a personal one, and the derivation can take us no further:

From the beginning I knew my destination, and I chose my route accordingly. But am I working toward an extreme of joy, or of pain? Will I achieve a minimum, or a maximum?

References

  1. Jean-Paul Sartre. Existentialism Is a Humanism. Translated by Carol Macomber. Yale University Press, 2007. On choice, responsibility, and becoming who we are through our actions. ↩
  2. Ted Chiang. “Story of Your Life.” In Stories of Your Life and Others. Vintage, 2016 edition. The story that prompted this essay. ↩
  3. Richard P. Feynman, Robert B. Leighton, and Matthew Sands. “Optics: The Principle of Least Time.” The Feynman Lectures on Physics, Vol. I, Chapter 26. Caltech online edition. Fermat’s principle and its connection to refraction. ↩
  4. Iain Stewart. “A Review of Analytical Mechanics.” 8.09 Classical Mechanics III, Chapter 1. MIT OpenCourseWare, Fall 2014. Further mathematical detail on variational principles and Lagrangian mechanics. ↩
  5. Richard P. Feynman, Robert B. Leighton, and Matthew Sands. “The Principle of Least Action.” The Feynman Lectures on Physics, Vol. II, Chapter 19. Caltech online edition. Stationary action, Newton’s law, and why “least” need not mean a minimum. ↩