The
Hamilton–Jacobi–Bellman (HJB) equation is a
partial differential equation which is central to
optimal control theory. Classical variational problems, for example, the
brachistochrone problem can be solved using this method. The HJB method can be generalized to
stochastic systems as well.
The solution of the HJB equation is the 'value function', which gives the optimal cost-to-go for a given
dynamical system with an associated cost function. The solution is open loop, but it also permits the solution of the closed loop problem.
The equation is a result of the theory of
dynamic programming which was pioneered in the 1950s by
Richard Bellman and coworkers.
[1] The corresponding discrete-time equation is usually referred to as the
Bellman equation. In continuous time, the result can be seen as an extension of earlier work in
classical physics on the
Hamilton-Jacobi equation by
William Rowan Hamilton and
Carl Gustav Jacob Jacobi.
Optimal control problems
Consider the following problem in deterministic optimal control
![\min \left\{ \int_0^T C[x(t),u(t)]\,dt + D[x(T)] \right\}](https://lh3.googleusercontent.com/blogger_img_proxy/AEn0k_thrPF5IZ5sruC-ngFOcOM1Zbq0bTdP7uP18AKiSS3mp7Zdksd5eyLexeOFNn_8q45oULrHqd4SntLxmWaKb9vZ0JCcBSQlRY4omKMFxPPLqKOfKl9kNygq7XR_o8z-c42bxOO7S2QIf1UaIJpECA=s0-d)
where C[] is the scalar cost rate function and D[] is a function that gives the salvage value or scrap value scalar at the final state,
x(t) is the system state vector,
x(0) is assumed given, and
u(t) for

is the control vector that we are trying to find.
The system must also be subject to
![\dot{x}(t)=F[x(t),u(t)]](https://lh3.googleusercontent.com/blogger_img_proxy/AEn0k_v1JZyc3VbTwGI0w2a4Wvheht0PL8SsUUw4RcvCDB7Xg6CAFPyczptqHjAgajKfd_yj8pe0rCmOYPUnpZxYXDvQcnX1WkfY6TDNQivWIqniI5LUkD8lCwmvvQ2comLpBE708DTxQTs1-5GrDYz_=s0-d)
where F[] gives the vector determining physical evolution of the state vector over time.
The partial differential equation
For this simple system, the Hamilton Jacobi Bellman partial differential equation is

subject to the terminal condition

where the

means the
dot product of the vectors a and b and

is the
gradient operator.
The unknown scalar
V(x,t) in the above PDE is the Bellman '
value function', which represents the cost incurred from starting in state
x at time
t and controlling the system optimally from then until time
T.
Deriving the equation
Intuitively HJB can be "derived" as follows. If
V(x(t),t) is the optimal cost-to-go function (also called the 'value function'), then by Richard Bellman's
principle of optimality, going from time
t to
t + dt, we have

Note that the
Taylor expansion of the last term is

where o(dt^2) denotes the terms in the Taylor expansion of higher order than one. Then if we cancel
V(x(t),t) on both sides, divide by
dt, and take the limit as
dt approaches zero, we obtain the HJB equation defined above.
Solving the equation
The HJB equation needs to be
solved backwards in time, starting from
t = T and ending at
t = 0.
[citation needed]
The HJB equation is a
necessary and sufficient condition for an optimum.
[2] If we can solve for
V then we can find from it a control
u that achieves the minimum cost.
In general case, the HJB equation does not have a classical (smooth) solution. Several notions of generalized solutions have been developed to cover such situations, including
viscosity solution (
Pierre-Louis Lions and
Michael Crandall),
minimax solution (
Andrei Izmailovich Subbotin), and others.
Extension to stochastic problems
The idea of solving a control problem by applying Bellman's principle of optimality and then working out backwards in time an optimizing strategy can be generalized to stochastic control problems. Consider similar as above

now with
![(X_t)_{t \in [0,T]}\,\!](https://lh3.googleusercontent.com/blogger_img_proxy/AEn0k_v2H8qbAn85oJSKJUGN0D9VloB7x08Yf1khTsmhTsZF7jKeOCDj401vsBWfRgGob4ln_2tGRSFBIRALiiIiLgE7M1uxhgUKMkTFzqJMJcvIODqLNtbsl-gbbe2JaKCludGzdO-VH1VPBTD41lgkaQ=s0-d)
the stochastic process to optimize and
![(u_t)_{t \in [0,T]}\,\!](https://lh3.googleusercontent.com/blogger_img_proxy/AEn0k_vfCn9QpmvLH49ZDkwXC0Pdbg8qAzkELcd7booXGoiomjJMpIJdNPFkou6gUTw6OtBz2gCdyoeBa-tE2BbCTuy3prP98rX30caAEeml8eG7iwcI1UFKf8QYF6HEfi9Y22Z_tQn2RGVFOdhAgask=s0-d)
the steering. By first using Bellman and then expanding
V(t,Xt) with
Itô's rule, one finds the deterministic HJB equation

where

represents the stochastic differentiation operator, and subject to the terminal condition

Note, that the randomness has disappeared. In this case a solution

of the latter does not necessarily solve the primal problem, it is a candidate only and a further verifying argument is required. This technique is widely used in Financial Mathematics to determine optimal investment strategies in the market (see for example
Merton's portfolio problem)