Skip to main content

Section 2.7 Inverse and Implicit Functions

The inverse and implicit function theorems explain when a differentiable map can be solved locally for some of its variables in terms of the others. The key hypothesis is that the derivative at the point of interest is invertible.

Definition 2.7.1. Continuously Differentiable.

Let \(U \subseteq \R^n\) be open and let \(f \colon U \to \R^m\text{.}\) We say that \(f\) is continuously differentiable on \(U\text{,}\) or that \(f\) is \(C^1\text{,}\) if \(f\) is differentiable at every point of \(U\) and the derivative map \(\va \mapsto Df(\va)\) is continuous on \(U\text{.}\)
For the inverse function theorem, it is convenient to express continuous differentiability in a two-point Caratheodory form.

Proof.

Suppose first that \(f\) is \(C^1\) on \(U\text{,}\) and choose an open ball \(B \subseteq U\) centered at \(\va\) such that the line segment joining any two points of \(B\) stays inside \(U\text{.}\) For \(\vx,\vy \in B\text{,}\) define
\begin{equation*} \Psi(\vx,\vy) := \int_0^1 Df(\vy+t(\vx-\vy))\,dt. \end{equation*}
Since \(Df\) is continuous, \(\Psi\) is continuous on \(B \times B\text{.}\) Let \(h(t)=f(\vy+t(\vx-\vy))\text{.}\) By the chain rule,
\begin{equation*} h'(t)=Df(\vy+t(\vx-\vy))(\vx-\vy). \end{equation*}
The one-variable Fundamental Theorem of Calculus gives
\begin{equation*} f(\vx)-f(\vy) = h(1)-h(0) = \int_0^1 Df(\vy+t(\vx-\vy))(\vx-\vy)\,dt = \Psi(\vx,\vy)(\vx-\vy). \end{equation*}
Conversely, assume such a function \(\Psi\) exists on a ball \(B\text{.}\) Fix \(\va \in B\text{.}\) Then
\begin{equation*} f(\vx)-f(\va)=\Psi(\vx,\va)(\vx-\va) \end{equation*}
for \(\vx \in B\text{.}\) Since \(\vx \mapsto \Psi(\vx,\va)\) is continuous at \(\va\text{,}\) the Caratheodory criterion shows that \(f\) is differentiable at \(\va\text{,}\) with
\begin{equation*} Df(\va)=\Psi(\va,\va). \end{equation*}
Because \(\va \mapsto \Psi(\va,\va)\) is continuous, the derivative map is continuous on \(B\text{.}\) Hence \(f\) is \(C^1\) on \(B\text{.}\)

Proof.

Let \(\vb=f(\va)\) and let \(A=Df(\va)\text{.}\) By PropositionΒ 2.7.2, after shrinking to a small open ball \(B_r(\va) \subseteq U\) we may assume there is a continuous matrix-valued function \(\Psi \colon B_r(\va)\times B_r(\va)\to M_{n\times n}\) such that
\begin{equation*} f(\vx)-f(\vy)=\Psi(\vx,\vy)(\vx-\vy) \end{equation*}
for all \(\vx,\vy \in B_r(\va)\text{,}\) and \(\Psi(\va,\va)=A\text{.}\) Shrink \(r\) further so that the closed ball \(\overline{B_r(\va)}\) is contained in \(U\) and
\begin{equation*} \norm{I-A^{-1}\Psi(\vx,\vy)}_{\mathrm{op}} \le \frac12 \end{equation*}
for all \(\vx,\vy \in \overline{B_r(\va)}\text{.}\)
We first show that \(f\) is injective on \(\overline{B_r(\va)}\text{.}\) If \(f(\vx)=f(\vy)\text{,}\) then
\begin{equation*} A^{-1}\Psi(\vx,\vy)(\vx-\vy)=0. \end{equation*}
Writing \(A^{-1}\Psi(\vx,\vy)=I-E\) with \(\norm{E}_{\mathrm{op}}\le \frac12\text{,}\) we get
\begin{equation*} \vx-\vy = E(\vx-\vy). \end{equation*}
Hence \(\norm{\vx-\vy}\le \frac12 \norm{\vx-\vy}\text{,}\) so \(\vx=\vy\text{.}\)
Now let
\begin{equation*} W := \left\{ \vz \in \R^n : \norm{A^{-1}(\vz-\vb)} \lt \frac{r}{2} \right\}. \end{equation*}
For each \(\vz \in W\text{,}\) define a map \(T_{\vz} \colon \overline{B_r(\va)} \to \R^n\) by
\begin{equation*} T_{\vz}(\vx):=\vx-A^{-1}(f(\vx)-\vz). \end{equation*}
If \(\vx,\vy \in \overline{B_r(\va)}\text{,}\) then
\begin{equation*} T_{\vz}(\vx)-T_{\vz}(\vy) = \bigl(I-A^{-1}\Psi(\vx,\vy)\bigr)(\vx-\vy), \end{equation*}
\begin{equation*} \norm{T_{\vz}(\vx)-T_{\vz}(\vy)} \le \frac12 \norm{\vx-\vy}. \end{equation*}
Thus \(T_{\vz}\) is a contraction. Also,
\begin{equation*} \norm{T_{\vz}(\vx)-\va} \le \norm{T_{\vz}(\vx)-T_{\vz}(\va)} + \norm{A^{-1}(\vz-\vb)} \lt \frac12 r + \frac12 r = r. \end{equation*}
Hence \(T_{\vz}\) maps the complete metric space \(\overline{B_r(\va)}\) to itself. By the contraction mapping theorem, \(T_{\vz}\) has a unique fixed point \(\vx_{\vz} \in \overline{B_r(\va)}\text{.}\) The fixed point equation \(T_{\vz}(\vx_{\vz})=\vx_{\vz}\) is exactly \(f(\vx_{\vz})=\vz\text{.}\) Therefore \(f\) maps \(\overline{B_r(\va)}\) onto \(W\text{.}\)
Let \(V=f^{-1}(W)\cap B_r(\va)\text{.}\) Then \(f \colon V \to W\) is bijective. To understand the inverse, let \(\vz,\vw \in W\) and set \(\vx=f^{-1}(\vz)\text{,}\) \(\vy=f^{-1}(\vw)\text{.}\) Then
\begin{equation*} \vz-\vw = \Psi(\vx,\vy)(\vx-\vy), \end{equation*}
\begin{equation*} f^{-1}(\vz)-f^{-1}(\vw) = \Psi(f^{-1}(\vz),f^{-1}(\vw))^{-1}(\vz-\vw). \end{equation*}
Since \(\norm{I-A^{-1}\Psi(\vx,\vy)}_{\mathrm{op}} \le \frac12\text{,}\) the matrices \(\Psi(\vx,\vy)\) are invertible, and
\begin{equation*} \norm{f^{-1}(\vz)-f^{-1}(\vw)} \le 2\norm{A^{-1}}_{\mathrm{op}} \norm{\vz-\vw}. \end{equation*}
Thus \(f^{-1}\) is continuous on \(W\text{.}\) The coefficient function
\begin{equation*} \Phi(\vz,\vw) := \Psi(f^{-1}(\vz),f^{-1}(\vw))^{-1} \end{equation*}
is therefore continuous on \(W \times W\text{,}\) and
\begin{equation*} f^{-1}(\vz)-f^{-1}(\vw)=\Phi(\vz,\vw)(\vz-\vw). \end{equation*}
By PropositionΒ 2.7.2, \(f^{-1}\) is \(C^1\) on \(W\text{.}\) Finally,
\begin{equation*} D(f^{-1})(\vb)=\Phi(\vb,\vb)=\Psi(\va,\va)^{-1}=A^{-1}. \end{equation*}

Example 2.7.4.

Let \(f \colon \R^2 \to \R^2\) be defined by
\begin{equation*} f(x,y)=\langle x+y^2, y \rangle. \end{equation*}
Use the inverse function theorem to see that \(f\) has a local inverse near \(\langle 0,0 \rangle\text{.}\)
Solution.
The derivative matrix is
\begin{equation*} Df(x,y) = \begin{bmatrix} 1 \amp 2y \\ 0 \amp 1 \end{bmatrix}, \end{equation*}
so \(\det Df(x,y)=1\) for every \(\langle x,y \rangle\text{.}\) In particular, \(Df(0,0)\) is invertible, so the inverse function theorem guarantees a \(C^1\) inverse near \(\langle 0,0 \rangle\text{.}\)
In fact, this inverse can be written explicitly:
\begin{equation*} f^{-1}(u,v)=\langle u-v^2, v \rangle. \end{equation*}
The implicit function theorem is a direct consequence of the inverse function theorem. We separate the variables into two blocks: \(\vx \in \R^n\) and \(\vy \in \R^m\text{.}\) For a map \(F(\vx,\vy)\text{,}\) the matrices \(D_{\vx}F\) and \(D_{\vy}F\) denote the derivatives with respect to the \(\vx\)-variables and the \(\vy\)-variables.

Proof.

Define \(H \colon U \to \R^n \times \R^m\) by
\begin{equation*} H(\vx,\vy)=\langle \vx, F(\vx,\vy) \rangle. \end{equation*}
Its derivative at \((\va,\vb)\) has block form
\begin{equation*} DH(\va,\vb) = \begin{bmatrix} I_n \amp 0 \\ D_{\vx}F(\va,\vb) \amp D_{\vy}F(\va,\vb) \end{bmatrix}. \end{equation*}
Because this matrix is block triangular and \(D_{\vy}F(\va,\vb)\) is invertible, \(DH(\va,\vb)\) is invertible. By TheoremΒ 2.7.3, there are open neighborhoods \(V\) of \((\va,\vb)\) and \(W\) of \(\langle \va,\vz \rangle\) such that \(H \colon V \to W\) is bijective, with \(C^1\) inverse \(J=H^{-1}\text{.}\)
Shrinking \(W\) if necessary, we may write \(W=A \times C\text{,}\) where \(A\) is an open neighborhood of \(\va\) and \(C\) is an open neighborhood of \(\vz\text{.}\) Let
\begin{equation*} J(\vu,\vw)=\langle P(\vu,\vw), Q(\vu,\vw) \rangle. \end{equation*}
Since the first component of \(H(\vx,\vy)\) is just \(\vx\text{,}\) the identity \(H(J(\vu,\vw))=\langle \vu,\vw \rangle\) forces \(P(\vu,\vw)=\vu\text{.}\) Define
\begin{equation*} g(\vx):=Q(\vx,\vz) \qquad (\vx \in A). \end{equation*}
\begin{equation*} H(\vx,g(\vx))=H(J(\vx,\vz))=\langle \vx,\vz \rangle, \end{equation*}
so \(F(\vx,g(\vx))=\vz\) for every \(\vx \in A\text{.}\) Since \(J(V)\) lies in \(V\text{,}\) the values of \(g\) lie in some open neighborhood \(B\) of \(\vb\text{.}\)
If \(\vx \in A\) and \(\vy \in B\) satisfy \(F(\vx,\vy)=\vz\text{,}\) then \(H(\vx,\vy)=\langle \vx,\vz \rangle = H(\vx,g(\vx))\text{.}\) Since \(H\) is injective on \(V\text{,}\) we get \(\vy=g(\vx)\text{.}\) This proves the uniqueness statement.
Finally, differentiate the identity \(F(\vx,g(\vx))=\vz\) at \(\vx=\va\text{.}\) By the chain rule,
\begin{equation*} D_{\vx}F(\va,\vb) + D_{\vy}F(\va,\vb)\,Dg(\va)=0. \end{equation*}
Multiplying on the left by \([D_{\vy}F(\va,\vb)]^{-1}\) gives
\begin{equation*} Dg(\va) = -[D_{\vy}F(\va,\vb)]^{-1}D_{\vx}F(\va,\vb). \end{equation*}

Example 2.7.6.

Show that near \(\langle 0,1 \rangle\text{,}\) the equation
\begin{equation*} x^2+y^2=1 \end{equation*}
determines \(y\) as a \(C^1\) function of \(x\text{.}\)
Solution.
Let \(F(x,y)=x^2+y^2-1\text{.}\) Then
\begin{equation*} F(0,1)=0, \qquad D_yF(x,y)=2y. \end{equation*}
Since \(D_yF(0,1)=2 \ne 0\text{,}\) the implicit function theorem applies at \(\langle 0,1 \rangle\text{.}\) Therefore there is a unique \(C^1\) function \(y=g(x)\) defined for \(x\) near \(0\) such that \(g(0)=1\) and \(x^2+g(x)^2=1\text{.}\)
In this example, the implicit function is the upper semicircle \(g(x)=\sqrt{1-x^2}\text{.}\) The derivative formula gives
\begin{equation*} g'(x)=-\frac{F_x(x,g(x))}{F_y(x,g(x))} = -\frac{2x}{2g(x)} = -\frac{x}{g(x)}. \end{equation*}