October 6th, 2026
Notes (PDF)We cover the Jordan Normal Form. (Personally, I do not like Artin's proof, and I will explain why at the bottom). We follow LADR's exposition to the Jordan Normal Form, and then consider Artin's.
Let be a linear operator on a finite-dimensional vector space.
.
For any such that , we have that
.
Let , and let be such that . Then, for any , note that .
Fix any : we seek to show . We already have that by , so we show the opposite. For any , note that
Thus . Thus as desired.
We simply show that , and the result follows by . Suppose they are not equal. By contrapositive of , this means that the inclusion chain
is strict. For to hold true, we need that . The most conservative assumption possible is that going from to increases the dimension by only . Following this, we have that , which is not possible since the kernel is a subspace of the domain, , which is only of dimension . Thus we have reached a contradiction.
It is noteworthy that we have stability of image chains, simply in the reverse direction. That is,
The result in and hold for this chain as well, so the image chain is as stable as the kernel chain. The proof of this is almost immediate by rank nullity, so we do not bother. Just note that Jordan theory can be approached either from the perspective of the kernel chain or the image chain. We use the kernel chain.
It is a desirable result that (we use internally, meaning that and ). Special conditions on , such as it being diagonalizable, make this statement true. We see, however, there is a slightly weaker statement that holds true for all operators.
Suppose is a linear operator on a finite-dimensional vector space. Then
Let . We show that . Indeed, suppose . Then , and there is some such that . Then note , and so by the last theorem. Then as desired.
Since these spaces don't intersect, we have that by the inclusion-exclusion principle, with the last equality by rank nullity. Thus we have the desired result.
In studying linear operators, one of our main desires is to make an operator as “simple” as possible. We want to be able to describe the action of an operator as cleanly as possible. One of the simplest ideas is that, for any vector space and a linear operator , we would like a decomposition of the form
where each is -invariant. Then the action of can be very quickly realized as the perfect combination of restricted to each of these , by invariance and the direct sum. Even more trivial would be if each were one-dimensional, in which case simply acts by scaling. This is the result diagonalization gives us: if is a finite-dimensional vector space and the operator is diagonalizable with eigenvalues , then we have that
where restricted to an eigenspace simply acts like the operator .
Of course, we must always keep in mind the inherent tradeoff between the generality of a theorem and the strength of its conclusions. If a result applies to a larger class of structures, it naturally cannot be as rigid as a result tailored to a smaller, more specialized class, since broadening our scope leaves us with fewer shared properties to exploit. Accommodating a wider variety of mathematical objects forces us to account for increasingly complex behaviors that simpler structures naturally avoid. Consequently, if we want a classification that applies universally across this broader space, we must compromise by accepting a slightly less pristine structural form.
Diagonalization is an extremely restrictive condition, but if we are able to show that an operator is diagonalizable, we get a beautiful structural result that lets us very quickly understand the structure of the operator. We aim to find a similar result for a broader class of operators, at the cost of a slightly weaker result. We will eventually show that every operator on a complex vector space admits this same type of decomposition, except the invariant subspaces need not be one-dimensional, and operators restricted to these subspaces are not as simple as just scaling.
Note that all operators are not necessarily diagonalizable, specifically because they may not have enough linearly independent eigenvectors to form a basis that diagonalizes the operator. We remedy this situation by generalizing the notion of eigenvectors.
Suppose is an eigenvalue of . We say some is a generalized eigenvector of if there exists some such that
Equivalently, we demand that .
Thus generalized eigenvectors are, fittingly, generalizations of eigenvectors. If , then we have a traditional eigenvector. Note that there is no such thing as a “generalized eigenvalue” since, if , then (if it were, then the kernel chain would stabilize at this point, since ).
By standard properties of our kernel chain, note that is a generalized eigenvector of corresponding to if and only if .
By extending our notion of “eigenvectors” to generalized eigenvectors, we find that operators on a complex vector space have enough linearly independent generalized eigenvectors to form a basis.
Let be an operator on a finite-dimensional, complex vector space. Then admits a basis of generalized eigenvectors.
Let . We proceed by induction on . Note this result is true for since every nonzero vector is an eigenvector.
Suppose this result is true for all values less than . Since is an operator on a complex vector space, it admits an eigenvalue (for a quick proof as to why, note that the characteristic polynomial of must split over ). Note that
as per the previous theorem. If it turns out that , then we have that , and so has a basis of generalized eigenvectors obviously. Otherwise, we have that both are of dimension less than . By the inductive hypothesis, each has a basis of generalized eigenvectors. Adjoining them gives a basis of generalized eigenvectors for , completing the proof.
Note that we require that be a complex vector space, since an operator on a complex vector space is guaranteed an eigenvalue (this is not true for operators on real vector spaces). being a complex vector space will consistently be required for future theorems due to this fact.
Eigenvectors have nice properties. There is exactly one eigenvalue corresponding to an eigenvector, and eigenvectors corresponding to different eigenvalues are linearly independent. These are not obviously true for generalized eigenvectors, but it turns out they are. We present the following theorems but omit the proof.
Let be an operator on a finite-dimensional, complex vector space. If is a generalized eigenvector of , then there is a unique such that .
Let be an operator on a finite-dimensional, complex vector space. Generalized eigenvectors of corresponding to different eigenvalues are linearly independent.
We now introduce the class of operators that will appear naturally on each generalized eigenspace.
A linear operator is nilpotent if for some . The smallest such is called the nilpotency index of .
If is nilpotent and , then in fact by stability of the kernel chain.
We now package the generalized eigenvectors corresponding to one eigenvalue into a single subspace.
Let be an eigenvalue of . The generalized eigenspace corresponding to is
By stability of kernel chains, if , then . In particular, is a subspace. It is also -invariant, since commutes with .
More importantly, note that on , the operator is nilpotent. Indeed,
Thus for some nilpotent operator . Rearranging gives . Therefore, on each generalized eigenspace, is simply a scalar operator plus a nilpotent operator, which reduces Jordan Normal Form to understanding nilpotent operators.
Let be an operator on a finite-dimensional, complex vector space, with distinct eigenvalues . Then
The sum is direct because generalized eigenvectors corresponding to distinct eigenvalues are linearly independent. By the Generalized Eigenbasis Theorem, has a basis consisting of generalized eigenvectors, and every such vector lies in one of the . Thus these spaces span .
Thus every operator on a finite-dimensional complex vector space decomposes into invariant subspaces on which it has the form for some nilpotent . Again, as a parallel to diagonalization, if is diagonalizable, then on every generalized eigenspace.
We now study the orbits of vectors under a nilpotent operator. Since repeated application of a nilpotent operator must eventually send every vector to , the orbit of a vector has the form
for some .
Let be nilpotent. If but , then
is called a Jordan chain of length . We call a generator of this Jordan chain.
Thus a Jordan chain is simply the nonzero part of the orbit of some vector under repeated application of . Notice that the final vector lies in , since .
When on , this means that
and hence is an ordinary eigenvector of corresponding to . Thus a Jordan chain begins with a generalized eigenvector and, under repeated application of , eventually terminates at an ordinary eigenvector.
Every Jordan chain is linearly independent.
Suppose
and let be the smallest index such that . Applying gives
since every later term contains a power of at least and therefore vanishes. This contradicts . Thus every coefficient is zero.
With respect to the ordered basis of its span, the matrix of is
One Jordan chain need not span the entire generalized eigenspace. A nilpotent operator can have several independent orbits which cannot be joined into one longer orbit. Jordan Normal Form amounts to decomposing the generalized eigenspace into the spans of these independent chains.
Every nilpotent operator on a finite-dimensional vector space admits a basis that is a union of Jordan chains.
We induct on . Let be the nilpotency index of , choose such that , and let . Then is -invariant and has a Jordan-chain basis.
If , we are done. Otherwise, LADR constructs an -invariant subspace such that . Since , the inductive hypothesis gives a Jordan basis for . Combining this with the chain spanning gives a Jordan basis for .
Thus a nilpotent operator generally decomposes into several Jordan chains,
where each is the span of one orbit
Each chain contributes exactly one vector to , namely its terminal vector . These terminal vectors are linearly independent because they belong to a Jordan basis, and together they span .
Let be nilpotent. In any Jordan basis for , the terminal vectors of the Jordan chains form a basis of . Consequently,
When , we have
Thus the number of Jordan chains inside is exactly .
The geometric multiplicity of an eigenvalue is equal to the number of Jordan chains in , and hence to the number of Jordan blocks corresponding to .
This gives a useful picture of a generalized eigenspace. If , then the generalized eigenspace decomposes into exactly independent Jordan chains. Each chain terminates at one of the independent eigenvector directions in , although the chains themselves may have different lengths.
We can now return to an arbitrary operator. If on some generalized eigenspace and a Jordan chain for has length , then with respect to that chain, has a particularly simple matrix.
The Jordan block corresponding to is
Thus one Jordan chain gives one Jordan block. If a generalized eigenspace contains several Jordan chains, then the restriction of to that generalized eigenspace contains several Jordan blocks, all having the same eigenvalue on their diagonals.
Let be an operator on a finite-dimensional, complex vector space. Then there exists a basis of such that
The Jordan blocks are unique up to their ordering. Note that the values appearing here need not be distinct.
By the Generalized Eigenspace Decomposition,
On each , the operator is nilpotent, so it admits a basis consisting of Jordan chains. Each such chain gives one Jordan block . Putting all of these chain bases together gives a basis of and produces the desired block diagonal matrix.
For a matrix , apply the theorem to the corresponding operator . If is a Jordan basis and is the matrix whose columns are the vectors of written in standard coordinates, then
so every complex matrix is similar to a matrix in Jordan Normal Form.
We now relate Jordan form back to the characteristic polynomial. A single Jordan block is triangular with appearing times on the diagonal, so
Now fix one eigenvalue . Suppose the Jordan chains inside have lengths . Then
where each is the span of one Jordan chain, and with respect to the resulting basis,
Thus the phrase “the Jordan blocks corresponding to ” simply means the Jordan blocks arising from the different Jordan chains inside the single generalized eigenspace . There is one generalized eigenspace corresponding to , but it may contain several Jordan chains, and therefore several Jordan blocks.
Since the chains together form a basis of ,
Each block contributes to the characteristic polynomial, so the total contribution of the entire generalized eigenspace is
Let be the distinct eigenvalues of . Then
In particular, the algebraic multiplicity of is
Thus there are two different numbers associated with each eigenvalue :
while
In terms of Jordan form, the algebraic multiplicity is the sum of the sizes of all Jordan blocks corresponding to , whereas the geometric multiplicity is the number of those blocks.
For example, suppose the -part of Jordan form is
Then , so the algebraic multiplicity of is . There are three Jordan chains, so , and hence the geometric multiplicity is .
The characteristic polynomial therefore tells us how much total dimension belongs to each eigenvalue, but it does not tell us how that dimension is divided into Jordan chains. For example, if has algebraic multiplicity , the -part could be
or another partition of . All of these have the same factor in the characteristic polynomial.
The powers of distinguish these possibilities.
Fix an eigenvalue , and let
with . Then is the number of Jordan blocks corresponding to having size at least .
Indeed, consider a single Jordan block of size . The operator kills exactly the last vectors of its Jordan chain, so that block contributes dimensions to . When we pass from to , this contribution increases by exactly when . Summing over all blocks gives the result.
Thus the sequence
completely determines the Jordan block sizes for . In particular, the first term gives the number of blocks, because
while the stabilized value gives their total size,
This explains how the three main pieces of information fit together: the characteristic polynomial determines the total dimension assigned to each eigenvalue, the eigenspace determines how many independent Jordan chains there are for that eigenvalue, and the full kernel chain determines the lengths of those chains.
Both LADR and Artin prove Jordan Normal Form by induction, but LADR separates the two main ideas much more cleanly. LADR first proves the generalized eigenspace decomposition, reducing the general case to nilpotent operators, and then proves the nilpotent case by choosing one longest Jordan chain and finding an invariant complement. Artin instead carries out much of the generalized eigenspace decomposition inside the Jordan-form proof itself: he shifts by an eigenvalue, studies stabilized kernel and image chains, restricts to the image, lifts Jordan generators through preimages, and finally adds kernel vectors. The arguments are fundamentally doing the same thing, but I find Artin's proof more mechanically involved because the decomposition and Jordan-chain construction occur simultaneously rather than as two separate structural steps.