Back to lab

What does a linear transformation actually do?

Matrices are often introduced as rectangular arrays of numbers, which is perhaps one of the least helpful ways to understand them geometrically.

A matrix can instead be viewed as describing a transformation of space. In fact, when one graduates from a first course in linear algebra to an advanced (or abstract) linear algebra course, we learn about so-called linear transformations in the more abstract sense and on more abstract spaces (more on this later).

The important fact is that a linear transformation is completely determined by what it does to a basis. If

v=αb1+βb2,v = \alpha b_1 + \beta b_2,

then linearity forces

T(v)=αT(b1)+βT(b2).T(v) = \alpha T(b_1) + \beta T(b_2).

Linearity of a function, call it ff, refers to it satisfying:

f(x+y)=f(x)+f(y)f(αx)=αf(x), f(x+y) = f(x) + f(y) \\ f(\alpha x) = \alpha f(x),

where x,yx,y are vectors in the space, and α\alpha is in the ground field (usually R, C\mathbb R, \ \mathbb C or similar) .

Since all vector spaces have a basis, except for infinite dimensional spaces when you disallow the axiom of choice , which let’s you write elements of the space v∈Vv \in V as a linear combination of basis elements

v=∑i=1nαiei,(again, we ignore the infinite dimensional case) v = \sum_{i=1}^n \alpha_i e_i, \\ \color{pink} (\text{again, we ignore the infinite dimensional case})

a linear transformation can be described entirely by what it does to the basis elements.

The tool below lets you look at the same idea from three slightly different perspectives.

We are working with the familiar vector space R2\mathbb R^2, as it is easy to visualize, though the idea generalizes to arbitrary vector spaces.

The demo let’s you do a couple of things:

basis-relative grid

Linear map

(
)

T(e₁) = (1, 0) · T(e₂) = (1, 1)

det(A) = 1

Basis

Valid basis · det = 1

Placed vectors keep their coefficients relative to this basis, so changing the basis moves them in real time.

Transformation

Click anywhere to place a vector.

Selected vector

[v]₍B₎ = (3, 2)

v = 3b₁ + 2b₂

Standard coordinates:

v = (3, 2)

What follows is for the more curious.

What is a basis?

In the demo I let you define a different basis from e1=(1,0), e2=(0,1)e_1 = (1,0), \ e_2 = (0,1). Now, perhaps it is intuitive from the demo and preceding exposition what this is supposed to mean, perhaps not.

Definition: Basis of a vector space.

Let VV be a vector space over a field K\mathbb K (feel free to search up the rigorous definition). A basis of VV is a linearly independent (feel free to search up the definition here also, though I think this one is included in a first course) set of vectors {b1,b2,…,bn}⊂V\{b_1,b_2,\ldots,b_n\} \subset V such that

span⁡{b1,b2,…,bn}=V, \operatorname{span}\{b_1,b_2,\ldots,b_n\} = V,

meaning every vector v∈Vv \in V can be written as a linear combination

∑i=1nαibi, αi∈K for i=1,2,…,n. \sum_{i=1}^n \alpha_i b_i, \ \alpha_i \in \mathbb K \ \text{for} \ i = 1,2,\ldots,n.

Note: The basis is finite here, but this can generalize to infinite dimensions.

Now what does this mean? For any linear transformation between spaces T:V→WT : V \to W,
have

T(x+y)=T(x)+T(y),T(αx)=αT(x), α∈K. T(x+y) = T(x)+T(y), \\ T(\alpha x) = \alpha T(x), \ \alpha \in \mathbb K.

Since x=∑iαieix = \sum_i \alpha_i e_i, and TT has the property above, we have

T(x)=T(∑iαiei)=∑iαiT(ei), T(x) = T\left( \sum_i \alpha_i e_i \right) = \sum_i \alpha_i T(e_i),

so we can determine where xx goes by seeing where the basis elements go!

Why is this cool/important?

So far, changing basis may look like a somewhat unnecessary exercise in describing the same space in different ways.

It’s useful because some bases make problems easier.

Observe:

Suppose that a linear transformation T:V→VT:V\to V has a basis consisting of eigenvectors

e1,e2,…,en,e_1,e_2,\ldots,e_n,

so that

T(ei)=λiei.T(e_i)=\lambda_i e_i.

Now take an arbitrary vector

x=∑i=1nαiei.x=\sum_{i=1}^n \alpha_i e_i.

By linearity,

T(x)=∑i=1nαiT(ei)=∑i=1nαiλiei.T(x) = \sum_{i=1}^n \alpha_i T(e_i) = \sum_{i=1}^n \alpha_i \lambda_i e_i.

Notice what happened.

Instead of the transformation mixing all the coordinates together, each basis direction simply gets multiplied by a number.

In this basis, the transformation is represented by a diagonal matrix

[T]=(λ10⋯00λ2⋯0⋮⋮⋱⋮00⋯λn).[T] = \begin{pmatrix} \lambda_1 & 0 & \cdots & 0\\ 0 & \lambda_2 & \cdots & 0\\ \vdots & \vdots & \ddots & \vdots\\ 0 & 0 & \cdots & \lambda_n \end{pmatrix}.

Very cool!

An example in PDEs

This becomes especially useful in engineering because many differential equations involve linear differential operators.

For example, differentiation is linear:

ddx(f+g)=dfdx+dgdx,\frac{d}{dx}(f+g)=\frac{df}{dx}+\frac{dg}{dx},

and

ddx(αf)=αdfdx.\frac{d}{dx}(\alpha f)=\alpha \frac{df}{dx}.

Likewise, the second derivative operator

d2fdx2 \frac{d^2f}{dx^2}

is linear.

So instead of thinking about a matrix acting on vectors in R2\mathbb R^2, we can think about a differential operator acting on a vector space whose “vectors” are functions.

For arbitrary differentiable functions this becomes infinite dimensional. In PDEs one usually works with function spaces and orthogonal/orthonormal systems rather than treating everything as a finite-dimensional vector space.

Example: heat flowing through a rod

Suppose a thin rod of length LL has temperature

u(x,t)u(x,t)

at position xx and time tt.

A standard model for heat flow is the heat equation

∂u∂t=κ∂2u∂x2,\frac{\partial u}{\partial t} = \kappa \frac{\partial^2 u}{\partial x^2},

where κ>0\kappa>0 describes how quickly the material conducts heat.

Suppose also that both ends of the rod are kept at temperature zero:

u(0,t)=u(L,t)=0.u(0,t)=u(L,t)=0.

At first sight this is rather more intimidating than multiplying a vector by a 2×22\times2 matrix.

But consider the functions

sin⁡(nπxL),n=1,2,3,…\sin\left(\frac{n\pi x}{L}\right), \qquad n=1,2,3,\ldots

If we apply the second derivative operator to one of these functions, we get

d2dx2sin⁡(nπxL)=−(nπL)2sin⁡(nπxL).\frac{d^2}{dx^2} \sin\left(\frac{n\pi x}{L}\right) = - \left(\frac{n\pi}{L}\right)^2 \sin\left(\frac{n\pi x}{L}\right).

That should look familiar.

Each sine function is an eigenfunction of the second derivative operator.

Instead of

Av=λv,Av=\lambda v,

we now have

D2f=λf.D^2f=\lambda f.

Now suppose the initial temperature distribution can be written as

u(x,0)=∑n=1∞cnsin⁡(nπxL).u(x,0) = \sum_{n=1}^{\infty} c_n \sin\left(\frac{n\pi x}{L}\right).

This is a Fourier sine expansion.

Because the heat equation is linear, we can study every one of these basis functions independently.

The solution becomes

u(x,t)=∑n=1∞cne−κ(nπ/L)2tsin⁡(nπxL).u(x,t) = \sum_{n=1}^{\infty} c_n e^{-\kappa(n\pi/L)^2t} \sin\left(\frac{n\pi x}{L}\right).

So an apparently complicated temperature distribution has been decomposed into simple modes.

Each mode evolves independently:

cn⟼cne−κ(nπ/L)2t.c_n \quad\longmapsto\quad c_n e^{-\kappa(n\pi/L)^2t}.

Higher-frequency modes decay faster, which is why sharp variations in temperature smooth out over time.