From Linear Algebra to OpenCV: Understanding Image Transformations
Every time we open our phone camera, resize an image, or ask GPT something, linear algebra just ran. We didn't call it. We didn't realize it. But we definitely used it unknowingly.
In this article, we will explore linear algebra concepts with real-world examples, especially focusing on how they are used in computer vision and OpenCV.
Why do we need to learn linear algebra?
Let's take an image and resize it, rotate it, zoom in, zoom out, and finally shear it.
By completing this article, you will be able to understand exactly how these transformations are done using simple linear algebra examples.
In high school, we just read definitions and solved for x and y. Even when we encountered application-based questions, we mostly just plugged numbers into formulas without seeing what was actually happening.
Later, when I really wanted to visualize what these linear algebra transformations looked like, I started exploring. That's when I found 3Blue1Brown's series. It started with vectors, followed by matrix multiplication, and built up from there.
I have compiled all these learnings into a single article. We will learn the core concepts, and I will also add some code for doing the basic image operations you saw above. I hope this helps anyone who is just starting to explore linear algebra.
The Building Blocks of Linear Algebra
From Vectors to Linear Combinations, followed by Linear Transformations—these are the fundamental building blocks of linear algebra.
What Is a Vector?
In simple terms, a vector is an arrow pointing in a specific direction with a specific length. It's a way to represent movement from one point to another.
In linear algebra, a vector is rooted at the origin.
Each vector value is calculated with respect to movement parallel to the axes.
For example, $\begin{bmatrix} 3 \\ 2 \end{bmatrix}$: we move 3 units along the x-axis, then 2 units along the y-axis.
Every vector is associated with a pair of numbers.
Every pair of numbers gives exactly one vector.
For 3D space, we add the z-axis.
We mainly perform two operations on vectors:
- Addition
- Scaling
Vector Addition
As the name itself suggests, vector addition is nothing but adding vectors.
Example:
We take 2 vectors $\vec{a} = \begin{bmatrix} 1 \\ 2 \end{bmatrix}$ and $\vec{b} = \begin{bmatrix} 3 \\ 2 \end{bmatrix}$ and add them.
Then final vector
$\vec{c} = \begin{bmatrix} 1+3 \\ 2+2 \end{bmatrix} = \begin{bmatrix} 4 \\ 4 \end{bmatrix}$.
Vector Scaling
We multiply a vector with a scalar value. This can be used to make the vector longer, shorter or even reverse the direction.
Example:
We take a vector $\vec{a} = \begin{bmatrix} 1 \\ 2 \end{bmatrix}$ and multiply it by a scalar value 2.
Then final vector
$\vec{c} = \begin{bmatrix} 1 \cdot 2 \\ 2 \cdot 2 \end{bmatrix} = \begin{bmatrix} 2 \\ 4 \end{bmatrix}$.
Note: The number used for multiplying the vector is called a scalar.
This multiplying of a number with a vector is called scaling.
Linear Combinations
Even though we know how vectors work, there is another way to think about them.
Think of each coordinate as a scalar that stretches a basis vector.
For example, in 2D, we have $\hat{i}$ for the x-axis and $\hat{j}$ for the y-axis, both unit vectors.
For $\vec{v} = \begin{bmatrix} 1 \\ 2 \end{bmatrix}$
we scale $\hat{i}$ by 1 and $\hat{j}$ by 2, then add the results to get $\vec{v}$.
It is pronounced as "i-hat" and "j-hat".
They add together to form the vector $1\hat{i} + 2\hat{j}$.
Both $\hat{i}$ and $\hat{j}$ are the "basis vectors" of the xy coordinate system.
Any time we scale vectors and add them, it is called a linear combination of $\vec{v}$ and $\vec{w}$.
Why are scaling and adding vectors called a linear combination?
If we fix the direction of a single vector and vary its scalar multiplier, it traces out a 1-dimensional line.
But if we take two independent vectors and freely change both of their scalars, their combinations can reach every single position in the 2D plane.
If the direction of the two vectors is exactly the same (or completely opposite), then no matter how we scale and add them, we end up trapped on a single line.
The set of all possible vectors we can reach from a given set of vectors by linear combination is called the span of those vectors.
This freedom to scale independently and add the results together is exactly why it is called a linear combination.
Vectors vs Points
Usually, to avoid visual clutter when dealing with many vectors, we just represent each vector by its tip (a point). Since every vector starts at the origin, a single point uniquely identifies a vector.
If we plot all possible linear combinations in a 2D plane, the collection of all these vector tips forms an infinite 2D plane (or grid).
When thinking about one or two vectors, it helps to picture them as arrows. But when thinking about an entire span of vectors, it's easier to picture their tips forming a continuous space of points.
If two vectors line up, their span is just a line.
If two vectors are not pointing in the same direction, their span will be the collection of all possible vectors we get by tuning all possible scalars, creating a full 2D space.
It can also be seen as chaining: adding vector tips to tails freely in two independent directions covers the entire flat sheet. This flat sheet is the span of those 2 independent vectors.
Basically, when vectors are "linearly dependent," it means one vector is just a scaled version of another (or lies in the same span). Adding it doesn't give you access to any new dimensions—you are trapped in the space you already had.
Linear Combination in 3D
For 3D, we scale 3 different vectors and add them together.
That is: $a\vec{v} + b\vec{w} + c\vec{u}$.
The span of this vector set is all possible linear combinations.
If all 3 vectors point along the exact same line (they are all dependent on each other), their combinations only give you a 1D line:
If 2 vectors are independent, but the 3rd vector lies perfectly flat in the plane they create (it is dependent), their span is limited to that single 2D sheet. The third vector is redundant because it doesn't unlock any new directions.
But if all 3 vectors are independent—meaning the third vector points in a brand new direction away from the plane formed by the first two—then we unlock the entire 3D space!
This situation—where a vector doesn't add any new direction and keeps us stuck in a lower dimension (like a line in 3D, or a flat plane in 3D)—is called linear dependence.
If each vector adds a brand new dimension that wasn't reachable before, we say the set of vectors is linearly independent.
Finally, the basis of a vector space is a set of linearly independent vectors that spans the full space.
Linear Transformations
A transformation is nothing but a function which transforms any input into an output.
With respect to linear algebra, we say a function which transforms an input vector into an output vector.
Why was it named "transformation" instead of "function"?
Transformation implies it has movement. So if we give an input vector and it gives an output vector, then we can say the vector is moving from the input state to the output state.
To understand it as a whole, we can say every possible input vector is moving to its corresponding output vector.
As we also discussed multiple vectors as points, we can say the movement of input points to output points.
Usually the effect of a transformation is cool because it changes the shape of the 2D grid, which is fascinating to watch.
But to understand why "linear" is used in linear transformation, we should understand these core rules to classify it as linear.
What Makes Linear Transformation Different From Normal Transformation?
To classify a transformation as strictly "linear", it must satisfy two mathematical properties:
- It preserves vector addition.
$$ T(\vec{u}+\vec{v}) = T(\vec{u}) + T(\vec{v}) $$
- It preserves scalar multiplication.
$$ T(c\vec{v}) = cT(\vec{v}) $$
These two properties are the formal definition of a linear transformation.
There is also an easier way to visualize this:
- Lines remain lines and parallel lines remain parallel.
- The Origin stays fixed at $(0,0)$.
If the origin moves, or if straight lines become curved, the transformation is not linear.
If either of these rules is broken, it is a non-linear transformation. For example, if grid lines become wavy, or if the grid shifts so the origin moves, it's no longer a linear transformation.
How can we describe it numerically?
The vector $\vec{v} = \begin{bmatrix} -1 \\ 2 \end{bmatrix}$ is formed by multiplying the unit vectors $\hat{i}$ and $\hat{j}$ by the scalar values -1 and 2.
We can write this as a matrix multiplying our scalar values. Since $\hat{i}$ is mathematically just $\begin{bmatrix} 1 \\ 0 \end{bmatrix}$ and $\hat{j}$ is $\begin{bmatrix} 0 \\ 1 \end{bmatrix}$, we simply use them as the columns of our matrix:
$$ \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix} \begin{bmatrix} -1 \\ 2 \end{bmatrix} = \begin{bmatrix} -1 \\ 2 \end{bmatrix} $$
The beautiful property of a linear transformation is that we only need to track where the basis vectors $\hat{i}$ and $\hat{j}$ land. Every other vector just follows them perfectly.
If a transformation moves $\hat{i}$ to $\begin{bmatrix} 1 \\ -2 \end{bmatrix}$ and $\hat{j}$ to $\begin{bmatrix} 3 \\ 0 \end{bmatrix}$, then our vector $\vec{v}$ will land at:
$$ \vec{v}_{\text{transformed}} = -1(\text{transformed } \hat{i}) + 2(\text{transformed } \hat{j}) $$
$$ \vec{v}_{\text{transformed}} = -1\begin{bmatrix} 1 \\ -2 \end{bmatrix} + 2\begin{bmatrix} 3 \\ 0 \end{bmatrix} = \begin{bmatrix} -1 \\ 2 \end{bmatrix} + \begin{bmatrix} 6 \\ 0 \end{bmatrix} = \begin{bmatrix} 5 \\ 2 \end{bmatrix} $$
We write this mathematically as a Matrix-Vector multiplication! A matrix is just a package containing where $\hat{i}$ and $\hat{j}$ land as its columns:
$$ \begin{bmatrix} 1 & 3 \\ -2 & 0 \end{bmatrix} \begin{bmatrix} -1 \\ 2 \end{bmatrix} = \begin{bmatrix} (1 \times -1) + (3 \times 2) \\ (-2 \times -1) + (0 \times 2) \end{bmatrix} = \begin{bmatrix} -1 + 6 \\ 2 + 0 \end{bmatrix} = \begin{bmatrix} 5 \\ 2 \end{bmatrix} $$
Linear transformations squeeze, stretch, rotate, or shear the space such that grid lines stay parallel and evenly spaced.
Matrix Multiplication
So far, we have seen how a matrix can transform a vector.
But what if we want to apply more than one transformation?
For example, we may want to first scale a vector and then rotate it.
We can represent each transformation using a matrix.
If $A$ represents the first transformation and $B$ represents the second transformation, then:
$$ \vec{v}' = A\vec{v} $$
and then:
$$ \vec{v}'' = B\vec{v}' $$
Substituting the first transformation:
$$ \vec{v}'' = B(A\vec{v}) $$
which can be written as:
$$ \vec{v}'' = (BA)\vec{v} $$
Here, $BA$ means $A$ is applied first, then $B$, so we read from right to left.
This means we can combine multiple transformations by multiplying their matrices together.
The order is important because matrix multiplication is generally not commutative:
$$ AB \neq BA $$
So applying transformation $A$ followed by transformation $B$ can give a different result from applying $B$ followed by $A$.
2D Transformation Matrices
Now that we understand how matrices represent transformations, we can look at some common 2D transformations.
Scaling
To scale a vector along the x and y directions, we use:
$$ \begin{bmatrix} s_x & 0 \\ 0 & s_y \end{bmatrix} $$
For a vector:
$$ \begin{bmatrix} x \\ y \end{bmatrix} $$
the transformation becomes:
$$ \begin{bmatrix} s_x & 0 \\ 0 & s_y \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} = \begin{bmatrix} s_x x \\ s_y y \end{bmatrix} $$
Rotation
To rotate a vector by an angle $\theta$, we use:
$$ \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} $$
For example, for a 90° rotation:
$$ \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix} $$
which transforms:
$$ \begin{bmatrix} x \\ y \end{bmatrix} \rightarrow \begin{bmatrix} -y \\ x \end{bmatrix} $$
Shearing
A shear transformation changes the position of points along one direction.
For example:
$$ \begin{bmatrix} 1 & k \\ 0 & 1 \end{bmatrix} $$
transforms:
$$ \begin{bmatrix} x \\ y \end{bmatrix} \rightarrow \begin{bmatrix} x+ky \\ y \end{bmatrix} $$
These are the same types of transformations we will later apply to images.
Applying These Concepts To An Image
Instead of digging deeper into matrix math, let's look at a practical example! We'll take a simple image of an apple and use OpenCV to apply these basic transformations.
In OpenCV, we can define our 2x3 transformation matrix M and pass it to a single magical function called cv2.warpAffine().
This function takes the mathematical concepts we just explored and applies them to every pixel in the image simultaneously.
Note: In OpenCV, the origin (0,0) is at the top-left corner, and the y-axis points down. This means scaling happens from the top-left. Because y points down, the textbook rotation matrix would turn the image clockwise on screen, so
getRotationMatrix2Dflips the signs: a positive angle rotates counter-clockwise.
Applying these transformations to an image
To get started, we first load our image using OpenCV and get its dimensions.
import cv2
import numpy as np
# Load our apple image
image = cv2.imread('apple.jpg')
rows, cols = image.shape[:2]
1. Translation (Shifting)
Note: Translation is technically an affine transformation (a linear transformation plus a shift). That's why OpenCV uses a 2×3 matrix: the left 2×2 is the linear part and the third column is the shift.
To move an image, we use a translation matrix. The matrix requires two parameters: tx (shift along the x-axis) and ty (shift along the y-axis). Here, we shift the apple right by 100 pixels and down by 50 pixels.
# Matrix: [1 0 tx]
# [0 1 ty]
M_translate = np.float32([
[1, 0, 100],
[0, 1, 50]
])
translated_apple = cv2.warpAffine(image, M_translate, (cols, rows))
2. Scaling (Zooming)
Scaling changes the size of the image. The matrix uses sx and sy to scale along the x and y axes. Let's zoom in on our apple by making it 1.5x larger. We also need to increase the output dimensions so the image isn't cropped! If we use a scale factor below 1 (e.g., 0.5), it will zoom out instead.
# Matrix: [sx 0 0]
# [0 sy 0]
M_scale = np.float32([
[1.5, 0, 0],
[0, 1.5, 0]
])
scaled_apple = cv2.warpAffine(image, M_scale, (int(cols * 1.5), int(rows * 1.5)))
3. Rotation
Rotating an image requires calculating sines and cosines for the matrix, which can be tedious. Luckily, OpenCV has a built-in helper cv2.getRotationMatrix2D() that generates the affine matrix for us based on the center point and angle. Let's rotate the apple by 45 degrees.
center = (cols // 2, rows // 2)
M_rotate = cv2.getRotationMatrix2D(center, angle=45, scale=1.0)
rotated_apple = cv2.warpAffine(image, M_rotate, (cols, rows))
4. Shearing
Shearing slides each row sideways in proportion to its height. Let's apply a shear along the x-axis.
# Matrix: [1 kx 0]
# [0 1 0]
M_shear = np.float32([
[1, 0.5, 0],
[0, 1, 0]
])
sheared_apple = cv2.warpAffine(image, M_shear, (int(cols + 0.5 * rows), rows))
By defining different transformation matrices, we can squash, stretch, rotate, and move our images however we want. Since it's all based on linear algebra, the computer can process millions of pixels in a fraction of a second using these simple matrix multiplications!
What's Next?
We've just scratched the surface of what OpenCV can do with linear algebra. Now that we've covered the foundational concepts of affine transformations—translating, scaling, rotating, and shearing our images—we've built the perfect stepping stone for more advanced topics.
In future posts, we'll dive deeper into perspective transformations, 3D projections, and how these exact same mathematical principles power everything from augmented reality filters to deep learning computer vision models.
This is just the beginning—stay tuned for more!