Showing posts with label group theory. Show all posts
Showing posts with label group theory. Show all posts

Walkthrough of Galois theory

Symmetries of polynomials; the Galois group

You have often seen that there is a certain fundamental symmetry between the roots of a polynomial. For example, there is an obvious symmetry between $i$ and $-i$, the roots of $x^2+1=0$. But even without imaginary numbers, there is a certain symmetry between $2+\sqrt{3}$ and $2-\sqrt{3}$, the roots of $x^2-4x+1=0$. 

What is this symmetry precisely? The idea is this: $i$ and $-i$ can only "relate" to the reals through squaring -- thus any equation in them with real coefficients remains true up to permutation of roots. In equations like $\alpha^2+1=0$, $\alpha+\beta=0$, etc. one may permute the roots $\alpha,\beta$ and maintain the equations. The only way you can break this symmetry is with an equation whose coefficients are themselves complex. 

It's not necessary that every permutation is a symmetry of the polynomial. For example, consider the equation $x^4-10x^2+1=0$, whose roots are $\sqrt2+\sqrt3$, $\sqrt2-\sqrt3$, $-\sqrt2+\sqrt3$, $-\sqrt2-\sqrt3$. In fact, the only permutations that are symmetries of these roots are: $e$ (which conjugates nothing), $(12)(34)$ (which conjugates $\sqrt3$), $(13)(24)$ (which conjugates $\sqrt2$), $(14)(23)$ (which conjugates both $\sqrt2$ and $\sqrt3$) -- these permutations are the Klein 4-group $\mathbb{Z}_2\times \mathbb{Z}_2$. 

As a more extreme example, consider $x^2-3x+2=0$, whose roots are $1$ and $2$. There is no non-trivial permutation of these roots that preserves e.g. the equation $\alpha-1=0$.

This group of symmetries is known as the Galois group.

How might we formalize this? Well, you may notice that none of what we described really depends on the polynomial -- just its roots. Fundamentally, we are "extending" the rational numbers with these roots -- e.g. to $\mathbb{Q}[i]$ or $\mathbb{Q}[\sqrt2]$ or whatever -- then, instead of saying "a permutation of these roots that preserves all $\mathbb{Q}$-algebraic equations in them", we can say "a field automorphism that fixes $\mathbb{Q}$" (because under such an automorphism, all arithmetical operations and $\mathbb{Q}$ coefficients remain the same); this is nice, because it is more natural to think of complex conjugation as a symmetry of the complex plane as a whole rather than just .

Thus our formal definition: the Galois group of some field extension $L/K$ is its automorphism group (the group of automorphisms of $L$ that fix $K$). 

Why is this actually useful or relevant? It's a natural construction so of course it's useful, go perforate your head IDC.

Why field theory? Constructible numbers

Geometrical constructability is defined as follows: given some initial geometric figure (collection of points with known distances), what geometric figure can be constructed with only a straightedge and compass? 

More formally, we have the following axioms -- to define constructible points, shapes and numbers. Where $G$ is is a collection of points in the plane:

  1. The points in $G$ are constructible points.
  2. A line through two constructible points is a constructible shape.
  3. A circle with its centre a constructible point and its radius a constructible line is a constructible shape.
  4. The intersections of two constructible shapes are constructible points. 
  5. The co-ordinate of a constructible point is a constructible number.

Pedantry: 

In fact, these are the axioms for constructability with a collapsible compass -- a compass whose legs collapse once it is taken off the paper -- so you cannot use it to mark an identical distance elsewhere. For constructability with a non-collapsible compass, Axiom 3 would instead read: A circle with its centre a constructible point and its radius equal to the length of a constructible line is a constructible shape. 

Exercise: show that these definitions are equivalent -- that under the axioms for a collapsible compass, you can move a line to an arbitrary point. Solution.

In particular, this means that constructability only depends on the initial set of distances, not the precise points themselves, because any arrangement of these lengths can be shifted around -- we can formulate the question as "what numbers are constructible from some initial set of numbers?".

Exercises: Given numbers $a$ and $b$, show that the following numbers are constructible from them:

  1. $a+b$
  2. $|a-b|$
  3. $ab$
  4. $a/b$
  5. $\sqrt{a}$

Numbers expressible through only these operations: $+,-,\times,/,\sqrt{\cdot}$ are called algebraically constructible from some base set of numbers. 

The above exercise demonstrates that all algebraically constructible numbers are geometrically constructible.

Conversely, as all geometrically constructible lengths can be constructed using these operations (you know, Pythagoras theorem and stuff), all geometrically constructible numbers are algebraically constructible.

Thus algebraical and geometrical constructability are equivalent. 

When no base figure is specified, the tacit assumption is that the base figure is a line segment of length 1, i.e. the base set is "1". We simply call the numbers constructible from this to be the constructible numbers. 

Other forms of constructability besides straightedge-and-compass are known -- such as conic constructability, solid constructability, neusis constructability, origami constructability. We want a theory that handles not only straightedge-and-compass (i.e. square roots), but also these more general forms.

This immediately demonstrates why:

  • It is impossible to double the cube.
  • It is impossible to trisect an arbitrary angle (because that is equivalent to constructing a cube root).
  • It is impossible to square the circle (because $\pi$ cannot be constructed with radicals). 
These are not proofs per se ... it still remains to be shown that you literally can't construct a cube root with some finite nesting of square roots, and so on. 

A set closed under $+,-,\times,/$ is a field. So constructible numbers are some type of fields. What kind of fields, and what do we do with them? Yeah, yeah, something.

Computing the Galois group

So how do we actually figure out the Galois group of a thing?

First of all, we said that it's really the roots that matter, not the polynomial. Does that mean that we can adjoin any elements to a field and calculate the Galois group of that extension?

Not really. There's nothing interesting to be said about the automorphism group of $\mathbb{Q}[\sqrt[4]{2}]$, for example. $x^4-2=0$ has imaginary roots and stuff, and we want to be able to permute them. The automorphism group of this field -- with only one root adjoined -- says nothing of the properties of $x^4-2=0$. 

The fundamental theorem of Galois theory -- which we haven't yet stated -- will not apply to such a shitty field extension.

No -- the field extensions we are interested in are those generated by adjoining all the roots of a polynomial. This is a "normal extension", "Galois extension", "splitting field", whatever (yeah, yeah, a Galois extension must be normal and separable but do I look like the kind of guy who'd bother with things that aren't separable? -- what even is separable, your FACE is separable, I'll split it in two).

We can "intuit" out the Galois group of things.

The splitting field of $x_4-1$ is $\mathbb{Q}[i]$; then the image of $i$ determines the Galois group, which is thus $S_2$.

The Galois group of $x^4+1$ is $K_4$. Prove it.

The splitting field of $x^4-2$ is $\mathbb{Q}[\sqrt[4]{2},i]$; then the images of $\sqrt[4]{2}$ (for which there are 4 options) and $i$ (for which there are 2 options) are sufficient to determine the Galois group, which is thus $D_4$.

The Galois group of $\mathbb{Q}[\sqrt{5},\sqrt{7}]$ is $K_4$. Prove it.

The Galois group of $x^3-1$ is $S_2$. Prove it.

The Galois group of $x^3+1$ is $S_2$. Prove it.

The Galois group of $x^3-2$ is $S_3$. Prove it.

The Galois group of $x^5-1$ is $C_4$. Prove it.

The Galois group of $x^5+1$ is $C_4$. Prove it.

The Galois group of $x^5-2$ is ... okay, just see this Math.SE question. It's some semidirect product, and you should be able to see why by now.

If you worked through the examples above, you should have a good intuitive grasp of the Fundamental theorem of Galois theory by now. We've been constructing the Galois group of a field extension by determining some number of "basis elements" whose images suffice to determine the automorphism; extensions by these basis elements (intermediate field extensions) correspond to specific subgroups (intermediate subgroups of the subgroup lattice).

Solvable groups

add.

Lie group topology

I'll assume you have a basic understanding of general topology -- if not, consult the topology articles here. Most of the abstract stuff and "weird" cases are not really important, because it is easy to see that Lie groups are manifolds.

We need to be careful while studying the topology of Lie groups, because we already have an intuitive picture of a Lie group, and we need to be careful to prove all the things we just "believe" to be true.

The main point of the topology of a Lie group is that the group elements define the "flows" on the manifold. What this means is that left-multiplication is a homeomorphism, and it's not absurd to say that inversion is a homeomorphism, because it represents a "reflection" of the manifold. That these conditions make sense is confirmed by looking at the proofs of the following "obvious" facts.

(1) In a connected group, a neighbourhood of the identity generates the entire group, i.e. $H\le G\land H\in N(1)\implies H=G$ for connected $G$.

Let's think about why this is true. Why does $H$ need to be a neighbourhood -- why must it contain an open set containing the identity? Suppose instead we just knew it contained a set $Q$ that looked like this:


Well, $H$ still contains the orange point, but we cannot say it contains the purple point, because it's perfectly happy not containing it -- it's not like we have some vertical element in the Lie group that if you multiplied to some point in $Q$, you'd get the purple point. But instead if $Q$ was an open neighbourhood of the identity:

Then the purple point has to be in $H$, because $Q$ contains flows in "all directions" on the group. To actually prove that every point will be contained in $H$ -- well, we know that the point is (will eventually be) that $H$ is the connected component of $G$ (and since $G$ is connected, $H=G$) -- let's just show that $H$ is both open and closed, i.e. nothing in $H$ touches its exterior, and nothing in its exterior touches $H$. Here's the proof:
  • Nothing in $H$ touches anything -- Suppose $\exists x\in H, x\in\mathrm{cl}(H')$. Then $xQ$ contains a point in $H'$.
  • Nothing outside $H$ touches it -- Suppose $\exists x\in H', x\in\mathrm{cl}(H)$. Then $xQ$ contains a point in $H$, so $x$ must be in $H$.
We're really just formalising the notion of "translating $Q$ to its edges to extend $H$ further and further". The key fact we've used here is, of course, that left-multiplication is a homeomorphism, so $xQ$ is still an open set.

(2) The connected component of the identity is a subgroup.

The idea is that taking two elements $g,h$ of the connected component, their product should remain in the connected component. Once again, this follows from the continuity of left-multiplication -- considering the action of left-multiplication by $g$ on the connected component, its continuity implies that the image must remain connected.

(3) If a subgroup contains a neighbourhood of the identity, it contains the connected component of the identity.

Corollary to (1) and (2).

(4) The connected component of the identity is a normal subgroup.

Conjugation is a continuous map.

(5) Open subgroups are closed.

Corollary to (3). Alternate proof: the complement is the union of some cosets, which are open sets too. A weaker theorem can be made of closed sets -- closed subgroups with finite index are open.

What this means: any open subgroup is a union of connected components.

(6) Intuition for compact subgroups

How can a Lie group possibly "close in on itself"? Surely we keep "extending" an open neighbourhood $W$ of the identity by observing that $xW$ must be in the subgroup? The idea is that these translations of $W$ form an open cover of the group, if it has a finite subcover, then it makes sense for the group to close in on itself. By playing around with different open neighbourhoods $W$ and taking some suitable unions, one can see that this is equivalent to the condition that every open cover has a finite subcover, i.e. the group is compact.


(7) A compact, connected Abelian Lie group is a torus.

This is a generalisation of "a finite Abelian group is the direct product of cyclic groups".

The idea behind the proof is that in the Abelian case, the exponential map is a homomorphism from the Lie algebra to the Lie group, but the Lie algebra cannot detect compactness in the Lie group -- the kernel of the exponential map can. We know from our study of the exponential map that it has a discrete kernel, and in the Abelian case is surjective -- thus the Lie group is homeomorphic to $\mathbb{R}^n/\mathbb{Z}^n$, which is an $n$-torus.

(8) A connected Abelian Lie group is a cylinder (direct product of a torus and an affine space)

Analogous to above, except $\mathbb{R}^m/\mathbb{Z}^n$ where $m\ge n$.

Intuition, analogies and abstraction

$$-1=\sqrt{-1}\sqrt{-1}=\sqrt{(-1)(-1)}=\sqrt{1}=1$$
I bet you've seen the fake "proof" above that minus one and one are equal. And the standard explanation as to why it's wrong is that the statement $\sqrt{ab}=\sqrt{a}\sqrt{b}$ only applies when $\sqrt{a}$ and $\sqrt{b}$ are real, or something like that (maybe only one of them needs to be real -- something like that -- who cares?).

But if you're like me, that isn't a very satisfactory proof. Why does the identity not hold for complex numbers? For that matter, why does it hold for real numbers? Well, that is a good question, and one way of answering it would be to try and prove the identity for real numbers, and see what properties of the real numbers (or of the real square root, in particular) you use. And if this article were being filed under "MAR1104: Introduction to formal mathematics", that's how I might explain things -- but that doesn't give us too much insight -- not about square roots and complex numbers, anyway.

Let's think about what $\sqrt{ab}=\sqrt{a}\sqrt{b}$ means.

What does the square root of a real number mean, anyway? It's some property related to multiplying a real number by itself. What does multiplication mean? What does a real number mean? The picture I have in my head of the real numbers is of a line. But what exactly is this line? -- the real numbers are just a set. Why did you put them on this line in this specific way? In doing so, you gave the real numbers a structure, a specific type of structure called an "order", defined by the operation $<$.

But there are other ways to think about/structure the real numbers. One way is to think of real numbers as (one-dimensional) scalings. You can scale things like mass, and volume, using real numbers, representing the scalings as real numbers. Scaling a mass by 2 is equivalent to multiplication by 2. So this gives the real numbers a multiplicative structure, defined by the operation $\times$ (or whatever notation -- or lack thereof -- you prefer). And the "real line" then just represents the image of "1" under all scalings.

So the way to think about square roots is to think of numbers as linear transformations called scalings, and think about the scaling that when done twice, gives you the number you're taking the square root of. So what's $\sqrt{-1}$? What's $-1$? $-1$, multiplicative, is a reflection. What's its square root? Try to think of a (linear!) transformation that when done twice gives you a reflection. It can't be done in one dimension. And can you think of another such transformation? Can you prove these are the only two? Are you sure -- what about if you add a dimension?

So the natural way to think about square roots of numbers that may or may not be complex, is with so-called "Argand diagrams", on the complex plane, the image of "1" under all complex numbers multiplicative.

Click "edit graph" to play with a and b!

To simplify things, consider only unit complex numbers (this is okay, because all complex numbers can be written as a real multiple of a unit complex number and a real number). The product of complex numbers $a$ and $b$ involves rotating by $a$, then rotating by $b$. The square roots of $a$ and $b$ involve going halfway around the circle as $a$ and $b$, and the square root of $ab$ goes halfway around the circle as $ab$.

So it seems like the identity should hold, doesn't it? $\sqrt{ab}$ goes half as much as $a$ and $b$ put together -- this seems to be exactly what $\sqrt{a}\sqrt{b}$ does -- go around half as much as $a$, then half as much as $b$. Isn't $\frac{\theta+\phi}2=\frac{\theta}2+\frac{\phi}2$?

The problem is that $\sqrt{ab}$ doesn't really go $\frac{\theta+\phi}2$ around the circle, if $\theta+\phi$ is greater than $2\pi$. You can see this in the diagram courtesy of Desmos above -- $ab$ has gone a full circle, and its square root is defined to halve the argument of $ab$, but the argument isn't $\arg (ab)=\arg (a) + \arg (b)$, rather:

$$\arg (ab) \equiv \arg (a) + \arg (b) \pmod{2\pi}$$
But halving is not an operation that the $\bmod$ equivalence relation respects -- not in general, anyway. It is not true that

$$\arg (ab)/2 \equiv (\arg (a) + \arg (b))/2 \pmod{2\pi}$$
Instead:

$$\arg (ab)/2 \equiv (\arg (a) + \arg (b))/2 \pmod{\pi}$$
Let's recall from basic number theory -- on integers, the general result regarding multiplication on mods. If $a\equiv b\pmod{m}$, then $na\equiv nb \pmod{nm}$, certainly, and also $na\equiv nb \pmod{m}$ iff $n$ is an integer*. But $1/2$ isn't an integer, which is why only the former result is relevant.

This is also why $(ab)^2=a^2b^2$ does hold for complex numbers.

*when $n$ isn't an integer, we need $na$, $nb$ to be integers for the statement to even be well-defined in standard number theory, and then you have a result for division on mods involving $\gcd(d,m)$, etc. This isn't a concern for us here because we're dealing with divisibility over the reals -- if you want to be formal, a real number is divisible by another real number if the former can be written as an integer multiple of the latter.

So there you have it -- I just demonstrated a very fundamental analogy between two seemingly incredibly unrelated ideas: complex numbers modular arithmetic -- square roots of complex numbers don't multiply naturally, because mod doesn't respect division. It's almost as if somehow, somewhere, somehow magically, exactly the same kind of math was used to derive results, to prove things, about these unrelated objects.

As if they're just two instances of the same thing.

I wonder what that thing could be.



Let's talk about something completely unrelated (no, genuinely -- completely unrelated -- I won't tell you this is an instance of the "same thing" too). Let's talk about logical operators, specifically: do $\forall$ and $\exists$ commute? I.e. is $\forall t, \exists s, P(s,t)$ equivalent to $\exists s, \forall t, P(s,t)$?

You just need to read the statements aloud to realise they don't. To use a classical example, "all men have wives" and "there is a woman who is the wife of all men" are two very different statements (okay, in this case both statements are false, so they're equivalent in that sense, so you get my point).

But let's think more deeply about why they don't commute. What do $\forall t, \exists s, P(s,t)$ and $\exists s, \forall t, P(s,t)$ mean, anyway? $\forall$ and $\exists$ are just infinite $\land$ and $\lor$ statements , i.e. $\forall t$ is just an $\land$ statement ranging over all possible values that $t$ can take and $\exists s$ is just an $\lor$ statement ranging over all possible values $s$ can take.

So $\forall t, \exists s, P_{st}$ just means (letting $s$ and $t$ be natural numbers for simplicity, but they don't have to):

$$({P_{11}} \lor {P_{21}} \lor ...) \land ({P_{12}} \lor {P_{22}} \lor ...) \land ...$$
And $\exists s, \forall t, P(s,t)$ means:

$$({P_{11}} \land {P_{12}} \land ...) \lor ({P_{21}} \land {P_{22}} \land ...) \lor ...$$
This is a bit complicated, so let's instead look at the simpler case where you have only 2 by 2 statements -- i.e. just construct the analogy between $\forall,\exists$ and actual $\land,\lor$ statements.

So the question is if:

$$({P_{11}} \lor {P_{21}}) \land ({P_{12}} \lor {P_{22}}) \Leftrightarrow ({P_{11}} \land {P_{12}}) \lor ({P_{21}} \lor {P_{22}})$$
This is interesting. Maybe you see where this is going. Let me just do a notation change -- I'll use "$\times$" for $\land$, "$+$" for $\lor$, "$=$" for $\Leftrightarrow$" and some new letters for the propositions. Under this new notation, where $\times$ is invisible as always, we're asking if:

$$(a + b)(c + d) = ac + bd$$

Aha! This is Freshman's dream, isn't it? And we know it's not true -- it's a dream, after all, don't be delusional -- and we know why it's not true too.

But wait -- we aren't talking about elementary algebra here. I just gave you some silly notation and made it look like Freshman's dream. But here's the thing: the proof (or algebraic proof -- a counter-example is also a proof, but that isn't so interesting... not here, anyway) that these propositions aren't equivalent is exactly the same as in algebra. We expand out the brackets (because we know that $\land$ distributes over $\lor$ -- we also know that $\lor$ distributes over $\land$, incidentally, something that is not true in standard algebra) and point out that there are extra terms, and point out that these extra terms change the value of the expression (they aren't zero).

So there's some kind of relationship between the boolean algebra and an elementary algebra. A lot of proofs that can be done in one of these algebras can be written almost identically in the other. Not all these proofs, mind you -- then the algebras would just be isomorphic to each other -- but some of them can. Maybe a lot of important ones can.

An abstraction that produces such proofs simultaneously for both elementary algebra and boolean algebra may be more complicated than you think -- there's no real sense in which a statement is "always zero" in boolean algebra. Take for instance, distributivity of $\lor$ over $\land$ -- $a+bc=(a+b)(a+c)$. This is not true in elementary algebra, because the extra term $ab+ac$ is not always equal to zero ($a^2\ne a$ is not really an example, because $a^2=a$ for $a\in\{0,1\}$ -- but $a(b+c)=0$ is not true for all $a,b,c\in\{0,1\}$). It's just that it leaves the value of the existing terms unchanged in this specific instance.



I've just illustrated two examples here -- the first one is a type of group, by the way, but you've probably seen dozens of other such "connections between different areas of mathematics" yourself. I've made these sorts of analogies fundamental to a lot of the articles I've written here (I think). You might've just thought of them as interesting insights, but in reality, abstract mathematics/abstract algebra -- or really just mathematics in general -- is all about these analogies.

In a sense, mathematics is largely about abstraction. I mean, that's not what mathematics fundamentally is -- fundamentally, math is just logic -- but it's how mathematics largely functions. Whenever one talks of axioms, you could think of them as fundamental defining ideas of mathematical objects, and you can also think of them as "interfaces" between mathematics and reality (see my introduction to linear transformations). There are a massive number of different physical phenomena that we can study, and rather than prove everything from scratch for each one of them, it is much better -- and more insightful in terms of understanding the connections between things -- to show that they satisfy a certain set of axioms that apply to a whole range of things, and then deduce that all the logical consequences of these axioms -- all theorems -- are satisfied by the objects.

If we can do that with physical phenomena, we can sure as well do it with mathematical phenomena too -- instead of proving something from scratch for every new mathematical object, we prove that it is a group, or a ring, or a field, or a module, or an algebra, or a topology, or a geometry of some sort, by verifying it matches the axioms -- and then use all the abstract knowledge we have about these things and deduce they must necessarily apply to our new object, because they are logical consequences of our axioms.

Abstract mathematics is, in this sense, all about generalising things by finding the "smallest set of axioms" the thing requires.

(Well, not really -- the most general statement is "true", and everything else is just a logical deduction from this statement. So in that sense mathematics is all about finding special cases. But in order to know what to take a special case of, and what special case that "what" is of "true", you need to generalise.)

List some weird analogies you've seen before in math. Something about divisibility sound familiar?

Quaternion introduction: Part I

I generally really like the content produced at 3blue1brown, but their recent video on quaternions was just downright terrible. It entirely lacked Grant Sanderson's signature "discover it for yourself" approach, i.e. motivating the idea from the ground-up, and focused too much on an arbitrary formalism (stereographic projections aren't necessary for visualising anything).

The right way to motivate quaternions is to start by thinking about generalising complex numbers to higher dimensions. Complex numbers are a remarkable and elegant idea -- if you don't understand why I'm saying this, you could either get off the grid and spend the rest of your life as a circus monkey, or you could read my posts "Null and row spaces, transpose and the dot product" and "Making sense of Euler's formula".

The key idea behind complex numbers is that they are an alternate, simple representation of a specific set of linear transformations, namely: two-dimensional spirals (scaling and rotations). Note, similarly, that the real numbers can also be considered an alternate representation of e.g. scaling in one dimension.

The natural way to generalise complex numbers to more than two dimensions may seem to be to have an imaginary unit for each possible rotation (or more precisely, each "basis rotation"). In three dimensions, the basis has three planes of rotation, and could be e.g. rotations in the xy-plane, rotations in the yz-plane and rotations in the zx plane (you may have heard these as rotations "around" the z, x and y axes respectively, referring to the axes that remain invariant during the rotation -- however, as it turns out, in a greater number of dimensions $n$, the number of dimensions held invariant is $n-2$, which is only equal to 1 -- i.e. a single axis -- in 3 dimensions. e.g. in 4 dimensions, an $xy$-rotation would leave the $zw$ plane invariant.)

So let's try out this formalism, because it seems promising. We could write, e.g. i for the yz rotation, j for the zx rotation and k for the xy rotation. Try to work out some of the algebra here for yourself. What does $ij=?$ equal? What does $jk = ?$ What does $i^2=?$ equal?

As it turns out, none of these transformations result in anything very interesting. It would have certainly been elegant if you'd gotten nice results, like $ij=k$, or something, but you don't. One of the neat things about the complex number system is that not only do all complex numbers together, or all unit complex numbers together, form a group -- even $\{1,i,-1,-i\}$ forms a group under multiplication. But $\{1,-1,,i,j,k,-i,-j,-k\}$ do not form a group.

How would one solve this problem? Well, the reason $i^2$ doesn't equal minus 1 is that it only offers a reflection across the $x$-axis. The matrix representing $i^2$ is:

$${\left[ {\begin{array}{*{20}{c}}1&0&0\\0&0&{ - 1}\\0&1&0\end{array}} \right]^2} = \left[ {\begin{array}{*{20}{c}}1&0&0\\0&{ - 1}&0\\0&0&{ - 1}\end{array}} \right]$$
(If you can't come up with the matrix for $i$, you should review the linear algebra series -- or the circus monkey thing.) What if you reflected across all three axes, in some order? You'd have:

$${i^2}{j^2}{k^2} = \left[ {\begin{array}{*{20}{c}}1&0&0\\0&{ - 1}&0\\0&0&{ - 1}\end{array}} \right]\left[ {\begin{array}{*{20}{c}}{ - 1}&0&0\\0&1&0\\0&0&{ - 1}\end{array}} \right]\left[ {\begin{array}{*{20}{c}}{ - 1}&0&0\\0&{ - 1}&0\\0&0&1\end{array}} \right] = \left[ {\begin{array}{*{20}{c}}1&0&0\\0&1&0\\0&0&1\end{array}} \right]$$
In other words, ${i^2}{j^2}{k^2} = 1$. Additionally you may have observed while crunching the numbers above that ${i^2}{j^2} = {k^2}$.

This may give you an idea*. Here's another thing that may give you an idea: the reason you had $i^2=-1$ with complex numbers was that $i$ rotated all the axes in the plane. By contrast, $i,j,k$ only each rotate two of the three axes in 3-dimensional space.

*the idea being that perhaps combinations of two rotations can give us more interesting results

Well, how do you solve this problem? How do you create a rotation that "rotates all the axes"? Seemingly, you can't. Sure, you can define a rotation that rotates all three of the x, y and z axes, but that would still leave some other axis invariant, which we call "the axis of rotation". Can we define a rotation that leaves no axis invariant?

In three dimensions, the answer is no. Any rotation leaves one axis invariant, and trying to rotate this axis requires rotating it with another axis, and the resulting product rotation still leaves some, calculable axis invariant.

Calculate this axis.

The key is to extend our thinking to four dimensions. Here, you can have pairs of rotations acting simultaneously on two different pairs of axes. Since there are only four dimensions in four dimensions, all four axes are transformed.

Now, the obvious thing to do here may be to define an imaginary number for each pair of rotations in four dimensions -- there are $\left( {\begin{array}{*{20}{c}}4\\2\end{array}} \right)=6$ rotations, and $\left( {\begin{array}{*{20}{c}}6\\2\end{array}} \right) = 15$ such pairs. But this would be too many "basis rotations", and the rotations would not be independent of each other, since rotations in 4 dimensions can be described with only 6 basis rotations.

So how could we make use of our idea of using pairs of rotations as our basis for describing rotations?

The key is to make one of our four axes "special" -- call this axis $t$, and the other three axes $x, y, z$. Instead of considering all 15 rotation-pairs, we only consider the following three:

$$\begin{array}{l}i = (tx,yz)\\j = (ty,\overline{xz})\\k = (tz,xy)\end{array}$$
This is not the only possible representation of the quaternions, of course. Even among complex numbers, you have two possible representations -- you could make $i$ a counter-clockwise rotation, as is conventional, or a clockwise one, i.e. there is a symmetry between $i$ and $-i$. For quaternions, it turns out there are 48 different possible representations -- prove this.

Where $tx$ represents a rotation that sends $t$ to $x$ (i.e. a counter-clockwise rotation on a plane where $t$ is the x-axis and $x$ is the y-axis) and $\overline{xz}$ represents a rotation that sends $z$ to $x$, i.e. the clockwise rotation on a plane where $x$ is the x-axis and $z$ is the y-axis.

It turns out that these pairs -- called quaternions -- in fact allow the representation of 3-dimensional rotations, since you need only a $\left( {\begin{array}{*{20}{c}}3\\2\end{array}} \right)=3$-dimensional basis to represent rotations in 3 dimensions.

Think: Are there any other dimensions that allow such a system to be defined? Can you have, e.g. "hexternions"?

Note that however tempting it may seem, there is no known natural description of special relativity in terms of quaternions. Sorry.

One may work through the algebra of these new quaternions by tracking the position of each axis through the multiplication, and as it turns out, it is indeed much more elegant than the more obvious representation detailed earlier:

$$\begin{array}{l}j = k,jk = i,ki = j\\{i^2} = {j^2} = {k^2} =  - 1\\ijk =  - 1\end{array}$$
In the next several articles, we will look at exactly how 3 dimensional rotations can be represented with quaternions, the relation between quaternions and the dot and cross products through the commutative and anti-commutative parts, and further extensions of the quaternions to higher dimensions.