Showing posts with label relativity. Show all posts
Showing posts with label relativity. Show all posts

Newton's laws in relativity

(Copy of my Quora answer to question "Are Newton's three laws false?" or something to that effect)

They are, but it's intriguing to think about in what way exactly Newton's three laws have been replaced or generalised in relativity.
  1. There are two ways to think about the first law -- the first is "inertial reference frames exist". This is unchanged in special relativity, but general relativity generalises the notion with that of geodesics. The law as it is typically stated -- "stuff moves in straight lines on spacetime unless forced" is generalised to the geodesic equation, $\frac{{{d^2}{x^\mu }}}{{d{s^2}}} = - {\Gamma ^\mu }_{\alpha \beta }\frac{{d{x^\alpha }}}{{ds}}\frac{{d{x^\beta }}}{{ds}}$.
  2. $F=dp/dt$ is generalised to $F=dp/d\tau$ in special relativity, and is replaced by a covariant derivative in general relativity. $F=ma$ has some weirder changes.
  3. The third law is the conservation of momentum. This is replaced in General Relativity by the statement $\nabla^\mu T_{\mu\nu}=0$ ($\nabla$ instead of $\partial$).

    Minkowski everything -- spacetime vectors, rapidity

    Four-vectors and energy-momentum analogies

    Let's look once more at the equation

    $$E=\frac{m}{\sqrt{1-v^2}}$$
    This looks an awful lot like the equation for time dilation. $E$ is the mass as measured by someone who sees the object moving at $v$ whereas $m$ is the mass as measured by someone who sees the object at rest, e.g. by the object itself.

    Similarly, we have the equation $p=vE$, which looks an awful lot like the equation $x=vt$. It therefore makes sense to wonder how far this analogy goes. We could start with analysing the invariant.

    Even if I measure the mass of a 1kg rock as 10kg because of my reference frame, I know that if I brought the bag to rest, I would measure it as 1kg. Much like I can tell people's biological age or look at their clocks to determine their proper time, I can look at the moving thing's mass balance and determine its proper mass $m$.

    If we just wanted $m$ in terms of the "co-ordinates" $E$ and $p$,

    $$m = E\sqrt {1 - {v^2}}  = \sqrt {{E^2} - {v^2}{E^2}}  = \sqrt {{E^2} - {p^2}}$$
    $${m^2} = {E^2} - {p^2}$$
    Or in 4 dimensions,

    $${m^2} = {E^2} - p_x^2 - p_y^2 - p_z^2$$
    We call $m$ the "proper mass". In general, "proper" means "as measured in the rest frame" -- proper time, proper length, proper mass, whatever. This equation is also useful because unlike the previous thing, this also works when $v=1$ (i.e. for light), and reduces to $E=pc$.

    But this looks an awful lot like the spacetime interval.

    That's not all. Consider an object with mass $E$, momentum $p$ and velocity $w=p/E$ in our reference frame $O$. Now boost to a reference frame $O'$ with relative velocity $v$ to $O$. Then the velocity of the object has transformed from $w$ to $\frac{{w - v}}{{1 - wv}}$. So

    $$\begin{array}{c}E' = \frac{m}{{\sqrt {1 - {{\left( {\frac{{w - v}}{{1 - wv}}} \right)}^2}} }}\\ = \frac{{m(1 - vw)}}{{\sqrt {{{(1 - wv)}^2} - {{(w - v)}^2}} }}\\ = \frac{{m(1 - vw)}}{{\sqrt {(1 - {w^2})(1 - {v^2})} }}\\ = \gamma (v)\left( {1 - vw} \right)\gamma (w)m\\ = \gamma \left( {1 - vw} \right)E\\ = \gamma (E - vwE)\\E' = \gamma (E - vp)\end{array}$$
    And

    $$\begin{array}{c}p' = \frac{{m\left( {\frac{{w - v}}{{1 - wv}}} \right)}}{{\sqrt {1 - {{\left( {\frac{{w - v}}{{1 - wv}}} \right)}^2}} }}\\ = \left( {\frac{{w - v}}{{1 - wv}}} \right)E'\\ = \left( {\frac{{w - v}}{{1 - wv}}} \right)\gamma \left( {1 - wv} \right)E\\ = \gamma (wE - vE)\\p' = \gamma (p - vE)\end{array}$$
    Or alternatively

    $$\left[ \begin{array}{l}{E'}\\{p'}\end{array} \right] = \gamma \left[ {\begin{array}{*{20}{c}}1&{ - v}\\{ - v}&1\end{array}} \right]\left[ \begin{array}{l}E\\p\end{array} \right]$$
    In 4 dimensions,

    $$\left[ \begin{array}{l}{E'}\\{{p'}_x}\\{{p'}_y}\\{{p'}_z}\end{array} \right] = \left[ {\begin{array}{*{20}{c}}1&{ - v}&{}&{}\\{ - v}&1&{}&{}\\{}&{}&1&{}\\{}&{}&{}&1\end{array}} \right]\left[ \begin{array}{l}E\\{p_x}\\{p_y}\\{p_z}\end{array} \right]$$
    Which is precisely the transformation for time and position.

    We call vectors that transform like this spacetime vectors or four-vectors. Four-vectors all share the same algebraic properties -- they transform in the same way, they follow vector addition, their norms and in general their dot products are invariant, etc. -- but not necessarily other properties. E.g. energy and momentum have conservation laws, but position and time do not.

    The norm of a spacetime vector is taken as:

    $${\left| {\left[ {\begin{array}{*{20}{c}}{{q_0}}\\{{q_1}}\\{{q_2}}\\{{q_3}}\end{array}} \right]} \right|^2} = q_0^2 - q_1^2 - q_2^2 - q_3^2$$
    Which is distinct from the Euclidean norm, once again telling us that the geometry of spacetime is not Euclidean.

    Four-vectors are perhaps the most beautiful example of the symmetry between space and time. They essentially allow you to replace ordinary pre-relativistic vectors like momentum with vectors that also have a time component alongside three spatial components, because the world is 4-dimensional. You just need to find a quantity that behaves with the vector like time behaves with position -- i.e. you need to show the two quantities transform between each other in a Lorentz transformation sort of way.

    You end up with truly mind-boggling results -- we already saw that mass is the time-component of momentum, which explains why mass produces inertia -- an object with mass already devotes a lot of its momentum to moving forward in time, so the more the mass, the more of this momentum you need to transform into the spatial direction. This is really what is meant by the transformation law $p'=\gamma(p-vE)$ for mass $E$, generalising the Galilean $p'=p-vE$ (change $E$ to $M$ if that makes you happy). It also explains why massless (meaning zero rest mass) things can move at the speed of light.

    Other such four-vectors include:
    • Four-force (time-component: $dE/dt$)
    • Four-current (time-component: charge density, space-component: current density)
    • Electromagnetic four-potential
    Other quantities, like the electric and magnetic fields, even though they follow similar invariants (in the electromagnetic field example $E^2-B^2$), do not combine to form four-vectors, but instead objects called "tensors", which we will eventually talk about.

    Note that during this transformation (giving something momentum), both mass and momentum increase. Similarly, time dilates when you move something around. This is again because $E^2-p^2$, not $E^2+p^2$ is invariant. The latter would correspond to a circular rotation, with invariant circles, whereas the former corresponds to a skew (a "hyperbolic rotation"), with invariant hyperbolae.



    Rapidity and hyperbolic rotations



    Points $(\cos\theta,\sin\theta)$, $(1,\tan\theta)$, $(\cosh\xi,\sinh\xi)$ and $(1,\tanh\xi)$ plotted for varying $\theta$ and $\xi$. While only $\theta$ can be interpreted as an angle too, both $\theta$ and $\xi$ can be interpreted as areas.

    This will be a bit of a DIY section, with some guidance.

    QUESTION 1

    (a) Consider the equation $v' = \frac{{v - w}}{{1 - vw}}$. What trigonometric identity does this remind you of? Could you resolve the differences somehow? (Hint: $v=\tanh\xi$)

    (b) Prove that the Lorentz transformations can be written as

    $$\begin{array}{l}t' = t\cosh \xi  - x\sinh \xi \\x' = x\cosh \xi  - t\sinh \xi \end{array}$$
    (c) Use the hyperbolic analog of angle-addition formulae to show that this is equivalent to, where $\phi=\mathrm{artanh}(x/t)$ is the rapidity of the point $(t,x)$ in the original reference frame.

    $$\begin{array}{l}t' = s\cosh (\phi  - \xi )\\x' = s\sinh (\phi  - \xi )\end{array}$$
    (d) The above result means that rapidity transforms as $\phi ' = \phi  - \xi $ (which is itself nice, because it tells you that velocity at low speeds is approximately equal to rapidity by a factor of $c$) and $(t,x) = (s\sinh \phi ,s\cosh \phi )$. Relate the former to the idea of invariant hyperbolae and the interpretation of rapidity as an area (hint, hint: area sweeped out by a conic section... Kepler).

    QUESTION 2

    (a) Results 1(b) and 1(c) are very similar to the effect of rotations on co-ordinate transformations. Here the linear transformations are skews, not rotations, which is why the formulae are different. Draw as many analogs as you can between rotations and skews in linear algebra. Refer to Article 1103-006. Think about the rotational transformation matrix, etc.

    (b) Consider (a) directly in the context of special relativity. Pretending that Lorentz boosts are simply rotations (which would imply a metric signature (+,+,+,+) and treat time exactly like space), explain transformations between time and position, etc. Relate this to the actual, skew-y Lorentz transformations. Describe how relativity would behave in this theory.

    (c) Write as many relativistic things as you can in the language of rapidity -- the Lorentz factor, the Doppler factor, components of a four-vector (how do $E$ and $p$ look in terms of rapidity), etc.

    (d) Graph the hyperbolic functions and explain why the graphs make the results in 2(b) make sense.

    (e) How does rapidity interpretation make certain things, like $c$ being the maximum speed, natural?

    QUESTION 3

    (a) Consider once again the transformation $\phi ' = \phi  - \xi $. What does this tell you about the relative rapidity $\Delta\phi$? Is this invariant, i.e. do all observers agree on what the relative rapidity between two objects is, like observers did on relative velocity in Galilean relativity?

    (b) Explain why it would be foolish to expect the quantity $\arctan{v}$, the Euclidean angle (as opposed to rapidity, which we may call the "Minkowskian angle"), to have any physical significance. Think about the quantity $r\arctan{v}$ where $r^2=\Delta t^2+\Delta x^2$ (no minus sign).

    It's therefore reasonable to define the dot product on spacetime as $\vec a \cdot \vec b = |\vec a||\vec b|\cosh \Delta \phi $ where $\Delta\phi$ is the relative rapidity/Minkowskian angle/difference in rapidity. This expression implies that $|\vec a|^2=\vec a\cdot\vec a$is manifestly (i.e. obviously) Lorentz invariant, since both norms and relative rapidity are invariant.

    (c) Translate this out of rapidity language, i.e. into a language where rapidity is not used as a parameterisation. You should get $a_0b_0-a_1b_1$ (where 0 and 1 are the temporal and spatial components respectively) in two dimensions.

    The fact that this modified dot product is invariant under a skew is analogous to how the standard dot product is invariant under rotations ("complex skews"). Indeed, it turns out see that the 4-dimensional Minkowski dot product

    $${a_0}{b_0} - {a_1}{b_1} - {a_2}{b_2} - {a_3}{b_3}$$
    Is invariant under skews (between the time axis and some other axis) as well as spatial rotations (and all combinations thereof -- i.e. a general Lorentz transformation), as it contains both a "skew-y" part and a "standard dot product-y" part.



    Some interesting things regarding 2(b):

    A circular Lorentz transformation would transform position and time something similar to this:

    $$\begin{array}{l}x' = \eta (x - vt)\\t' = \eta (t + vx)\end{array}$$
    One can also talk about transforming the positive and negative sides of the axes separately.

    $$\begin{array}{l}x' = \eta (x - vt)\\t' = \eta (t + vx)\\ - x' = \eta ( - x - v( - t))\\ - t' = \eta ( - t - v( - x))\end{array}$$
    Whereas with hyperbolic functions, there is no sign difference, so you only need to transform twice to return. This is linked to you having to differentiate circular functions four times to return, as opposed to twice for hyperbolic functions, all the sign differences between trigonometric and hyperbolic identities, the whole $ie^{i\theta}$ proof of Euler's formula, etc.

    Relativistic dynamics

    There are numerous "proofs" you will finding online of mass-energy equivalence and other dynamic equations in relativity, most of which are wrong ("let's accelerate an object near the speed of light"), circular ("let's prove momentum from energy and vice versa"), simplistic and insufficiently motivated ("it's just empirical"), or just plain inelegant (some really weird collision).

    The motivation for studying relativistic dynamics comes from thinking about conservation of the standard forms of energy and momentum with our new relativistic dynamics. It is easy to demonstrate that $mv$ cannot be conserved in all inertial frames of reference in special relativity. Consider two balls of equal mass colliding inelastically with equal speed $v$ in opposite directions, $+v$ and $-v$. They smash into each other and remain stationary.

    Now boost into one of the balls' frames, say $v$. Now the velocity of the other ball is $2v/(1+v^2)$, so the total initial momentum is $-2mv/(1+v^2)$. But after the collision, we see the thing moving at a velocity of $-v$ (we know this because it was 0 in the original frame), which means the final total momentum is $-2mv$, so momentum is not conserved.

    But we don't like this! If this expression isn't conserved, we can't use it so nicely in calculations and stuff. We want to define momentum in a way that it is conserved. Similar arguments can be used to show that $mv^2/2$ is not conserved, either.

    You may try to derive a conserved expression via similar arguments as the symmetry-based arguments we use in non-relativistic mechanics, swapping Galilean symmetry with Lorentz symmetry where appropriate. The resulting functional equations would be ludicrously complicated, though, and we'd much rather use a different symmetric argument.

    We've made several arguments so far based on known properties of light, and it would make sense to assume other, quantum mechanical properties of light as well. Two such properties are:

    $$\begin{array}{l}p = hf/c\\E = hf\end{array}$$
    This means that we know the behavior of $p$ and $E$ at low velocities, as well as at velocities close to the speed of light. Surely, we're smart enough to fill in the stuff in between?

    Consider the following set-up: a stationary mass m lets out two equal flashes of light in opposite directions, each with energy = momentum (since $c=1$) E/2. We then analyse the same set-up from a boosted reference frame with velocity $v$. This involves a doppler shift in the frequency of each light beam.

    We'll consider this set-up in the following three examples:

    (a) v is small, momentum conservation

    We first consider the case where v is small enough to allow the usage of non-relativistic mechanics. Formally, this means taking the limit as $v\to0$.

    Then the doppler shift factor $\sqrt{\frac{1+v}{1-v}}$ approaches $1+v$ and $\sqrt{\frac{1-v}{1+v}}$ approaches $1-v$. Both energy and momentum are scaled by the same factor since they're proportional to frequency. Now you know why we choose momentum conservation instead of energy conservation -- the total energy is clearly conserved anyway.

    The reason we consider low velocities is that we know the formula for momentum must reduce to the Newtonian $p=mv$, i.e. the initial momentum of the system was $-mv$. The total momentum of the two flashes of light is $((1+v)E/2-(1-v)E/2)=vE$. Since momentum must be conserved, this means the momentum of the mass itself is no longer $-mv$. But its velocity is constant, and still low, so this means some of the mass must have been converted into the energy of the photons. Specifically,

    $$-m_fv-(-m_iv)=vE$$
    Giving us the celebrated equation

    $$E=m$$
    Where $m$ is the amount of mass that was converted into energy. You could, of course, write this in inelegant ways such as $E=c^2m$ or even $E=mc^2$.

    this change in mass is not linked to the whole "relativistic mass" thing we'll be doing later. This decrease in mass is absolute, mass is not conserved, it is also seen in the rest frame, and is required to produce that bit of energy. It's only the derivation that requires boosting into another reference frame, to ensure conservation in all reference frames.

    On a related note, note that conserved and invariant are by no means the same thing, or even related. A quantity is conserved if it doesn't change with time when taken of the whole system. It is invariant if it is the same from all reference frames. The difference isn't even subtle -- proper mass is an invariant in special relativity, but Energy and momentum are conserved.

    Something to think about: why doesn't our argument work in a non-relativistic frame? I mean, we even assumed that v is small. Try to perform the same arguments without relativity -- you will see that since there is no relativistic doppler shift, the result will have a unit of mass being worth an infinite amount of energy -- something you get in the limit $c\to\infty$ -- useless anyway.

    (b) v is not small, energy conservation

    We said the decrease in mass exists in all reference frames. If we found what exactly the decrease in mass $\Delta m$ is in each reference frame, then we'd be able to see how mass transforms under a Lorentz transformation.

    In the rest frame, energy $E$ is released, therefore by energy conservation the energy (or equivalently, the mass) of the object decreases by $E$.

    In the moving frame, one of the beams transforms as $\sqrt {\frac{{1 + v}}{{1 - v}}} \frac{E}{2}$ while the other transforms as $\sqrt {\frac{{1 - v}}{{1 + v}}} \frac{E}{2}$. So the total energy released (i.e. the energy loss of the object) is:

    $$\left( {\sqrt {\frac{{1 - v}}{{1 + v}}}  + \sqrt {\frac{{1 + v}}{{1 - v}}} } \right)\frac{E}{2} = \gamma E$$
    So the mass has transformed as $\gamma m$ under a Lorentz boost of significant velocity.

    We call this mass the "relativistic mass" $M$, and distinguish it from the rest mass $m$.

    Then the following are immediately true:
    • $E = m$ is only true when an object is at rest. In general, $E = \gamma m$. We may call $E_0=m$ the rest energy.
    • $E=M$
    • $M=\gamma m$
    • The increase in mass is essentially the kinetic energy. One may Taylor (or Newton's Binomial) expand out $m/\sqrt{1-v^2}$ to see that the terms start as $m+\frac12mv^2+3/8mv^4+...$, and the higher-order terms vanish at low speeds. Therefore the relativistic kinetic energy is generally $M-m=(\gamma-1)c^2m$.

    In general, we will denote the relativistic mass as $E$ and the rest mass as $m$ unless otherwise stated.

    It is a fad among modern relativity textbooks to claim the phrase "relativistic mass" is a misnomer or even a mnemonic to help kids understand relativity and simply call it the energy, reserving the word "mass" to mean the rest mass. However, this obscures some of the best analogies between spacetime and momentum-energy, as we will soon see -- for instance, the relativistic mass is actually analogous to the co-ordinate time and the rest mass to the proper time/spacetime interval.

    Therefore, we will use the word "mass" to refer to the relativistic mass $E$ and "proper mass" and "momentum-energy interval" to refer to the rest mass $m$. This is a convention in our course only.

    (c) v is not small, momentum conservation

    We may do a similar analysis as above with momentum to arrive at the expression for relativistic momentum.

    The total/net momentum of the light beams in the boosted frame is

    $$\left( {\sqrt {\frac{{1 + v}}{{1 - v}}}  - \sqrt {\frac{{1 - v}}{{1 + v}}} } \right)\frac{E}{2} = \gamma vE$$
    (Note that $E$ represents the total rest energy of the light beams here, as was defined in the question.)

    Therefore $p=\gamma mv$, or $p=vE$.

    You may use this to calculate the relativistic calculation for $F=dp/dt$, but it's simply computation from this point, so I'll just direct you to wikipedia. Come up with an expression for a general directional inertia (simple).

    some people are surprised by the relation $p=vE$, or even remember it wrongly as $E=vp$ because of the seeming resemblance with $E=pc$ at the speed of light (this confusion is because of people not getting the hang of $c=1$ natural units). But it's really nothing new. $E$ is simply the mass. We know momentum equals mass times velocity. This is not new.

    Continued in Minkowski everything -- spacetime vectors, rapidity.

    Minkowski everything -- invariants

    Some philosophers often say silly things like "truth is relative" or worse, "relativity implies that truth is relative".

    Even before relativity, there would be people who gave obviously insincere explanations of this axiomatically incorrect statement -- e.g. "the number 6 viewed from the opposite direction looks like the number 9, therefore truth is relative" or "some people like doughnuts, some people don't, therefore truth is relative". The answer to these kinds of arguments is "someone who sees the number as 6 agrees the other guy sees it as 9, and vice versa", "someone who likes donuts agrees the other person doesn't". The statement donuts are good is not meaningful, except in terms of the donut-liker's neurobiology -- it's equivalent to saying "when you put a donut in his mouth, dopamine is released in his brain". All observers agree that this is the case with him, it's just that dopamine isn't released in the donut-disliker's brain. These statements of absolute truth are absolute.

    Perhaps this gives too much credit to these nonsensical arguments, but the response is similar with relativity. If your parents were bored of raising two children so decided to send your twin brother to Trappist-1 at close to the speed of light, then you would be 80 years old when he returns as a newborn baby. But you do see him as a newborn baby, not an old man, and if you could understand his unintelligible babbling, you would hear that he sees you as an old man on the verge of death, not a kid his age he can play with.

    So biological age is an invariant. Even though you see him as having lived 80 years, you also think that his clock moved a lot slower, which is why he's still an infant.

    But there's nothing special about human biology or biological clocks. Even if the newborn took a clock with him, the time recorded on that clock is an invariant -- all observers agree on what it is.

    Let's try to extract this biological time -- we will call this the "proper time" from the co-ordinate measurements of any arbitrary observer.

    We have:

    $$\Delta t = \frac{{\Delta t'}}{{\sqrt {1 - {v^2}} }}$$
    We write ${\Delta t'}$ as ${\Delta \tau }$, the general proper time according to the moving observer himself.

    $$\begin{array}{l}\Delta \tau  = \Delta t\sqrt {1 - {v^2}} \\\Delta \tau  = \sqrt {\Delta {t^2} - {v^2}\Delta {t^2}} \\\Delta \tau  = \sqrt {\Delta {t^2} - \Delta {x^2}} \\\Delta {\tau ^2} = \Delta {t^2} - \Delta {x^2}\end{array}$$
    One may check that this result is always invariant by Lorentz-transforming $t$ and $x$ and showing $t'^2-x'^2=t^2-x^2$. In a general orthonormal co-ordinate system of spatial co-ordinates (i.e. we don't necessarily take $x$ to be the direction of motion), we may write:

    $$\Delta {\tau ^2} = \Delta {t^2} - \Delta {x^2} - \Delta {y^2} - \Delta {z^2}$$
    Note the resemblance to the Euclidean norm/Pythagorean theorem! If only the minus signs were pluses, this would be the Euclidean norm. This norm is called the Minkowski norm, and the proper time $\Delta\tau$ (or sometimes $\Delta s=c\Delta\tau$, which is the same thing when we set $c=1$) is called the spacetime interval.

    This equation summarises the non-dynamical results of special relativity, and can be treated as an alternative axiomatic foundation for the theory (the "Minkowskian formulation", as opposed to the Einsteinian one we've been discussing so far) -- it's the Pythagorean theorem on spacetime. Unlike in Galilean relativity, where time and space are individually invariant, in special and general relativity, spacetime is invariant -- time and space simply transform between each other leaving the norm of $(\Delta t,\Delta x,\Delta y,\Delta z)$ invariant. This is indeed a rotation ("skew") of this vector, but in Minkowski spacetime, rotations are across hyperboloids, called invariant hyperboloids (or in 2D, hyperbolae), not spheres (or circles). Changing the observer changes the spacetime vector (called four-position), but doesn't take it off this invariant hyperbola.

    Indeed, this means that Minkowski spacetime doesn't have the geometry of Euclidean geometry -- instead, it has a geometry called "hyperbolic geometry", which cannot be embedded in Euclidean space (i.e. we have no way to visualise it).

    Here's another possible motivation for studying invariants:
    Lorentz boosts are essentially rotations in the t-x plane (hyperbolic rotations, actually, or skews, but stick with the analogy for now), so it's often useful to get an intuitive feel for them in special relativity by comparing boosts to rotations on some other plane, like the x-y plane. So let's do that.

    Consider if you were measuring the y-length of a stick on the x-y plane -- clearly, this depends on your frame of reference. A co-ordinate system in which the stick lies on the y-axis clearly gives you the maximum value of this y-length, a co-ordinate system in which it lies on the x-axis clearly gives you a value of 0.



    So the specific co-ordinate dimensions $(x, y)$ of the stick depend on your reference frame. But we can also be interested in the real lengths of sticks, because this is invariant in all reference frames. This can be calculated easily using the Pythagorean theorem:

    $$\psi=\sqrt{x^2+y^2}$$
    (Note that the invariance is not the only thing that is important, but also that it allows you to define a polar co-ordinate system where $x=\psi\cos\theta$, $y=\psi\sin\theta$.)

    If you accept that it can be useful to know the dimensions of objects on their own axes, it's clear that the same principle applies on the t-x plane. Here, the "rotations" are skews, the trigonometry is hyperbolic trigonometry, the Pythagoras theorem is $\tau=\sqrt{t^2-x^2}$ and instead of the proper time being the highest point of a circle it is the lowest point of a hyperbola.

    But the same principles still apply -- if you see someone blast a toddler off into outer space at a high speed then return, you might measure the toddler as having taken a hundred years to return, but you and the toddler both agree (assuming he isn't dead yet from starvation) that he's only aged a year. This biological time, or proper time, is an invariant.
    (From my answer on Physics Stackexchange to Why is invariance important?)

    A related fact is an intuitive explanation for the speed of light being the maximum achievable speed -- all observers have a fixed speed ($ds/d\tau$) through spacetime, which is the speed of light -- this is essentially a tautology. A stationary object has no speed through space, so $dx^2+dy^2+dz^2=0$ so it moves at $c$ through time ("co-ordinate time" $t$ -- as opposed to proper time), i.e. $d(ct)/d\tau=c$. On the other hand, when an object moves at the speed of light, its clock has stopped -- we see $d(ct)/d\tau=0$. The velocity cannot exceed the speed of light, because the object simply doesn't have that much speed -- it doesn't have any more speed to take from its time-speed. Another way of saying this is that an invariant hyperboloid never crosses the light cone.

    It's important to keep in mind that in our argument above, time, position and velocity are always with respect to some other observer (again, this is also implied by the Minkowskian formulation, as $dx$, $dt$ etc. are in the frame of some observer). So the point is really that "no observer can see an object going faster than light, because to keep the speed through spacetime fixed, the Lorentz transformation would have to map the time to an imaginary number ($\Delta t^2 < 0$).

    We will see later that there are other quantities that transform between each other like time and space. Then we will see that the four-position is just another vector among a class of vectors called four-vectors.

    (Note of caution: often, $\Delta s^2$ instead of $\Delta s$ is called the spacetime interval. When you hear the phrase "negative spacetime interval", this is typically what is being referred to.)

    (Note: Because both $\Delta s^2$ and $-\Delta s^2$ are invariants, sometimes $- {c^2}d{t^2} + d{x^2} + d{y^2} + d{z^2}$ is called the spacetime interval instead. This choice is called the "metric signature" and is denoted by $(+---)$ and $(-+++)$ respectively. The first is also called the particle physics convention, the quantum field theory convention, the West coast convention, the time-like convention and the mostly-minus convention. The second is also called the cosmology convention, the general relativity convention, the East Coast convention, the space-like convention and the mostly-plus convention. However, $\Delta\tau^2$ is always defined via the time-like convention, as it is the proper time.)

    You might be tempted to say that Minkowski spacetime is simply 4-dimensional Euclidean spacetime with one of the dimensions being $ict$ instead of $ct$. However, this doesn't actually make Minkowski spacetime Euclidean -- for instance, Minkowski spacetime allows distinct points in spacetime to have a zero spacetime interval between them, something not possible with a Euclidean distance function. After all, the norm of a complex number $t + ix$ is still $\sqrt{t^2+x^2}$, not $\sqrt{t^2-x^2}$.

    You might be tempted to rewrite the equation as $d{t^2} = d{\tau ^2} + d{x^2} + d{y^2} + d{z^2}$. But since $d{t^2}$ is not an invariant, this obscures the true geometry of Minkowksi spacetime, which is hyperbolic, not Euclidean. Similarly, equations like $m^2 = E^2-p^2$ (where $m$, $E$ and $p$ are the proper mass, relativistic mass and momentum respectively -- we will later derive this) should not be written as $E^2=m^2+p^2$.

    You might recall some equations in physics that seem to exhibit the same kind of symmetry between space and time as the spacetime interval -- $-c^2t^2$ and $x^2$ showing a symmetry. An example is the wave equation for light, $\frac{1}{c^2}\frac{\partial^2u}{\partial t^2}-\frac{\partial^2u}{\partial x^2}=0$. This is actually the reason why Maxwell's equations are already Lorentz invariant, and indeed, we will see that this symmetry will be our criterion for Lorentz invariance.

    (Technical note: Formally speaking, Minkowski spacetime doesn't actually have hyperbolic geometry itself. What it does have are sub-manifolds with a hyperbolic geometry.)

    We may divide spacetime intervals into three categories: space-like (outside the light cone), light-like (on the light cone) and time-like (inside the light cone), corresponding to the cases $\Delta s^2<0$, $\Delta s^2=0$ and $\Delta s^2>0$ respectively (in the cosmology convention, it is exactly disrespectively). The fact that you cannot influence space-like separated events, i.e. cannot travel faster than light is the same as saying "you cannot transverse an imaginary proper time".

    Saying the speed of light is fixed for all observers is equivalent to saying that the statement $\Delta s^2=0$ is invariant, since $\Delta s= \sqrt{c^2\Delta t^2-\Delta x^2}$ and $x=ct$. We now know that $\Delta s^2=n$ is invariant for all $n$, not just 0.



    The image above shows invariant some hyperbolae plotted -- $\Delta s^2=-3$, $\Delta s^2=-2$, $\Delta s^2=-1$, $\Delta s^2=0$, $\Delta s^2=1$, $\Delta s^2=2$, $\Delta s^2=3$. Note how the hyperbolae never cross the light cone -- implying the existence of an absolute future, an absolute past, an absolute left and an absolute right.

    The correct resolution of the twin paradox

    (Note: in this article, we will use the phrase "seeing" to mean "considers simultaneous to". E.g. "We see Betelgeuse explode today" doesn't literally mean that we can observe Betelgeuse going supernova today, but rather that we will see it happening ~600 years later, thus calculating that Betelgeuse exploded today, i.e. in a time simultaneous to the present.)

    The time dilation equation, $\Delta t = \gamma \Delta \tau$, often looks utterly wrong.

    "If Observer $A$ observes $A'$'s clocks as being slower and $A'$ observes $A$ as being slower, who's right? Surely, we cannot have $\Delta t = \gamma^2 \Delta t$!"

    Hopefully, you should be able to answer this paradox based on your understanding of the relativity of simultaneity.

    This is in fact exactly why we started out with a discussion of simultaneity being relative -- Observer $A$ is seeing a different point on the worldline of $A'$, and Observer $A'$ is seeing a different point on the worldline of $A$.

    Again, a spacetime diagram is instructive:


    The point that Observer A sees on A' is not looking back at him at all -- it's seeing his past. The point on A's worldline that A thinks is simultaneous to himself -- at that point, A' is thinking of a different point on A's worldline to be simultaneous to himself.

    OK. But what if $A'$ returns to meet $A$. What if, for example, $A$ and $A'$ are two twins, and $A'$ goes to Betelgeuse at a non-zero speed and comes back? Who's older?

    The key is the change in velocity when $A'$ turns around. It's not that special relativity doesn't apply any longer (this is a common myth), but rather that the axis of simultaneity changes here.


    In other words, $A'$ sees $A$'s clock suddenly speed up rapidly for the time that he turns around. If the turnaround is instantaneous, he sees $A$'s clock suddenly skip ahead a few years and continue at a dilated rate. So $A'$ agrees: $A$ is older -- he rapidly aged at a certain point during $A'$'s journey, compensating for his otherwise apparently slow aging. When $A'$ comes back, he does come to a futuristic world where time travelers are culled at sight.

    This is the resolution to the twin's paradox. The key point is that while observers may disagree on the simultaneity of spatially separated points, they cannot disagree on the simultaneity of points that lie on each other (i.e. when the twins meet).

    Think: What if $A'$ were going in a circle to return to his original starting point? Or worse -- what if the universe itself had a spherical geometry that $A'$ comes back to the same spot without being acted on by a force?

     In the latter case mentioned above, you truly need general relativity, because we made spacetime spherical. But what if we use a cylindrical spacetime instead? (It turns out cylinders do not have "intrinsic curvature", because distance functions are not affected when a cylinder is made from a flat sheet.) I.e. where only one spatial dimension is chosen as curved?

    It turns out that the first postulate of special relativity (you can't do an experiment to determine your absolute velocity) gets broken in this context (because the radius of the cylinder gets Lorentz contracted when you're moving with respect to it and you can measure this by sending light signals across the cylinder).

    To answer the question in an arbitrary geometry, including a curved one, we have to compare the "spacetime interval" tranversed by each twin. We will define this in a coming article, and show that it is invariant. In general relativity, it is simply the calculation of this interval that changes a bit.

    There are some other, less trivial paradoxes in special relativity. An example is Rindler's grid paradox -- whose resolution requires realising that rigid objects do not exist in special relativity.

    Lorentz transforms lives

    Duration

    In your years as an infant reading up stuff on wikipedia, you might've seen formulae such as

    $$\Delta t = \frac{{t'}}{{\sqrt {1 - {v^2}} }}$$
    Or simply $t=\gamma t$. From our knowledge of the Lorentz transformations, we certainly know that the scale on the time axis changes. It would be interesting to find out exactly how this might be observed in real life -- I mean, we know how time as a co-ordinate transforms, but how does duration -- the interval between two points in time -- transform?

    You might be tempted to do calculations like ${t'_1} - {t'_2} = \gamma \left( {{t_1} - vx} \right) - \gamma \left( {{t_2} - vx} \right)$, much like people are tempted to sign up for "get rich quick" scams. Doing so woulbe reckless and stupid.

    What we need to do is first precisely formulate what we're looking for. We ask:

    Suppose there is a clock moving at a constant velocity v relative to me. In my time, how long does it take for the moving clock to tick by 1 second? Assume that we synchronised our clocks in the beginning, i.e. the moving clock and my own clock showed exactly the same time at t = 0 when our positions coincided.

    Let's draw a spacetime diagram.


    Point A represents the event "moving clock ticks the one second mark". Since lines parallel to the x-axis link points that we (i.e. the stationary observer) consider simultaneous, we draw a horizontal line connecting Point A and the t-axis (remember, we want to find out what tick of our clock is simultaneous, according to us, with 1 second elapsing on the moving clock). Mark this point of intersection B. Then we are interested in finding the duration OB, which we call $t$ in terms of OA, which we call $t'$.

    Well, from the Lorentz transformations we know that $t' = \gamma \left( {t - vs} \right)$. We also know, geometrically, that $s = vt$, so we may write $t' = \gamma t \left( {1 - v^2} \right)$, i.e. $t'=t\sqrt(1-v^2)$, or $t=\gamma t'$.

    In general, for the duration between two events (where stuff might not pass through the origin at the right time), we may say $\Delta t = \gamma \Delta t'$. This phenomenon is called time dilation.

    Distance

    We do the same sort of calculation for distances, first operationalising what we mean:

    If I hold out a ruler to measure the length of a metre-stick (i.e. something that is 1 metre in its own reference frame) moving at speed v relative to me, what would be the length I measure?

    Once again, we draw a spacetime diagram.


    This is a little trickier -- when measuring the length of an object, we do so by measuring the two ends of the object simultaneously (or rather, what is simultaneous according to us). However, what is simultaneous for us is not what is simultaneous for the rod. While the rod's reference frame holds O and L as simultaneous, we actually choose another point on the worldline -- K -- as simultaneous with O, because it lies on the x-axis.

    Then:

    $$x'=\gamma\left(x+SK-vh\right)=x'=\gamma\left(x+vh-vh\right)=\gamma x$$
    Hence $x=x'/\gamma$, i.e. length/distance in the direction of motion is contracted under a Lorentz transformation.

    Back when I was an infant, I was confused about why it was that time got dilated (multiplied by $\gamma$), while length got contracted (divided by $\gamma$). Well, now you know -- the two phenomena aren't temporal-spatial analogs of each other at all! Length contraction is a result of measuring the two ends of a distance simultaneously

    Speed

    We have been interested, since the beginning of this series, in finding out how velocities and speeds transform under a Lorentz transformation. Once again, we formulate our question precisely as follows (if you've done DIDYMEUS, you should understand how this forces us to accept logical positivism):

    Suppose O' is moving at velocity v with respect to O. In O', the velocity of object K is w. What is the velocity of K in O?

    Once again, we draw a spacetime diagram.


    So given $x'/t'$, how would we find $x/t$?

    Well, here's an idea: we know the Lorentz transformation associated with the velocity $w$. So we just use simple matrix multiplication to find the compound transformation, and figure out what velocity is associated with this transformation.

    In other words, we write $L(v)L(w)$ as the co-ordinate system of $K$ with respect to $O$. Performing the matrix product,

    $$\begin{array}{c}\gamma (v)\left[ {\begin{array}{*{20}{c}}1&v\\v&1\end{array}} \right]\gamma (w)\left[ {\begin{array}{*{20}{c}}1&w\\w&1\end{array}} \right] = \frac{1}{{\sqrt {\left( {1 - {v^2}} \right)\left( {1 - {w^2}} \right)} }}\left[ {\begin{array}{*{20}{c}}{1 + vw}&{v + w}\\{v + w}&{1 + vw}\end{array}} \right]\\ = \frac{{1 + vw}}{{\sqrt {\left( {1 - {v^2}} \right)\left( {1 - {w^2}} \right)} }}\left[ {\begin{array}{*{20}{c}}1&{\frac{{v + w}}{{1 + vw}}}\\{\frac{{v + w}}{{1 + vw}}}&1\end{array}} \right]\\ = \frac{1}{{\sqrt {\frac{{{v^2}{w^2} + 1 - \left( {{v^2} + {w^2}} \right)}}{{{v^2}{w^2} + 1 + 2vw}}} }}\left[ {\begin{array}{*{20}{c}}1&{\frac{{v + w}}{{1 + vw}}}\\{\frac{{v + w}}{{1 + vw}}}&1\end{array}} \right]\\ = \frac{1}{{\sqrt {\frac{{{v^2}{w^2} + 1 + 2vw - {{\left( {v + w} \right)}^2}}}{{{v^2}{w^2} + 1 + 2vw}}} }}\left[ {\begin{array}{*{20}{c}}1&{\frac{{v + w}}{{1 + vw}}}\\{\frac{{v + w}}{{1 + vw}}}&1\end{array}} \right]\\ = \frac{1}{{\sqrt {1 - {{\left( {\frac{{v + w}}{{1 + vw}}} \right)}^2}} }}\left[ {\begin{array}{*{20}{c}}1&{\frac{{v + w}}{{1 + vw}}}\\{\frac{{v + w}}{{1 + vw}}}&1\end{array}} \right]\\ = \gamma \left( {\frac{{v + w}}{{1 + vw}}} \right)\left[ {\begin{array}{*{20}{c}}1&{\frac{{v + w}}{{1 + vw}}}\\{\frac{{v + w}}{{1 + vw}}}&1\end{array}} \right]\\ = L\left( {\frac{{v + w}}{{1 + vw}}} \right)\end{array}$$
    Interestingly, this product is commutative. We may thus write:

    $$L\left( v \right)L\left( w \right) = L\left( {\frac{{v + w}}{{1 + vw}}} \right)$$
    The reason this is a useful form to write the velocity addition formula is that it conveys the precise positivist sense in which velocity is transformed: as it is observed in the Lorentz transformation of things associated with it.

    One may let one of the velocities be $c$ and confirm that $c$ is the same in all reference frames.

    What happens when the Lorentz boost is in the direction perpendicular to the direction of motion? Well, distance is not contracted, but time is still dilated, and the velocity is reduced by a factor of $1/\gamma(v)$ where $v$ is the velocity of the observer. This ensures, and you can verify, that the resultant velocity in the new frame doesn't exceed $c$ even by the Pythagorean sum.)

    Relativistic doppler shift

    This is a surprisingly important lemma to our future derivation of the equation $E=mc^2$, so make sure you're clear with it. Also, it tells you that speeding through a red light might cause it to turn into gamma radiation, if you go fast enough.

    We're interested in finding out how the frequency (i.e. colour) of light changes with respect to a moving observer, accounting for all relativistic effects. Frequency is just the inverse of the time period, which is the time interval between two wavefronts.


    The red vertical line is the worldline of the source, the blue line is the worldline of a moving observer and the black vertical line is of course the worldline of the observer we consider stationary. The purple lines are the wavefronts emitted by the source. Suppose one wavefront hits the worldlines of both the stationary and moving observers at the origin. Another wavefront hits quite later.

    We first find the co-ordinates of the point of intersection between the blue worldline and the worldline of the second wavefront in the stationary co-ordinate system. We simply find the equations of the lines and set: $x=vt$, $t=T-x$ so that:

    $$\begin{array}{l}t = T - vt\\(1 + v)t = T\\t = \frac{1}{{1 + v}}T\\x = \frac{v}{{1 + v}}T\end{array}$$
    Now we may easily calculate the co-ordinate $t'$, which is the same as $T'$:

    $$T' = t' = \gamma \left( {\frac{1}{{1 + v}}T - vx} \right) = \gamma T\left( {\frac{1}{{1 + v}} - \frac{{{v^2}}}{{1 + v}}} \right) = \gamma T\left( {1 - v} \right) = \sqrt {\frac{{1 - v}}{{1 + v}}} T$$
    Then

    $$f' = \sqrt {\frac{{1 + v}}{{1 - v}}} f$$
    This is an important result! It means that even though a photon has the same speed however fast you chase it, you do see it getting less and less energetic.

    Sometimes you will see the inverse coefficient $\sqrt{\frac{1-v}{1+v}}$ -- this involves an observer moving away from the source.

    How fast would you need to go for a red light to become gamma radiation? Well, it means that $\sqrt {\frac{{1 + v}}{{1 - v}}} =f'/f=10^{19}/(4*10^{14})=2.5*10^4$, i.e. $(1+v)/(1-v)=6.25\times10^8$. Solving for v, one sees that it must be within 1m/s of the speed of light.

    (yes, the title is a joke)

    Light cones and causality

    Before we do anything further, I'd like to prove a point I made earlier about causal events.

    Recall that we earlier showed that the simultaneity of two events depends on the observer. Similarly, we know that the order of two events depends on the observer.

    In the above diagram, we have chosen too reference frames, Lorentz boosted with respect to each other, where the event A is the origin of both. Now while the black reference frame (call it $O$) measures B's time as being after A's time, the red reference frame ($O'$) measures B as occurring earlier. The figure below illustrates this phenomenon:


    OK.

    Now notice how $x'$ is the set of all points that $O'$ measures as simultaneous to $A$, and $x$ is the set of all points $O$ measures as simultaneous to $A$. What this means is that if and only if the event B is between the x-axis and the x'-axis will the two observers disagree on which happened first.


    Why is this significant? Well, no matter how fast $O'$ moves, it can never move faster than light, so $x'$ can never cross the blue boundary. If $O'$ moves in the opposite direction, $x'$ will go below the x-axis, but can once again not cross the blue boundary below. Therefore, the set of all events that observers may disagree on whether they happened before or after A (the origin) are those shaded in purple below:

    Think: What about the events on the blue lines? Are those included or excluded?

    These are all the events outside the light cone of the event A. The light cone is the set of elements not shaded in the diagram above. The part with positive t is called the "future light cone", and the part with negative t is called the "past light cone" of the event.

    Why is this significant? Well, the past light cone is the set of all events that could have possibly caused (i.e. influenced) the event at A, and the future light cone is the set of all events that A could possibly cause (influence).

    In other words, the light cone is the set of all events causally connected to A! You should be able to explain why by now -- the borders of the light cone are paths of light beams sent away from or coming towards A. For A to cause or influence another event, it must be able to send some message, object or any form of information to that event. Since nothing can move faster than light, this message cannot move faster than light, therefore it cannot influence anything outside its future light cone. An analogous argument can be easily made for why nothing outside the past light cone can have influenced or caused A.

    The result is extremely significant! All observers agree on the order of two events A and B if A and B are in each other's light cones (if A is in B's future light cone, B is in A's past light cone, and vice versa). So the issue we earlier mentioned, is solved.

    (Note: in 2+1 dimensions, the light cone would be an actual 3-dimensional cone, and in 3+1 dimensions, the light cone would be a 4-dimensional cone. If we could actually see a light cone in three spatial dimensions, it would be a sphere expanding from the event at the speed of light.)

    ---

    We can also call the events in the future light cone the "absolute future" of $A$, and the events in the past light cone the "absolute past" of $A$. Similarly, one may call events outside the light cone the "absolute left" and "absolute right" of $A$ -- there are no observers whose worldlines pass through $A$ such that these events are to their left but to the right of $A$, or vice versa.

    This is a pretty obvious conclusion of the fact that nothing can move faster than light. Indeed, we will often find that conclusions regarding time are obvious in the context of space, or vice versa. This fact, and the symmetry between space and time, might help us "guess" many conclusions in special relativity from what we already know.

    ---

    Note that our proof is equivalent to requiring that no physical reference frame can move faster than the speed of light, i.e. the x' axis cannot cross into the light cone.

    In fact, it is equivalent to our earlier proof that nothing can move faster than light, where events A and B are "light hitting the hi-tech wall" and "train hitting the high-tech wall". There is a causal link between the two, since what happens when the train hits the high-tech wall is affected by whether or not the light hits the high-tech wall.

    This gives us an important insight regarding the thought experiment: we implicitly assumed that time cannot run backwards. This principle is known as causality, and we have just demonstrated that it is equivalent to the statement that things cannot move faster than light (i.e. locality).

    Introduction to special relativity

    Often, one wonders why some major paradigm shift took so long to occur. We ponder this in the context of political economy, for instance -- with regards to the Neolithic Revolution (the invention of agriculture, circa 9000 BC) and the Industrial Revolution (which one may trace as the ultimate conclusion of a series of events that began with the end of feudalism in Europe in the 1400s).

    We also often ponder this in the context of scientific achievements. Why, for instance, did it take till Einstein for an insight as key as relativity to be discovered? 

    While this question is sometimes tricky to answer, the question is very clear in the context of relativity.

    Special Relativity was developed as a resolution to the failure of Galilean Relativity to accomodate the predictions of Maxwell's electromagnetism. It turns out that while Maxwell's electromagnetism was fine (it is "Lorentz invariant"), mechanics itself needed to be fixed. A key insight from relativity with regards to electromagnetism is, in fact, that magnetism is a relativistic effect. Magnetism is what you get when electricity undergoes a Lorentz transformation, i.e. when the charge starts moving. It is just like the effect of velocity on mass, for instance, or distance, or duration.

    (Source: XKCD - 1489)
    The precise contradiction was as follows: Maxwell gives an absolute value for the speed of light, but Galileo says that no absolute speed exists -- it depends on the reference frame. For instance, if a train travelling at v with respect to the ground) blares light at speed c (in its own reference frame), then according to the ground's reference frame, the light should be travelling at c + v.

    Light is not important! Even though the initial result that spurred relativity came from the theory of electromagnetism and light, relativity itself produces the same predictions for anything that travels at the same speed that light does -- any massless particle, in general. I.e. the curious case of light in relativity is not a result of its particle-physics-y properties, but its kinematic properties.

    This prediction (by Galilean relativity) is fundamentally a result of the nature of the Galilean transformation (and this is the transformation that Einstein sought to change). This is the transformation that tells you how to transform co-ordinates (or really anything) between inertial reference frames with the same origin.

    Suppose Observer $O'$ is moving at speed $v_{O'}$ with respect to Observer $O$. Now consider the time and position of some event $P$ to be $(t, x)$ in reference frame $O$. If we're dealing with four dimensions, then $x$ and $v$ are of course, three-vectors. Then what's the position of event $P$ according to $O'$? Well, at time $t$, $O'$ would be $v_{O'}t$ to the right of $O$, hence the position of $P$ would be measured as $(t,x-v_{O'}t)$.

    This is the Galilean transformation

    $$G(w):\,\,\left[ {\begin{array}{*{20}{c}}t\\x\end{array}} \right] \to \left[ {\begin{array}{*{20}{c}}t\\{x - wt}\end{array}} \right]$$
    One may also write this as the matrix:

    $$G(w) = \left[ {\begin{array}{*{20}{c}}1&0\\{ - w}&1\end{array}} \right]$$
    Perhaps the asymmetry of the matrix bothers you. It bothers me too. And fortunately for us, it bothered Einstein too, and he actually did something about it, rather than rant about it on a blog. In fact, the asymmetry of this matrix corresponds directly to the asymmetry of time and space in Galilean relativity. This is pretty clear from the form of the Galilean transformation, and should also be obvious from your knowledge of linear algebra (if it's not, you should go and read up the first few chapters of the linear algebra course). As we will see, the symmetry between space and time will arise quite neatly from our postulates.

    One may plot this transformation on a spacetime diagram. Below shows a spacetime diagram viewed from the perspective of $O$, where the transformed reference frame $O'$ is shown as well.


    A spacetime diagram is essentially a displacement-time graph, where the displacement function is considered a transformation of the t-axis.We make the following observations:
    • The t' curve is the worldline of observer $O'$, i.e. the path taken by $O'$ in spacetime. The $x$-axis is not transformed (this is the asymmetry we were talking about earlier).
    • The x-axis is essentially the set of all events in spacetime such that $t=0$, i.e. are simultaneous to "the present". Surely, these points must be the same to all observers, since whether $t=0$ or $t=1$, or whatever value t holds, is independent of the observer? 
    • Within the reference frame $O$, $O'$'s reference frame seems squished up. But since no reference frame is special, within $O'$'s reference frame, $O'$ will look normal, and $O$ will look transformed, specifically by the velocity $-v$ (so the axis t is tilted from the "normal" $t'$ by the same angle in the opposite direction). This is just the inverse transformation.

    So those were the Galilean transformations, which we know are incorrect (differentiate $x' = x-wt$ so you get $\dot{x}'=\dot{x}-w$ -- we know from Maxwell that this is incorrect with regards to light). Before we derive the correct transformations (called the "Lorentz transformations"), we'll first take a detour to prove some significant results in special relativity, which will also give ourselves an idea of how powerfully predictive our two axioms are.

    (A note on notation: we will use units of distance and time such that the value of c is unity. For instance, lightseconds and seconds, etc. This is useful, because it eliminates c from our formulae and helps expose the symmetry between space and time.)

    1. Nothing can travel faster than light

    We consider a thought experiment, where an object $O'$ travelling at speed v (in reference frame $O$) releases light in the positive x-direction. The speed of this light is, of course, c in all reference frames. According to $O$ (i.e. an observer in $O$), the speed of this light is $c$, but the speed of light relative to the object is $c-v$.

    That's okay. But now consider if $v>c$. Then $c-v$ is negative, i.e. $O$ observes the light ray being emitted at some speed in the other direction with respect to the object, i.e. it sees $O'$ farting out the light ray, rather than vomiting it up. Whereas according to $O'$, he is stationary, and the velocity of light is still $c$ in the positive x-direction.

    Simulation of how one would observe a "tachyon" -- a hypothetical particle that can move faster than light.
    Why is this inconsistency problematic? Well, suppose there is some hi-tech wall somewhere further down the positive x-direction, which functions in the following way:
    • If light is shone on it, the hi-tech wall stops working.
    • If the object collides into the wall while it is working, it sets off blaring alarms and sends planes flying into buildings so everyone knows about the event.

    According to Observer $O$, the object collides with the wall first, before the wall stops working. His children die in a plane crash and he ends up drunk and homeless.

    According to Observer $O'$, however, he can never catch up with the light ray, and he bangs into a dysfunctional wall, and nothing happens. He turns around and waves at $O$, who is sober.

    We have ended up at a logical contradiction, and our only solution is to say that such an object that travels faster than light, $O'$, does not exist.

    When I narrate this proof to people, they are quick to ask "If we're talking about the response of the wall, shouldn't we only care about the wall's reference frame?" Well, no, because that privileges a reference frame. The laws of physics -- including the laws of the wall -- are valid in all reference frames, and an external observer shouldn't see the wall giving a response inconsistent with the laws of the wall. I mean, it's possible to have a wall that functions in the way demonstrated in the question in all reference frames.

    This is a rather surprising result. If you're travelling at .99999999c, can't you just supply a bit of energy to go 4m/s faster? As we will see, it turns out the laws of dynamics also change in special relativity, and this "bit" of energy is infinite.

    "Aha!" you say, "Maybe you can never choose an inertial reference frame travelling faster than light with respect to your reference frame, but what if you choose a reference frame travelling at 0.6c, then another reference frame travelling at 0.6c with respect to that reference frame. Then wouldn't this third reference frame be travelling at 1.2c with respect to us?" Again, it turns out that even the velocity addition formula is changed in special relativity.

    When we say "velocity addition formula", we mean "co-ordinate transformation between reference frames in relative motion with each other", i.e. if an observer on the train moving at speed v relative to the ground measures the speed of something to be w, then what's the speed of that thing wrt the ground? That's the velocity addition we're talking about.

    We're not talking about velocity addition within the same reference frame. If we see two light beams shining towards each other, we do see the space closing at speed 2c.

    2. Relativity of simultaneity

    We know from linear algebra that a linear transformation in $\mathbb{R}^n$ can be described fully by the images of $n$ linearly independent vectors. These images, written next to each other, form the matrix of the transformation in the basis comprised of these vectors. One such set of linearly independent vectors is the standard basis.

    We're working towards an expression for the Lorentz transformation. To find out how the unit vectors in the t, x, y and z axes transform, it is sufficient to find out how the axes themselves transform, and what the scale is on these transformed axes (e.g. the identity transformation and a scaling of two both leave the axes unchanged, but the scales on the transformed axes are different).

    For readers with a reasonable knowledge of linear algebra: we know two sets of eigenvectors of the Lorentz transformation. The fact that these are not eigenvectors of the Galilean transformation is equivalent to the problem of Galilean transformation not respecting the invariance of the speed of light. The problem of special relativity is therefore equivalent to finding a matrix with these eigenvectors, with eigenvalues that respect some symmetry properties we will see later.)

    We consider a general 1+1-dimensional spacetime diagram in the reference frame of $O$. Obviously, the t-axis is the worldline of our observer/the origin of our observer.

    What, exactly is the x-axis? Well, any P-axis is essentially the set of points such that all co-ordinates except P are 0. The x-axis is the set of points such that t = 0.

    In other words, the x-axis is what the observer regards as the present. If the x-axis were transformed in any way, then it would mean that the idea of what's the present and what's the past also depends on the observer. In general, any line parallel to the x-axis is a line of simultaneity (i.e. events that occur at the same point in time), and if the x-axis is transformed, the conception of simultaneity depends on the observer.

    So it makes sense to study simultaneity in our quest to find the Lorentz transformation.

    The relativity of simultaneity can be illustrated with the following thought experiment: suppose we have two sources of light, $S_1$ and $S_2$, which (in the reference frame of Observer O) release a pulse of light at the same instant $t=0$. How does Observer O know this? Well, he is situated at the midpoint of the two sources, and knows the distance between him and each source to be s, so when he sees the two pulses simultaneously at $t=s/c$, he knows that each pulse was released $s/c$ time earlier, i.e. at $t=0$.

    Now consider another observer $O'$, moving parallel to the light coming from $S_1$ at speed $v$. It happens to be that at the instant the two pulses collide at the origin of $O$, $O'$ also crosses this point.

    It is important to note that he observes the collision of the two pulses at the same instant as $O$ does. This occurs at a single event (a single point in spacetime), and all observers must agree on what happens at this event (this is from the principle of relativity).

    However, what $O'$ disagrees with $O$ on is on the simultaneity of the release of the light pulses itself. Observer $O'$ considers himself to be an ordinary, stationary observer. He has seen the point of intersection running away from the light emanating from $S_2$ and towards the light emanating from $S_1$. Therefore light -- whose speed is $c$ -- emanating from $S_2$ has to catch up with the intersection point, the distance closing at a speed of $c-v$, while light emanating from $S_1$ meets the intersection point with the distance closing at a speed of $c+v$.

    So for them to meet at the intersection point the same distance away from each source, $S_2$ must have released its pulse earlier than $S_1$. How much earlier? Well, let's not go there too fast -- we still don't know if distances themselves change in the reference frame, i.e. what the scale on the transformed x-axis is.

    Being simultaneous to doesn't mean "see". To see something, light (or anything else) must travel from that event to the observer's worldline. For instance, if Betelgeuse were to go supernova today, we are not simultaneous with the event, which we calculate to have happened hundreds of 600 years ago.

    You might wonder, then: what if the two events are causally connected? I.e. what if there is a link that ensures $S_2$ turns on in response to $S_1$ turning on? Well, it turns out that in such a case, the order of the two events is in fact preserved. We will see why later -- the reason has to do with the connection between causality and light cones.

    3. Transformation of the x-axis

    Let's think about how one would actually determine some event to be simultaneous to us right now. Well, obviously one must observe the event, for which we must detect the light coming from that event. Suppose we just observed the light we get from that event right now, at $t=0$. Could we say that gives us an event simultaneous with the present? Well, of course not. We know that the light traveled some distance to get here, so we're observing an event some time into the past. We would figure out how much into the past by determining the distance of the event from us. How would we do this? Well, we would reflect light off the object and see how long it takes to return.

    So suppose we use this method to determine which event is simultaneous to us. Releasing a light ray now would be too late -- we would get an event in the future in the reflected ray. Instead, we should have released a light ray d/c seconds ago, and if the light ray returns d/c seconds into the future,  .,the object is d away from us, and the reflected ray shows the event simultaneous to us right now.

    For instance,  if we shot a light ray at Betelgeuse in 1400, then the reflection we get in 2600 will be an image of how Betelgeuse looks today, in the year 2000 (on the scale of the universe, 17 years is no big deal -- for example, it is clearly an insufficient age to learn the difference between 2000 and 2017), because the star is 600 lightyears away.


    (Note that the slope of a light ray on a spacetime diagram is always 1/c in some direction -- since we're using natural units, this is just a slope of 1.)

    So here we have a general property -- and in fact a defining property -- of the x-axis: it is the set of all points such that if you sent a ray to bounce off the point $-a$ seconds ago, it will return to us $a$ seconds later.

    Why is this useful? Well, if this reference frame were viewed in some other observer's reference frame, it would still be true (by the principle of relativity).

    What do we mean?

    Label the axes of this co-ordinate system as $t'$ and $x'$:


    Then how would the points of spacetime this reference frame map into another reference frame? Well, perhaps something like this:


    What do we want to know about this diagram? Well, the direction of the x' axis relative to the x-axis, i.e. the angle between them.

    What do we know about this diagram?
    • The slopes of the blue lines (the paths of the light rays) are of magnitude one (because the speed of light is the same in this reference frame, too).
    • AO = OD.
    • The angle between t and the t' axes, which is simply a function of the velocity (you should be able to calculate this angle with respect to the velocity by now -- remember, it's just a distance-time graph).
    How would you calculate angle BOC?

    Note: this is simply a geometric problem at this point. I encourage you to try it out on your own.

    SPOILERS AHEAD.

    Well, if you look at the diagram hard enough, you might have noticed that ABD is a right-angled triangle with right angle B. Additionally, AO = OD. Well, any triangle can be inscribed in a circle, and in the case of a right-angled triangle, AD becomes the diameter. Thus AO = OD = OB is the radius.

    Then ODB is an isoceles triangle, and angle ODB = angle OBD. Meanwhile angles OCE and OEC are both pi/4, thus OED and OCB are equal as well (they are 3pi/4). Triangle OBC is thus congruent to triangle ODE (since two angles and a side are equal), and angle BOC = angle DOE. Since angle DOE = angle FOA, this means angle BOC = angle FOA.

    This conclusion is tremendously significant: the x-axis is rotated by precisely the same angle as the t-axis is, towards each other. This creates a brilliant symmetry between space and time in relativity. We are also very close to our final expression for the Lorentz transformation.

    By the way, this proof also illustrates the beauty of natural units: by choosing a system of units such that the slope of light's path is one, the angle ABD became a right angle, and we were able to exploit the property of an angle subtended by the diameter being 90 degrees.

    Think about what happens in a reference frame moving at the speed of light. The axes then coincide.

    To be fair, we already expected this. Since the speed of light is constant, the null vectors (vectors pointing along the path of light in spacetime, i.e. along the diagonals) are eigenvectors of the Lorentz transformation. The only way for this to be true when you have a linear transformation is for the x-axis to be tilted inwards by the same angle.

    4. Scale on the transformed axes

    We now know how the axes transform, and must determine the scale on each axis.

    First of all, we may assume that the Lorentz transformation is linear. Why? Well, a linear transformation is one which ensures that all straight lines remain straight lines, and the origin remains fixed. The origin remains fixed in the Lorentz transformation by definition (since the observer is at the same spot -- translations are not considered), and lines must not turn into curves, since curves represent non-inertial reference frames and an inertial reference frame must be seen as inertial in all reference frames.

    So how do our unit vectors look like? Well, we know the image of the x-unit vector is a multiple of the vector $\left[ {\begin{array}{*{20}{c}}
      1 \\
      v
    \end{array}} \right]$, where we're of course using natural units. The t-unit vector, meanwhile, is a multiple of the vector $\left[ {\begin{array}{*{20}{c}}
      v \\
      1
    \end{array}} \right]$.

    So the transformation matrix, which is itself a function of $v$, takes the form

    $$L(v)=\left[ {\begin{array}{*{20}{c}}
      \alpha &{\beta v} \\
      {\alpha v}&\beta
    \end{array}} \right]$$
    For some constants $\alpha$ and $\beta$. Note that this is the transformation matrix which maps the original co-ordinate system to the new one -- the actual Lorentz transformation is a co-ordinate transformation, and thus the inverse of this matrix.

    How would we find the values of $\alpha$ and $\beta$? Well, one way would be to consider the product $L(v)L(-v)$. Since you are simply boosting by a velocity of $v$ then boosting back by $-v$, this product must equal the identity matrix $I$. This is "Einstein's principle of velocity reciprocity". We impose this condition:

    $$\begin{gathered}
      \left[ {\begin{array}{*{20}{c}}
      1&0 \\
      0&1
    \end{array}} \right] = \left[ {\begin{array}{*{20}{c}}
      \alpha &{\beta v} \\
      {\alpha v}&\beta
    \end{array}} \right]\left[ {\begin{array}{*{20}{c}}
      \alpha &{ - \beta v} \\
      { - \alpha v}&\beta
    \end{array}} \right] = \left[ {\begin{array}{*{20}{c}}
      {{\alpha ^2} - \alpha \beta {v^2}}&{{\beta ^2}v - \alpha \beta v} \\
      {{\alpha ^2}v - \alpha \beta v}&{{\beta ^2} - \alpha \beta {v^2}}
    \end{array}} \right] \hfill \\
      {\alpha ^2}v - \alpha \beta v = 0 = {\beta ^2}v - \alpha \beta v \Rightarrow {\alpha ^2} = \alpha \beta  = {\beta ^2} \Rightarrow \alpha  = \beta  \hfill \\
      {\alpha ^2} - \alpha \beta {v^2} = 1 = {\beta ^2} - \alpha \beta {v^2} \Rightarrow {\alpha ^2} = 1 + \alpha \beta {v^2} = {\beta ^2} \Rightarrow {\alpha ^2} = 1 + {\alpha ^2}{v^2} \hfill \\
       \Rightarrow \alpha  = \beta  = \frac{1}{{\sqrt {1 - {v^2}} }} \hfill \\
    \end{gathered} $$
    We call this coefficient the "Lorentz factor", and denote it by $\gamma$. From linear algebra, we know then that the co-ordinates of any point can then be transformed into the reference frame $O'$ as follows:

    $$\begin{gathered}
      \left[ {\begin{array}{*{20}{c}}
      {x'} \\
      {t'}
    \end{array}} \right] = {L^{ - 1}}\left[ {\begin{array}{*{20}{c}}
      x \\
      t
    \end{array}} \right] = {\gamma ^{ - 1}}{\left[ {\begin{array}{*{20}{c}}
      1&v \\
      v&1
    \end{array}} \right]^{ - 1}}\left[ {\begin{array}{*{20}{c}}
      x \\
      t
    \end{array}} \right] \\
       = \sqrt {1 - {v^2}}  \cdot \frac{1}{{1 - {v^2}}}\left[ {\begin{array}{*{20}{c}}
      1&{ - v} \\
      { - v}&1
    \end{array}} \right]\left[ {\begin{array}{*{20}{c}}
      x \\
      t
    \end{array}} \right] \\
       = \frac{1}{{\sqrt {1 - {v^2}} }}\left[ {\begin{array}{*{20}{c}}
      1&{ - v} \\
      { - v}&1
    \end{array}} \right]\left[ {\begin{array}{*{20}{c}}
      x \\
      t
    \end{array}} \right] \\
       = \gamma \left[ {\begin{array}{*{20}{c}}
      1&{ - v} \\
      { - v}&1
    \end{array}} \right]\left[ {\begin{array}{*{20}{c}}
      x \\
      t
    \end{array}} \right] \\
    \end{gathered} $$
    We may write this without matrices as:

    $$\begin{gathered}
      x' = \gamma \left( {x - vt} \right) \\
      t' = \gamma \left( {t - vx} \right) \\
    \end{gathered} $$
    Which updates the Galilean transformation discussed previously, which was $x'=x-vt,\ \ t' = t$.

    How does this look without natural units? Well, first of all,

    $$\gamma  = \frac{1}{{\sqrt {1 - \frac{{{v^2}}}{{{c^2}}}} }}$$
    And

    $$\begin{gathered}
      x' = \gamma \left( {x - \frac{v}{c}ct} \right) \hfill \\
      ct' = \gamma \left( {ct - \frac{v}{c}x} \right) \hfill \\
    \end{gathered} $$
    You can see why we prefer to set $c=1$, but this is also instructive -- it presents a symmetry between $x$ and $ct$, and $v/c$ is the important "ratio factor" between these dimensions.

    The transformation we've been calling "Lorentz transformations" are actually Lorentz boosts. Lorentz transformations are a broader set of transformations which includes boosts as well as spatial rotations -- essentially all linear transformations under which special relativity is invariant. An even broader set, called the Poincaire transformations, is the set of all affine transformations under which special relativity is invariant, i.e. it includes translations. As we will learn, General Relativity is only invariant under Lorentz transformations, not translations.

    We imposed the condition $L(v)L(-v)=L(0)$. Do you think one may impose, in general, that $L(v)L(w)=L(v+w)$? Why or why not? ... Answer is "no", because the velocity addition formula is not, in general, $v+w$.

    5. Zero orthogonal action of the Lorentz transformation

    Something we haven't considered so far is how a Lorentz boost treats spatial directions orthogonal to a Lorentz boost. We've been considering a Lorentz boost in the x-direction -- what happens to the y- and z- coordinates under this boost?

    Well, turns out, the answer is nothing. The explanation for this is pretty simple: attach a paintbrush to a train and let it paint the walls of the tunnel as the train drives through. Now send another train in the opposite direction and attach a paintbrush to it at the same height. Neither paintbrush can be "higher" than the other -- the paintbrushes must overlap in all reference frames.