Chapter 2: Foundations of Quantum Mechanics

Chapter 2: Physical Foundations of Quantum Mechanics

In Chapter 1, we built a complete mathematical toolbox—from complex numbers and vector spaces, to linear operators, eigenvalue decomposition, and tensor products. The task of this chapter is to “physically instantiate” these abstract mathematical structures: we will see that quantum mechanics is not a mysterious theory conjured from thin air, but rather the inevitable outcome of classical physics breaking down at the microscopic scale; and those seemingly abstract mathematical postulates are, in fact, the most precise description of physical reality.


2.1 From Classical to Quantum: Motivation & History

The Triumphs and Crises of Classical Physics

Toward the end of the nineteenth century, classical physics appeared nearly complete. Newtonian mechanics precisely described the trajectories of macroscopic objects, Maxwell’s equations elegantly unified electricity, magnetism, and light, and thermodynamics together with statistical physics successfully explained macroscopic thermal phenomena and phase transitions. From planetary orbits to steam engine efficiency, from the prediction of electromagnetic waves to spectroscopic analysis, classical theory achieved remarkable success in nearly every known domain of physics. Physicists widely embraced an optimistic outlook, believing the main edifice of physics had been completed, with only “two clouds”—black-body radiation and the ether drift—remaining to be clarified. In his famous 1900 lecture, Lord Kelvin even declared that the edifice of physics was essentially built, and that future physicists need only perform minor patchwork.

History, however, proved this optimism premature. It was precisely these two seemingly insignificant clouds that stirred the most magnificent revolutionary storm of twentieth-century physics, forever altering humanity’s understanding of the natural world.

Black-body Radiation and the Ultraviolet Catastrophe

Black-body radiation refers to the electromagnetic radiation emitted by an object in thermal equilibrium. The Rayleigh–Jeans law, derived from classical electromagnetic theory and statistical physics, predicted that the radiation energy density would grow unboundedly with the square of frequency ν\nu:

u(ν,T)=8πν2c3kBTu(\nu, T) = \frac{8\pi\nu^2}{c^3} k_B T

where kBk_B is the Boltzmann constant and cc is the speed of light. This result implies that any hot object should radiate infinite energy in the ultraviolet and higher frequency regions—clearly in gross contradiction with experimental observation. This is the famous Ultraviolet Catastrophe.

In 1900, Max Planck, in order to fit the experimental curve, proposed a bold hypothesis: the energy of electromagnetic radiation is not continuous, but is emitted and absorbed in discrete “energy quanta,” each quantum having magnitude:

E=hν=ωE = h\nu = \hbar\omega

where h6.626×1034Jsh \approx 6.626 \times 10^{-34} \, \text{J}\cdot\text{s} is Planck’s constant, and =h/2π\hbar = h/2\pi is the reduced Planck constant. Based on this hypothesis, Planck derived a radiation formula in perfect agreement with experiment:

u(ν,T)=8πhν3c31ehν/kBT1u(\nu, T) = \frac{8\pi h\nu^3}{c^3} \frac{1}{e^{h\nu/k_B T} - 1}

This discovery marks the birth of quantum theory—energy, in the microscopic world, is no longer a continuously varying quantity, but possesses “granularity.”

The Photoelectric Effect and the Particle Nature of Light

In 1905, Albert Einstein further developed the quantum concept to explain the photoelectric effect. Classical wave theory predicted: when light shines on a metal surface, no matter how weak the intensity, as long as the illumination time is sufficiently long, electrons will eventually accumulate enough energy to escape. But experiments showed: electrons are ejected only when the light frequency ν\nu exceeds a threshold ν0\nu_0; moreover, the electrons’ kinetic energy depends only on the light frequency, not on the light intensity.

Einstein proposed: light itself consists of particles (later called photons), each photon having energy E=hνE = h\nu. When a photon strikes an electron, the energy is transferred all at once, and the electron acquires kinetic energy:

Ek=hνWE_k = h\nu - W

where WW is the work function of the metal. This explanation matched experiment perfectly and earned Einstein the 1921 Nobel Prize in Physics. Light—that purely wave phenomenon of classical physics—exhibits particle-like behavior at the microscopic scale.

The Hydrogen Spectrum and the Bohr Model

Another phenomenon that classical theory could not explain was the spectrum of the hydrogen atom. Experiments observed that hydrogen atoms emit light only at specific wavelengths, forming discrete spectral lines. Classical electrodynamics predicted that an electron orbiting a nucleus would continuously radiate energy, eventually spiraling into the nucleus—atoms could not be stable.

In 1913, Niels Bohr proposed a semi-classical quantization model: electrons can only move in specific orbits, and the orbital angular momentum must be an integer multiple of \hbar:

L=n,n=1,2,3,L = n\hbar, \quad n = 1, 2, 3, \ldots

When an electron “jumps” between different orbits, it absorbs or emits a photon whose energy equals the difference between the two orbital energies:

ΔE=EmEn=hν\Delta E = E_m - E_n = h\nu

The Bohr model successfully explained the experimental regularities of the hydrogen spectrum, but the physical meaning of its theoretical foundation remained unclear—it was more like a “cobbled-together” empirical rule than a result derived from first principles.

Matrix Mechanics and Wave Mechanics

The genuine theoretical breakthrough came in 1925–1926. Werner Heisenberg developed matrix mechanics in 1925, representing physical quantities as matrices and mechanical relations by matrix equations. Almost simultaneously, Erwin Schrödinger proposed wave mechanics in 1926, describing the electron as a “matter wave” surrounding the atomic nucleus.

At first, these two formalisms appeared utterly distinct. But soon, rigorous mathematical proofs showed that matrix mechanics and wave mechanics are completely equivalent—they are simply the same physical theory expressed in different representations. The unifying framework behind them is precisely the Hilbert space and linear operator theory we studied in Section 1.2. The wave function ψ\psi is a vector in Hilbert space, and the operators corresponding to physical observables are linear operators on that space.

Wave–Particle Duality: The Quantum Puzzle of the Double-Slit Experiment

No thought experiment more vividly captures the counterintuitive character of quantum mechanics than the double-slit experiment.

Imagine a source that emits electrons (or photons) one at a time toward an opaque screen with two narrow slits. Behind the screen, place a detection screen to record where each electron lands. The experiment proceeds in three stages:

  1. Only slit A open: Electrons form a single-slit diffraction pattern on the detection screen, brightest at the center and tapering off to the sides. The probability distribution of landing positions is PA(x)=ψA(x)2P_A(x) = |\psi_A(x)|^2.

  2. Only slit B open: Similarly, another single-slit diffraction pattern forms, with probability distribution PB(x)=ψB(x)2P_B(x) = |\psi_B(x)|^2.

  3. Both slits open: According to classical probability theory (recall Section 1.7), if electrons were classical particles, we would expect the total probability to be the sum of the individual-slit probabilities: Pclassical(x)=PA(x)+PB(x)=ψA(x)2+ψB(x)2P_{\text{classical}}(x) = P_A(x) + P_B(x) = |\psi_A(x)|^2 + |\psi_B(x)|^2

    Yet what is actually observed is an interference pattern—alternating bright and dark fringes! The total probability is: Pquantum(x)=ψA(x)+ψB(x)2=ψA(x)2+ψB(x)2+2Re[ψA(x)ψB(x)]P_{\text{quantum}}(x) = |\psi_A(x) + \psi_B(x)|^2 = |\psi_A(x)|^2 + |\psi_B(x)|^2 + 2\text{Re}[\psi_A^*(x)\psi_B(x)]

The key is the cross term 2Re[ψAψB]2\text{Re}[\psi_A^*\psi_B]. This term can be positive or negative, causing the probability of an electron arriving at certain positions to increase (constructive interference) and at other positions to decrease (destructive interference). This is a direct manifestation of the complex-number arithmetic we learned in Section 1.1: the squared modulus of a sum of complex numbers is not equal to the sum of the squared moduli, a+b2a2+b2|a + b|^2 \neq |a|^2 + |b|^2.

Even more astonishing: even if the electron emission rate is reduced to the extreme—ensuring that only a single electron traverses the apparatus at any time—after accumulating detections over a sufficiently long period, the interference fringes still appear on the detection screen! This means that a single electron simultaneously “passes through” both slits and interferes with itself. This is a phenomenon that classical physics is utterly incapable of explaining.

If detectors are placed behind the double slit in an attempt to determine which slit the electron actually passed through, the interference pattern immediately vanishes, and the detection screen reverts to the classical probability sum PA+PBP_A + P_B. The very act of measurement changes the state of the system—this is the measurement problem of quantum mechanics, and the physical root of the central property in quantum computing that “reading a quantum state destroys the superposition.”

Recap: “Classical Probability” vs “Quantum Probability”

Let us precisely contrast classical and quantum probability in mathematical language. In classical probability theory (Section 1.7), the probabilities of mutually exclusive events are additive: if events AA and BB are mutually exclusive, then P(A or B)=P(A)+P(B)P(A \text{ or } B) = P(A) + P(B). This corresponds to the direct addition of real-valued probabilities.

In quantum mechanics, probability is given by the squared modulus of a probability amplitude—a complex number ψ\psi. Amplitudes from different paths are first added, and only then is the squared modulus taken:

P=ψA+ψB2=ψA2+ψB2+2Re(ψAψB)P = |\psi_A + \psi_B|^2 = |\psi_A|^2 + |\psi_B|^2 + 2\text{Re}(\psi_A^*\psi_B)

The extra cross term 2Re(ψAψB)2\text{Re}(\psi_A^*\psi_B) is the source of quantum interference. This is the fundamentally new physical effect produced by the combination of complex-number arithmetic (Section 1.1) and probability theory (Section 1.7). In quantum computing, this interference effect is the essential mechanism behind quantum algorithmic speedup—through carefully designed quantum circuits, the amplitudes of “correct answers” are constructively enhanced while those of “wrong answers” are destructively suppressed.

Summary: Classical physics suffered systematic failure on the problems of black-body radiation, the photoelectric effect, and atomic spectra. The pioneering work of Planck, Einstein, and Bohr revealed the discreteness and quantization of the microscopic world. Heisenberg’s matrix mechanics and Schrödinger’s wave mechanics established the mathematical framework of quantum mechanics from two perspectives, later proven to be different representations of the same theory on a Hilbert space. The double-slit experiment is the fundamental watershed between quantum mechanics and classical physics: quantum probability obeys the complex-amplitude superposition rule ψA+ψB2ψA2+ψB2|\psi_A + \psi_B|^2 \neq |\psi_A|^2 + |\psi_B|^2; a single particle can exhibit interference behavior; and measurement irreversibly changes the quantum state.

Connection to Quantum Computing: The core resources of quantum computing are precisely the quantum superposition and quantum interference revealed by the double-slit experiment. A quantum bit (qubit) exists in a superposition state ψ=α0+β1|\psi\rangle = \alpha|0\rangle + \beta|1\rangle of 0|0\rangle and 1|1\rangle, analogous to an electron passing through both slits simultaneously. Quantum algorithms manipulate these complex amplitudes through unitary operations, employing constructive interference to amplify the probability of the correct result and destructive interference to suppress that of incorrect results. Without quantum interference, there would be no quantum computational speedup.


2.2 Quantum Mechanical Postulates

Following the historical review of Section 2.1, we now systematically establish the axiomatic framework of quantum mechanics. Just as Euclidean geometry rests upon five postulates, quantum mechanics can be derived in its entirety from five fundamental postulates. More importantly, each postulate directly corresponds to a mathematical structure we painstakingly constructed in Chapter 1—this is the crucial step of “grounding” mathematics in physics.

Postulate 1: State Space Postulate

The complete set of possible states of an isolated quantum system is described by unit vectors ψ|\psi\rangle in a complex Hilbert space H\mathcal{H}. This vector is called the state vector or wave function of the system.

Recall from Section 1.2: a Hilbert space is a complete complex vector space equipped with an inner product. In quantum mechanics, this space is called the state space. The state of a system is not given by definite numerical values (such as position and momentum in classical mechanics), but is fully described by a vector.

The requirement of “unit vector” means the state vector must satisfy the normalization condition (recall inner products and norms from Section 1.4):

ψψ=ψ2=1\langle\psi|\psi\rangle = ||\,|\psi\rangle\,||^2 = 1

The physical meaning of this condition is: the total probability of the system being in some state is unity. In Section 1.4, we learned that the inner product ϕψ\langle\phi|\psi\rangle measures the “degree of overlap” between two vectors; physically, ϕψ2|\langle\phi|\psi\rangle|^2 gives the probability that a system in state ψ|\psi\rangle is measured to be in state ϕ|\phi\rangle.

Example: The state space of a two-dimensional quantum system (spin-1/2 system) is C2\mathbb{C}^2. A normalized state can be written as:

ψ=α0+β1|\psi\rangle = \alpha|0\rangle + \beta|1\rangle

where α,βC\alpha, \beta \in \mathbb{C} and α2+β2=1|\alpha|^2 + |\beta|^2 = 1 (recall complex modulus arithmetic from Section 1.1). Here 0|0\rangle and 1|1\rangle form an orthonormal basis.

Postulate 2: Evolution Postulate

The state evolution of a closed (unmeasured and unperturbed) quantum system is described by a unitary operator UU. If the system is in state ψ(t0)|\psi(t_0)\rangle at time t0t_0, then at time tt the state is:

The definition of a unitary operator (Section 1.3) requires $U^\dagger U = U U^\dagger = I$, i.e., the Hermitian conjugate of $U$ equals its inverse. This mathematical property carries profound physical meaning: 1. **Norm preservation**: $\langle\psi(t)|\psi(t)\rangle = \langle\psi(0)|U^\dagger U|\psi(0)\rangle = \langle\psi(0)|\psi(0)\rangle = 1$. The state vector remains normalized at all times; total probability is conserved. 2. **Reversibility**: Unitary evolution is reversible—given the final state, the initial state can be uniquely determined. This stands in sharp contrast to the irreversibility of measurement. In Section 1.3, we studied the Pauli matrices $X, Y, Z$ and general unitary matrices. All logic gates in quantum computing—Hadamard gate, phase gate, CNOT gate, etc.—are unitary operators. The essence of a quantum algorithm is to design a unitary operator $U$ that transforms the initial state $|\psi_0\rangle$ to a target state $|\psi_f\rangle = U|\psi_0\rangle$, such that measurement yields the correct answer with high probability. **Example**: Applying the Hadamard gate (the matrix $H = \frac{1}{\sqrt{2}}\begin{pmatrix}1 & 1 \\ 1 & -1\end{pmatrix}$ introduced in Section 1.3) to the state $|0\rangle$: $$H|0\rangle = \frac{1}{\sqrt{2}}\begin{pmatrix}1 & 1 \\ 1 & -1\end{pmatrix}\begin{pmatrix}1 \\ 0\end{pmatrix} = \frac{1}{\sqrt{2}}\begin{pmatrix}1 \\ 1\end{pmatrix} = \frac{|0\rangle + |1\rangle}{\sqrt{2}} \equiv |+\rangle$$ The new state $|+\rangle$ is still normalized: $\langle +|+\rangle = \frac{1}{2}(1 + 1) = 1$. **Postulate 3: Observable Postulate** > Every physical **observable** corresponds to a **Hermitian operator** $A$. The **eigenvalues** $a_i$ of the operator give all possible measurement outcomes of that observable. The corresponding **eigenvectors** $|a_i\rangle$ form a complete orthonormal basis of the state space. This is a direct application of the **spectral decomposition theorem** from Section 1.5! The three key mathematical properties of Hermitian operators—eigenvalues are real, eigenvectors corresponding to distinct eigenvalues are orthogonal, and eigenvectors constitute a complete basis—correspond respectively to three physical facts: 1. **Eigenvalues are real**: Measurement results must be real numbers (we can read "3.5 cm" from a ruler, but not "2+i cm"). 2. **Eigenvectors are orthogonal**: States corresponding to different measurement outcomes are mutually distinguishable. 3. **Completeness**: Any state can be expressed as a superposition of these eigenstates. According to the spectral decomposition theorem, any observable can be written as: $$A = \sum_i a_i \, |a_i\rangle\langle a_i|$$ where $|a_i\rangle\langle a_i|$ is the **projection operator** onto the eigenvector $|a_i\rangle$—an object we discussed in detail in Section 1.4. **Example**: The $z$-direction spin operator of a spin-1/2 system is $S_z = \frac{\hbar}{2}Z$, where $Z = \begin{pmatrix}1 & 0 \\ 0 & -1\end{pmatrix}$ is the Pauli $Z$ matrix (Section 1.3). The eigenvalues of $Z$ are $\pm 1$, with corresponding eigenvectors $|0\rangle = \begin{pmatrix}1 \\ 0\end{pmatrix}$ and $|1\rangle = \begin{pmatrix}0 \\ 1\end{pmatrix}$. Hence the possible measurement outcomes of $S_z$ are $\pm \hbar/2$. **Postulate 4: Measurement Postulate / Born Rule** > When measuring an observable $A$ on a system in state $|\psi\rangle$: > 1. The probability of obtaining result $a_i$ is $P(a_i) = |\langle a_i|\psi\rangle|^2 = \langle\psi|a_i\rangle\langle a_i|\psi\rangle = \langle\psi|P_i|\psi\rangle$, where $P_i = |a_i\rangle\langle a_i|$. > 2. If the measurement outcome is $a_i$, the system immediately collapses into the corresponding eigenstate $|a_i\rangle$. This is the most iconic postulate of quantum mechanics. The **Born rule** provides the mapping from complex amplitudes to real probabilities: first compute the **inner product** $\langle a_i|\psi\rangle$ (a complex number) between the state $|\psi\rangle$ and the eigenstate $|a_i\rangle$, then take the squared modulus to obtain the probability. Let us reformulate this using the projection operator language of Section 1.4: the probability of measurement outcome $a_i$ equals the squared length of the component of $|\psi\rangle$ after being projected onto the direction of $|a_i\rangle$. The projection operator $P_i = |a_i\rangle\langle a_i|$ maps any state $|\psi\rangle$ to its component along $|a_i\rangle$: $P_i|\psi\rangle = |a_i\rangle\langle a_i|\psi\rangle$. The squared norm of this component is: $$||P_i|\psi\rangle||^2 = \langle\psi|P_i^\dagger P_i|\psi\rangle = \langle\psi|P_i|\psi\rangle$$ (The last step uses the Hermiticity and idempotence of $P_i$: $P_i^\dagger = P_i$, $P_i^2 = P_i$.) The post-measurement collapse means: quantum measurement is not a passive "reading out" of information, but an active alteration of the system state. This irreversible process cannot be described by a unitary operator—it is one of the deepest mysteries of quantum mechanics. **Numerical Example (Born rule demonstrated on a spin-1/2 system)**: Let the system be in state $|+\rangle = \frac{|0\rangle + |1\rangle}{\sqrt{2}}$, and measure its $z$-direction spin using $S_z = \frac{\hbar}{2}Z$. The eigenstates of $S_z$ are $|0\rangle$ (eigenvalue $+\hbar/2$) and $|1\rangle$ (eigenvalue $-\hbar/2$). Probability of obtaining $+\hbar/2$: $$P(+\hbar/2) = |\langle 0|+\rangle|^2 = \left|\frac{\langle 0|0\rangle + \langle 0|1\rangle}{\sqrt{2}}\right|^2 = \left|\frac{1 + 0}{\sqrt{2}}\right|^2 = \frac{1}{2}$$ Probability of obtaining $-\hbar/2$: $$P(-\hbar/2) = |\langle 1|+\rangle|^2 = \left|\frac{\langle 1|0\rangle + \langle 1|1\rangle}{\sqrt{2}}\right|^2 = \frac{1}{2}$$ If $+\hbar/2$ is obtained, the system instantaneously collapses to state $|0\rangle$. **Postulate 5: Composite System Postulate** > For a composite system consisting of $N$ subsystems, the state space is the **tensor product** of the individual subsystem state spaces: > $$\mathcal{H} = \mathcal{H}_1 \otimes \mathcal{H}_2 \otimes \cdots \otimes \mathcal{H}_N

This is precisely the content of Section 1.6! If subsystem AA is in state ψA|\psi\rangle_A and subsystem BB in state ϕB|\phi\rangle_B, then the composite system is in state ψAϕB|\psi\rangle_A \otimes |\phi\rangle_B, often abbreviated as ψAϕB|\psi\rangle_A|\phi\rangle_B or ψϕ|\psi\phi\rangle.

The dimension of the composite system grows exponentially with the number of subsystems: if each subsystem has dimension dd, then the composite space of NN subsystems has dimension dNd^N. This is the mathematical root of the enormous power of quantum computing—the state space of NN qubits has dimension 2N2^N, requiring exponential resources for a classical computer to simulate.

However, not all composite states can be written as tensor products of subsystem states. States of the form

ΨAB=0A0B+1A1B2|\Psi\rangle_{AB} = \frac{|0\rangle_A|0\rangle_B + |1\rangle_A|1\rangle_B}{\sqrt{2}}

cannot be factorized into ψAϕB|\psi\rangle_A \otimes |\phi\rangle_B. Such states are called entangled states, and they are the central resource for quantum communication and quantum cryptography. We will further understand the nature of entanglement through density matrices and reduced density matrices in Section 2.6.

Example: The standard computational basis for a two-qubit system is 00,01,10,11|00\rangle, |01\rangle, |10\rangle, |11\rangle. A general two-qubit state is:

Ψ=α00+β01+γ10+δ11|\Psi\rangle = \alpha|00\rangle + \beta|01\rangle + \gamma|10\rangle + \delta|11\rangle

where α2+β2+γ2+δ2=1|\alpha|^2 + |\beta|^2 + |\gamma|^2 + |\delta|^2 = 1 (normalization condition).

Summary: The five postulates of quantum mechanics constitute a complete theoretical framework: (1) The state of a system is described by a unit vector in Hilbert space—corresponding to the vector spaces of Section 1.2 and the normalization of Section 1.4; (2) The evolution of a closed system is driven by a unitary operator—corresponding to the unitary matrices of Section 1.3; (3) Observables correspond to Hermitian operators, whose eigenvalues are the possible measurement outcomes—corresponding to the eigenvalue problem of Section 1.5; (4) Measurement probabilities are given by the squared modulus of inner products, and measurement causes state collapse—corresponding to the projection operators of Section 1.4 and the probability theory of Section 1.7; (5) The state space of a composite system is a tensor product—corresponding to the tensor products of Section 1.6. These five postulates weave all the mathematical tools of Chapter 1 into a unified physical theory.

Connection to Quantum Computing: Every operation of a quantum computer directly corresponds to these five postulates. Qubit initialization corresponds to Postulate 1; quantum gate operations correspond to the unitary evolution of Postulate 2; computational basis measurements correspond to Postulates 3 and 4; multi-qubit systems correspond to the tensor product structure of Postulate 5. The art of quantum algorithm design lies in skillfully arranging quantum interference (Postulate 2) while maintaining unitary evolution, so that upon final measurement (Postulate 4), the correct result appears with high probability. Entangled states (a special product of Postulate 5) are the core resource for quantum communication, quantum key distribution, and quantum error correction.


2.3 Wave Functions and the Schrödinger Equation

Within the postulate framework of Section 2.2, the state vector ψ|\psi\rangle is an abstract vector in Hilbert space. In this section, we “project” it onto a concrete representation—the position representation—introducing the wave function ψ(x)\psi(x) and establishing its evolution equation: the Schrödinger equation.

The Wave Function as the Position Representation of a State Vector

Consider a particle of mass mm moving in one-dimensional space. In the position representation, we introduce position eigenstates x|x\rangle, satisfying the orthonormality relation xx=δ(xx)\langle x|x'\rangle = \delta(x - x') (the Dirac delta function) and the completeness relation +xxdx=I\int_{-\infty}^{+\infty} |x\rangle\langle x| dx = I.

An arbitrary state vector ψ|\psi\rangle can be expanded in the position basis:

ψ=+ψ(x)xdx|\psi\rangle = \int_{-\infty}^{+\infty} \psi(x) \, |x\rangle \, dx

where the expansion coefficient ψ(x)=xψ\psi(x) = \langle x|\psi\rangle is precisely the wave function. It is a complex-valued function, mapping each position xx to a complex number ψ(x)C\psi(x) \in \mathbb{C}. In Section 1.2, we discussed finite-dimensional vector spaces; the wave function is a vector in the infinite-dimensional (continuous-index) Hilbert space L2(R)L^2(\mathbb{R})—i.e., the space of square-integrable functions satisfying ψ(x)2dx<\int |\psi(x)|^2 dx < \infty.

Probability Interpretation

The core physical interpretation of the wave function was proposed by Max Born: ψ(x)2|\psi(x)|^2 represents the probability density of finding the particle at position xx. Specifically, the probability of finding the particle in the interval [x,x+dx][x, x+dx] is:

P(xXx+dx)=ψ(x,t)2dxP(x \leq X \leq x+dx) = |\psi(x,t)|^2 \, dx

The normalization condition ψψ=1\langle\psi|\psi\rangle = 1 becomes, in the position representation:

+ψ(x)2dx=1\int_{-\infty}^{+\infty} |\psi(x)|^2 \, dx = 1

This means the particle must exist somewhere in space—the total probability is unity.

The phase of the wave function also carries physical information. Although ψ(x)2|\psi(x)|^2 depends only on the modulus, the relative phase between two superposed wave functions determines the interference pattern. This is the physical manifestation of the complex phase concept from Section 1.1: ψ(x)=ψ(x)eiϕ(x)\psi(x) = |\psi(x)|e^{i\phi(x)}, where a global phase eiθe^{i\theta} does not affect physical observations, but the relative phase is crucial.

The Schrödinger Equation

How does the wave function evolve in time? In 1926, Schrödinger proposed the equation describing this evolution:

itψ(t)=Hψ(t)i\hbar \frac{\partial}{\partial t} |\psi(t)\rangle = H |\psi(t)\rangle

where HH is the Hamiltonian operator, representing the total energy of the system. In the position representation, for a single particle of mass mm under a potential V(x)V(x), the Hamiltonian is:

H=22m2x2+V(x)H = -\frac{\hbar^2}{2m}\frac{\partial^2}{\partial x^2} + V(x)

The first term is the kinetic energy operator (momentum operator p=ixp = -i\hbar\frac{\partial}{\partial x} squared divided by 2m2m); the second term is the potential energy.

The Schrödinger equation is the fundamental dynamical equation of quantum mechanics, analogous to Newton’s second law F=maF = ma in classical mechanics. But it is not “derived” from any more fundamental principle—it is one of the fundamental assumptions of quantum mechanics, its correctness validated by experiment.

The Time-Independent Schrödinger Equation

When the Hamiltonian does not depend explicitly on time (H/t=0\partial H/\partial t = 0), we can use separation of variables. Let ψ(t)=eiEt/ψ|\psi(t)\rangle = e^{-iEt/\hbar}|\psi\rangle, and substitute into the Schrödinger equation:

it(eiEt/ψ)=H(eiEt/ψ)i\hbar \frac{\partial}{\partial t} \left(e^{-iEt/\hbar}|\psi\rangle\right) = H \left(e^{-iEt/\hbar}|\psi\rangle\right)

EeiEt/ψ=eiEt/HψE \, e^{-iEt/\hbar}|\psi\rangle = e^{-iEt/\hbar} H |\psi\rangle

Canceling the time-phase factor from both sides yields the time-independent Schrödinger equation:

Hψ=EψH |\psi\rangle = E |\psi\rangle

This is exactly the eigenvalue equation we studied in detail in Section 1.5! The eigenvalues EnE_n of the Hamiltonian operator HH are the energy eigenvalues of the system, and the corresponding eigenstates ψn|\psi_n\rangle are called energy eigenstates or stationary states. For a system in a stationary state, the probability density ψn(x,t)2=ψn(x)2|\psi_n(x,t)|^2 = |\psi_n(x)|^2 does not change with time—only the phase rotates.

This correspondence is profoundly deep: solving for the energy levels of a quantum system is, mathematically, solving the eigenvalue problem of the Hamiltonian operator. The techniques we learned in Section 1.5—spectral decomposition, diagonalization, etc.—become practical tools here for computing atomic energy levels and molecular vibrational frequencies.

The One-Dimensional Infinite Square Well

To demonstrate the process of solving the Schrödinger equation concretely, we consider a textbook example: the one-dimensional infinite square well.

The potential well is defined as follows:

V(x)={00<x<LotherwiseV(x) = \begin{cases} 0 & 0 < x < L \\ \infty & \text{otherwise} \end{cases}

The particle is completely confined within the interval 00 to LL. Outside the well, V=V = \infty, so the wave function must be zero (otherwise the energy would be infinite); inside the well, V=0V = 0, and the time-independent Schrödinger equation simplifies to:

22md2ψdx2=Eψ-\frac{\hbar^2}{2m}\frac{d^2\psi}{dx^2} = E\psi

or equivalently:

d2ψdx2=k2ψ,where k=2mE\frac{d^2\psi}{dx^2} = -k^2\psi, \quad \text{where } k = \frac{\sqrt{2mE}}{\hbar}

This is the standard simple harmonic equation, with general solution ψ(x)=Asin(kx)+Bcos(kx)\psi(x) = A\sin(kx) + B\cos(kx).

Boundary conditions: The wave function is continuous at the boundaries, so ψ(0)=0\psi(0) = 0 and ψ(L)=0\psi(L) = 0.

From ψ(0)=B=0\psi(0) = B = 0, we get B=0B = 0. From ψ(L)=Asin(kL)=0\psi(L) = A\sin(kL) = 0, we require kL=nπkL = n\pi, where n=1,2,3,n = 1, 2, 3, \ldots (n=0n = 0 gives the trivial solution ψ=0\psi = 0; negative integers give the same physical state up to an overall minus sign).

Thus, the wave number is quantized:

kn=nπLk_n = \frac{n\pi}{L}

and the corresponding energy eigenvalues are:

En=2kn22m=n2π222mL2,n=1,2,3,E_n = \frac{\hbar^2 k_n^2}{2m} = \frac{n^2\pi^2\hbar^2}{2mL^2}, \quad n = 1, 2, 3, \ldots

The energy is discrete—only specific allowed values! This is a hallmark feature of quantum mechanics, with no counterpart in classical physics. The integer nn is called the quantum number. The ground state (n=1n = 1) energy E1=π222mL2E_1 = \frac{\pi^2\hbar^2}{2mL^2} is called the zero-point energy—even if the system is cooled to absolute zero, the particle retains this minimum energy, a direct consequence of the uncertainty principle.

The normalized wave functions are:

ψn(x)=2Lsin(nπxL)\psi_n(x) = \sqrt{\frac{2}{L}} \sin\left(\frac{n\pi x}{L}\right)

Verification of normalization:

0Lψn(x)2dx=2L0Lsin2(nπxL)dx=2LL2=1\int_0^L |\psi_n(x)|^2 dx = \frac{2}{L}\int_0^L \sin^2\left(\frac{n\pi x}{L}\right) dx = \frac{2}{L} \cdot \frac{L}{2} = 1

Different energy eigenstates are mutually orthogonal:

0Lψm(x)ψn(x)dx=δmn\int_0^L \psi_m^*(x)\psi_n(x) dx = \delta_{mn}

This is the embodiment of the orthogonality concept from Section 1.4 in the space of continuous functions.

The wave function ψn(x)\psi_n(x) has n1n-1 nodes (points where the wave function is zero) inside the well, with wavelength λn=2L/n\lambda_n = 2L/n. The magnitude of the particle’s momentum is pn=kn=nπ/Lp_n = \hbar k_n = n\pi\hbar/L, fully consistent with the de Broglie relation λ=h/p\lambda = h/p.

Summary: The wave function ψ(x)=xψ\psi(x) = \langle x|\psi\rangle is the representation of the state vector in the position representation; ψ(x)2|\psi(x)|^2 gives the position probability density. The Schrödinger equation itψ=Hψi\hbar\partial_t|\psi\rangle = H|\psi\rangle describes the time evolution of quantum states, analogous to Newton’s equation in classical mechanics. When the Hamiltonian is time-independent, the time-independent Schrödinger equation Hψ=EψH|\psi\rangle = E|\psi\rangle reduces to an eigenvalue problem—a direct application of the spectral theorem from Section 1.5. The solution of the one-dimensional infinite square well demonstrates the natural emergence of quantization: boundary conditions force the wave number (and hence the energy) to take discrete values; normalization and orthogonality correspond respectively to the norm and inner product concepts of Section 1.4.

Connection to Quantum Computing: Although quantum computing typically operates on finite-dimensional discrete systems (qubits), the framework of the Schrödinger equation still applies. Quantum gate operations correspond to the short-time unitary evolution generated by a Hamiltonian, U=eiHt/U = e^{-iHt/\hbar}. Continuous-variable quantum computing directly uses position or momentum eigenstates as information carriers. More importantly, the spirit of the eigenvalue problem in the Schrödinger equation runs throughout quantum algorithms: at the heart of the quantum phase estimation algorithm lies the encoding of the desired eigenvalue into the phase of a quantum state, followed by readout via the quantum Fourier transform (the quantum version of the DFT from Section 1.8). The discrete energy-level structure of the infinite square well also inspired the design of quantum dot qubits—electrons confined in nanoscale artificial potential wells, whose discrete energy levels naturally constitute the two levels of a qubit.


2.4 Two-Level Systems and Spin

After the continuous-system discussion of Section 2.3, we turn to the simplest yet most profound object of study in quantum mechanics: the two-level system. The state space of such a system is merely the two-dimensional complex vector space C2\mathbb{C}^2, yet it contains all the core features of quantum mechanics. More importantly, the two-level system is the physical prototype of the qubit—the fundamental unit of information in quantum computing.

The Two-Level System: The Simplest Quantum System

A two-level system has only two distinguishable energy states, conventionally denoted 0|0\rangle and 1|1\rangle (or ground|\text{ground}\rangle and excited|\text{excited}\rangle). Its state space is:

H=C2={α0+β1α,βC,α2+β2=1}\mathcal{H} = \mathbb{C}^2 = \{\alpha|0\rangle + \beta|1\rangle \,|\, \alpha, \beta \in \mathbb{C}, \, |\alpha|^2 + |\beta|^2 = 1\}

An arbitrary normalized state can be parameterized as:

ψ=cosθ20+eiϕsinθ21|\psi\rangle = \cos\frac{\theta}{2}\,|0\rangle + e^{i\phi}\sin\frac{\theta}{2}\,|1\rangle

where θ[0,π]\theta \in [0, \pi], ϕ[0,2π)\phi \in [0, 2\pi). The two real parameters correspond to a point on the Bloch sphere (see Section 2.5).

Two-level systems are ubiquitous: two specific energy levels of an atom, two polarization directions of a photon, two charge states of a superconducting qubit—and the protagonist of this section: spin-1/2.

The Stern–Gerlach Experiment

In 1922, Otto Stern and Walther Gerlach conducted a landmark experiment. They passed a beam of silver atoms (each silver atom contains one unpaired electron) through an inhomogeneous magnetic field. Classical physics predicted: since the orientation of the electron magnetic moment is continuously distributed, the atomic beam should spread into a continuous distribution on the detection screen.

Yet the experimental result was astonishing: the atomic beam split into two distinct beams, leaving two sharp spots on the detection screen!

This result demonstrated:

  1. The electron possesses an intrinsic angular momentum—spin—independent of any orbital motion.
  2. Spin can take only two discrete values along any spatial direction: ±/2\pm\hbar/2.

Spin is a purely quantum-mechanical concept, with no classical counterpart. The angular momentum of a classical particle can take any value and point in any continuous direction; but the electron’s spin has only two possibilities, regardless of which direction is measured.

Spin Operators and the Pauli Matrices

The mathematical description of spin directly corresponds to the Pauli matrices we studied in Section 1.3. The spin operators for the three spatial directions are:

Sx=2X=2(0110),Sy=2Y=2(0ii0),Sz=2Z=2(1001)S_x = \frac{\hbar}{2} X = \frac{\hbar}{2}\begin{pmatrix} 0 & 1 \\ 1 & 0 \end{pmatrix}, \quad S_y = \frac{\hbar}{2} Y = \frac{\hbar}{2}\begin{pmatrix} 0 & -i \\ i & 0 \end{pmatrix}, \quad S_z = \frac{\hbar}{2} Z = \frac{\hbar}{2}\begin{pmatrix} 1 & 0 \\ 0 & -1 \end{pmatrix}

These three operators satisfy the angular momentum commutation relations:

[Si,Sj]=iεijkSk[S_i, S_j] = i\hbar\varepsilon_{ijk} S_k

where εijk\varepsilon_{ijk} is the Levi-Civita symbol and [A,B]=ABBA[A, B] = AB - BA is the commutator.

Note that Sx,Sy,SzS_x, S_y, S_z are all Hermitian operators, consistent with the requirement of Postulate 3. Their eigenvalues are all ±/2\pm\hbar/2, in complete agreement with the Stern–Gerlach experimental observation.

Spin States

The eigenstates of the ZZ operator constitute the standard computational basis:

z=0=(10),z=1=(01)|\uparrow_z\rangle = |0\rangle = \begin{pmatrix} 1 \\ 0 \end{pmatrix}, \quad |\downarrow_z\rangle = |1\rangle = \begin{pmatrix} 0 \\ 1 \end{pmatrix}

The action of the Pauli matrices on these basis states:

  • Z0=+10=0Z|0\rangle = +1 \cdot |0\rangle = |0\rangle, Z1=11=1Z|1\rangle = -1 \cdot |1\rangle = -|1\rangle
  • X0=1X|0\rangle = |1\rangle, X1=0X|1\rangle = |0\rangle
  • Y0=i1Y|0\rangle = i|1\rangle, Y1=i0Y|1\rangle = -i|0\rangle

The XX operator flips the spin and is therefore called the bit-flip operator; the ZZ operator changes the sign of 1|1\rangle and is called the phase-flip operator.

The eigenstates of XX (the xx-direction spin eigenstates) are:

+=0+12,=012|+\rangle = \frac{|0\rangle + |1\rangle}{\sqrt{2}}, \quad |-\rangle = \frac{|0\rangle - |1\rangle}{\sqrt{2}}

with eigenvalues +1+1 and 1-1 (i.e., SxS_x eigenvalues +/2+\hbar/2 and /2-\hbar/2). Verification: X+=X0+X12=1+02=+X|+\rangle = \frac{X|0\rangle + X|1\rangle}{\sqrt{2}} = \frac{|1\rangle + |0\rangle}{\sqrt{2}} = |+\rangle.

The eigenstates of YY are:

+i=0+i12,i=0i12|+i\rangle = \frac{|0\rangle + i|1\rangle}{\sqrt{2}}, \quad |-i\rangle = \frac{|0\rangle - i|1\rangle}{\sqrt{2}}

Parameterization of an Arbitrary State

The most general pure state of a spin-1/2 system can be explicitly written as:

ψ=cosθ20+eiϕsinθ21|\psi\rangle = \cos\frac{\theta}{2}\,|0\rangle + e^{i\phi}\sin\frac{\theta}{2}\,|1\rangle

where θ[0,π]\theta \in [0, \pi], ϕ[0,2π)\phi \in [0, 2\pi). These two angles have clear physical meaning:

  • θ\theta determines the relative weights of 0|0\rangle and 1|1\rangle: when θ=0\theta = 0, ψ=0|\psi\rangle = |0\rangle; when θ=π\theta = \pi, ψ=1|\psi\rangle = |1\rangle; when θ=π/2\theta = \pi/2, the two weights are equal.
  • ϕ\phi is the relative phase between the two components. It is precisely this phase that causes the quantum interference in the double-slit experiment of Section 2.1.

This parameterization covers all normalized pure states. Note that a global phase eiγψe^{i\gamma}|\psi\rangle produces no observable physical effect, because probability depends only on ϕψ2|\langle\phi|\psi\rangle|^2, and the global phase cancels under the squared modulus. But the relative phase eiϕe^{i\phi} is physical—it determines the probability distribution of measurement outcomes.

Example: Measuring the zz-direction spin of the +i|+i\rangle state. The eigenstates 0|0\rangle and 1|1\rangle correspond to the SzS_z measurement.

P(Sz=+/2)=0+i2=122=12P(S_z = +\hbar/2) = |\langle 0|+i\rangle|^2 = \left|\frac{1}{\sqrt{2}}\right|^2 = \frac{1}{2} P(Sz=/2)=1+i2=i22=12P(S_z = -\hbar/2) = |\langle 1|+i\rangle|^2 = \left|\frac{-i}{\sqrt{2}}\right|^2 = \frac{1}{2}

If +/2+\hbar/2 is obtained, the state collapses to 0|0\rangle.

Summary: The two-level system is the simplest quantum system, with state space C2\mathbb{C}^2, and is the physical prototype of the qubit. The Stern–Gerlach experiment proved that the electron possesses intrinsic spin, and that spin along any direction takes only two values, ±/2\pm\hbar/2. The spin operators Sx,Sy,SzS_x, S_y, S_z correspond directly to the Pauli matrices X,Y,ZX, Y, Z of Section 1.3, whose eigenstates constitute the measurement bases for each direction. The computational basis 0|0\rangle and 1|1\rangle are eigenstates of ZZ; +|+\rangle and |-\rangle are eigenstates of XX; +i|+i\rangle and i|-i\rangle are eigenstates of YY. An arbitrary pure state is characterized by two parameters (θ,ϕ)(\theta, \phi); the relative phase ϕ\phi is the source of quantum interference.

Connection to Quantum Computing: A qubit is precisely an abstract two-level system. In physical implementation, it can be: two energy levels in a superconducting circuit (transmon qubit), two internal states of a trapped ion (ion-trap qubit), two polarization directions of a photon (photonic qubit), two spin states in a semiconductor quantum dot (spin qubit). All qubit operations—initialization, single-qubit gates, two-qubit gates, measurement—proceed within the five-postulate framework of Section 2.2. Single-qubit quantum gates (X,Y,Z,HX, Y, Z, H, etc.) correspond to unitary matrices generated by the Pauli matrices and their linear combinations; measurement corresponds to a spin measurement along some direction. The complete mathematical description of a spin-1/2 system is the entire theoretical framework for a single qubit.


2.5 The Bloch Sphere

In Section 2.4, we parameterized an arbitrary pure state of a two-level system as ψ=cos(θ/2)0+eiϕsin(θ/2)1|\psi\rangle = \cos(\theta/2)|0\rangle + e^{i\phi}\sin(\theta/2)|1\rangle. In this section, we map these parameters onto a three-dimensional sphere—the Bloch sphere—providing an intuitive geometric picture for the state space of a qubit.

Parameterization of the Bloch Sphere

Given the parameterization of a pure state:

ψ=cosθ20+eiϕsinθ21|\psi\rangle = \cos\frac{\theta}{2}\,|0\rangle + e^{i\phi}\sin\frac{\theta}{2}\,|1\rangle

where θ[0,π]\theta \in [0, \pi], ϕ[0,2π)\phi \in [0, 2\pi). We map it to a point (x,y,z)(x, y, z) in three-dimensional space:

x=sinθcosϕ,y=sinθsinϕ,z=cosθx = \sin\theta\cos\phi, \quad y = \sin\theta\sin\phi, \quad z = \cos\theta

Verification: x2+y2+z2=sin2θ(cos2ϕ+sin2ϕ)+cos2θ=sin2θ+cos2θ=1x^2 + y^2 + z^2 = \sin^2\theta(\cos^2\phi + \sin^2\phi) + \cos^2\theta = \sin^2\theta + \cos^2\theta = 1. Thus the point lies on the unit sphere. This mapping from quantum states to points on the unit sphere in three-dimensional space provides an intuitive and easily visualized geometric picture for the abstract two-level system.

The Bloch sphere provides a one-to-one mapping (states differing by a global phase eiγe^{i\gamma} map to the same point): each pure state of a qubit corresponds to a unique point on the sphere, and each point on the sphere corresponds to an equivalence class (differing by global phase) of pure states.

Positions of Standard States on the Sphere

Let us calculate the Bloch sphere coordinates of several important quantum states:

Quantum Stateθ\thetaϕ\phi(x,y,z)(x, y, z)Position on Sphere
$0\rangle$00arbitrary(0,0,+1)(0, 0, +1)
$1\rangle$π\piarbitrary(0,0,1)(0, 0, -1)
$+\rangle = \frac{0\rangle+1\rangle}{\sqrt{2}}$π/2\pi/2
$-\rangle = \frac{0\rangle-1\rangle}{\sqrt{2}}$π/2\pi/2
$+i\rangle = \frac{0\rangle+i1\rangle}{\sqrt{2}}$π/2\pi/2
$-i\rangle = \frac{0\rangle-i1\rangle}{\sqrt{2}}$π/2\pi/2

These states on the Bloch sphere form three mutually orthogonal axes: the zz-axis corresponds to eigenstates of the ZZ operator, the xx-axis to eigenstates of the XX operator, and the yy-axis to eigenstates of the YY operator.

Physical Meaning of the Phase ϕ\phi

The parameter ϕ\phi is the relative phase between the 0|0\rangle and 1|1\rangle components. On the Bloch sphere, ϕ\phi corresponds to the azimuthal angle about the zz-axis.

  • When ϕ=0\phi = 0, the state lies in the xx-zz plane (y=0y = 0), e.g., +|+\rangle.
  • When ϕ=π/2\phi = \pi/2, the state is shifted toward the +y+y direction, e.g., +i|+i\rangle.
  • Changing ϕ\phi is equivalent to rotating the quantum state about the zz-axis, corresponding to the ZZ gate or more general phase gates.

The relative phase is the core resource for quantum interference in quantum computing. For example, in the Deutsch algorithm, it is through carefully designed phase operations that the cases “the function is constant” and “the function is balanced” are mapped to different, distinguishable states.

The Bloch Vector

For a pure state ψ|\psi\rangle, define the Bloch vector n=(x,y,z)\vec{n} = (x, y, z), where:

x=ψXψ,y=ψYψ,z=ψZψx = \langle\psi|X|\psi\rangle, \quad y = \langle\psi|Y|\psi\rangle, \quad z = \langle\psi|Z|\psi\rangle

Verification: for ψ=cos(θ/2)0+eiϕsin(θ/2)1|\psi\rangle = \cos(\theta/2)|0\rangle + e^{i\phi}\sin(\theta/2)|1\rangle,

z=ψZψ=cos(θ/2)2eiϕsin(θ/2)2=cos2θ2sin2θ2=cosθz = \langle\psi|Z|\psi\rangle = |\cos(\theta/2)|^2 - |e^{i\phi}\sin(\theta/2)|^2 = \cos^2\frac{\theta}{2} - \sin^2\frac{\theta}{2} = \cos\theta

The other components xx and yy can be verified similarly. The three components of the Bloch vector are precisely the spin expectation values along the three Pauli directions (divided by /2\hbar/2).

Rotations on the Sphere Correspond to Unitary Operations

The geometric picture of the Bloch sphere makes unitary operations intuitive: any single-qubit unitary operation corresponds to a rotation on the Bloch sphere!

Specifically:

  • XX gate: X=eiπX/2X = e^{i\pi X/2} (up to a global phase), corresponds to a rotation of π\pi about the xx-axis. It takes 0|0\rangle (north pole) to 1|1\rangle (south pole), leaves +|+\rangle unchanged, and takes +i|+i\rangle to i|-i\rangle.

  • ZZ gate: Corresponds to a rotation of π\pi about the zz-axis. It leaves 0|0\rangle and 1|1\rangle unchanged (merely changing the sign of 1|1\rangle; the global phase does not affect the point on the Bloch sphere), takes +|+\rangle to |-\rangle, and takes +i|+i\rangle to i|-i\rangle.

  • YY gate: Corresponds to a rotation of π\pi about the yy-axis.

  • Hadamard gate HH: Corresponds to a rotation of π\pi about the axis n=(1,0,1)/2\vec{n} = (1, 0, 1)/\sqrt{2} (the diagonal direction in the xx-zz plane). It takes 0|0\rangle to +|+\rangle and 1|1\rangle to |-\rangle.

More generally, a unitary operation corresponding to a rotation by angle α\alpha about an arbitrary axis n=(nx,ny,nz)\vec{n} = (n_x, n_y, n_z) (a unit vector) is given by the Pauli rotation operator:

Rn(α)=eiα(nσ)/2=cosα2Iisinα2(nxX+nyY+nzZ)R_{\vec{n}}(\alpha) = e^{-i\alpha(\vec{n}\cdot\vec{\sigma})/2} = \cos\frac{\alpha}{2}I - i\sin\frac{\alpha}{2}(n_x X + n_y Y + n_z Z)

where σ=(X,Y,Z)\vec{\sigma} = (X, Y, Z) is the vector of Pauli matrices. This is a direct application of the matrix exponential from Section 1.3.

Pure States and Mixed States

All points on the surface of the Bloch sphere correspond to pure states. These pure states are quantum states that can be fully described by a single state vector, possessing maximal quantum coherence. However, points in the interior of the sphere, x2+y2+z2<1x^2 + y^2 + z^2 < 1, also have clear physical meaning—they correspond to mixed states. Mixed states cannot be described by a single state vector and must be characterized using the density matrix (Section 2.6). The center of the sphere (0,0,0)(0,0,0) corresponds to the maximally mixed state ρ=I/2\rho = I/2, representing complete ignorance about the system.

From the center to the surface, the “purity” of the state gradually increases. The square of the Bloch vector length r2=x2+y2+z2r^2 = x^2 + y^2 + z^2 is directly related to Tr(ρ2)\text{Tr}(\rho^2) of the density matrix: for a pure state, r=1r = 1 and Tr(ρ2)=1\text{Tr}(\rho^2) = 1; for a mixed state, r<1r < 1 and Tr(ρ2)<1\text{Tr}(\rho^2) < 1.

Summary: The Bloch sphere provides an elegant geometric representation for the pure states of a two-level system: each pure state ψ=cos(θ/2)0+eiϕsin(θ/2)1|\psi\rangle = \cos(\theta/2)|0\rangle + e^{i\phi}\sin(\theta/2)|1\rangle corresponds to a point (x,y,z)=(sinθcosϕ,sinθsinϕ,cosθ)(x, y, z) = (\sin\theta\cos\phi, \sin\theta\sin\phi, \cos\theta) on the unit sphere. The north and south poles correspond to 0|0\rangle and 1|1\rangle respectively; the xx-axis and yy-axis correspond to eigenstates of the XX and YY operators. The relative phase ϕ\phi corresponds to the azimuthal angle, determining the orientation of the quantum state in the equatorial plane. Quantum gate operations correspond to rotations on the Bloch sphere: the X,Y,ZX, Y, Z gates are rotations of π\pi about the three coordinate axes, and the Hadamard gate is a rotation of π\pi about a diagonal axis. Points in the interior of the sphere correspond to mixed states, which will be fully discussed through the density matrix in Section 2.6.

Connection to Quantum Computing: The Bloch sphere is an intuitive tool for single-qubit operations. Any single-qubit unitary gate corresponds to a rotation on the sphere, with its rotation axis and angle determined by a linear combination of Pauli matrices. Quantum state tomography reconstructs the Bloch vector by measuring expectation values along the X,Y,ZX, Y, Z directions, thereby fully characterizing an unknown quantum state. In quantum error correction, we need to encode quantum information into a higher-dimensional space so that Bloch vector deviations caused by noise can be detected and corrected. Understanding the Bloch sphere is the foundation for designing single-qubit gate sequences (such as optimal control pulses).


2.6 Measurement Theory and Density Matrices

In the preceding sections, we assumed the system is in a definite pure state ψ|\psi\rangle. In physical reality, however, we often face situations where “we do not know exactly which state the system is in”—the system may be in one of several possible states with a certain probability distribution. Furthermore, when we observe only part of a composite system, even if the whole is in a pure state, the subsystem can exhibit “mixed” behavior. The density matrix is the unified mathematical tool for describing such situations, and is key to understanding quantum entanglement and quantum error correction.

Projective Measurement

Let us begin from the measurement postulate of Section 2.2 and give a more systematic exposition using the projection operator language of Sections 1.4 and 1.5.

An observable MM corresponds to a Hermitian operator; by the spectral decomposition theorem (Section 1.5):

M=imiPiM = \sum_i m_i P_i

where mim_i are distinct real eigenvalues, and Pi=mimiP_i = |m_i\rangle\langle m_i| are the projection operators onto the eigenstates mi|m_i\rangle. Projection operators satisfy:

  • Hermiticity: Pi=PiP_i^\dagger = P_i
  • Idempotence: Pi2=PiP_i^2 = P_i (projecting again does not change the result—this is the property discussed in Section 1.4)
  • Orthogonality: PiPj=δijPiP_i P_j = \delta_{ij} P_i (projections onto different eigenspaces are mutually exclusive)
  • Completeness: iPi=I\sum_i P_i = I (all eigenspaces together span the entire state space)

When measuring MM on a system in state ψ|\psi\rangle:

  1. Probability of obtaining result mim_i: P(mi)=ψPiψ=Piψ2P(m_i) = \langle\psi|P_i|\psi\rangle = ||P_i|\psi\rangle||^2

    This is precisely the squared length of the component of ψ|\psi\rangle projected onto the direction of mi|m_i\rangle.

  2. Post-measurement state: ψobtain miPiψP(mi)|\psi\rangle \xrightarrow{\text{obtain } m_i} \frac{P_i|\psi\rangle}{\sqrt{P(m_i)}}

    The state vector is “projected” onto the corresponding eigenspace and then re-normalized.

Expectation Value

Upon repeated measurements, the average of the outcomes mim_i—the expectation value—is:

M=imiP(mi)=imiψPiψ=ψ(imiPi)ψ=ψMψ\langle M \rangle = \sum_i m_i P(m_i) = \sum_i m_i \langle\psi|P_i|\psi\rangle = \langle\psi|\left(\sum_i m_i P_i\right)|\psi\rangle = \langle\psi|M|\psi\rangle

This compact formula directly connects measurement statistics to operators.

Numerical Example (measuring +|+\rangle with ZZ):

Let ψ=+=0+12|\psi\rangle = |+\rangle = \frac{|0\rangle + |1\rangle}{\sqrt{2}}, and measure the observable Z=0011Z = |0\rangle\langle 0| - |1\rangle\langle 1|.

The projection operators are P0=00P_0 = |0\rangle\langle 0| and P1=11P_1 = |1\rangle\langle 1|, corresponding to measurement outcomes +1+1 and 1-1.

Computing probabilities:

P(+1)=+P0+=12(0+1)00(0+1)=120000=12P(+1) = \langle +|P_0|+\rangle = \frac{1}{2}(\langle 0| + \langle 1|)|0\rangle\langle 0|(|0\rangle + |1\rangle) = \frac{1}{2}\langle 0|0\rangle\langle 0|0\rangle = \frac{1}{2}

P(1)=+P1+=12P(-1) = \langle +|P_1|+\rangle = \frac{1}{2}

Post-measurement state: if +1+1 is obtained,

ψ=P0+1/2=00+1/2=0(1/2)1/2=0|\psi'\rangle = \frac{P_0|+\rangle}{\sqrt{1/2}} = \frac{|0\rangle\langle 0|+\rangle}{1/\sqrt{2}} = \frac{|0\rangle \cdot (1/\sqrt{2})}{1/\sqrt{2}} = |0\rangle

If 1-1 is obtained, ψ=1|\psi'\rangle = |1\rangle.

Expectation value:

Z=(+1)12+(1)12=0\langle Z \rangle = (+1)\cdot\frac{1}{2} + (-1)\cdot\frac{1}{2} = 0

Direct verification: +Z+=12(1,1)(1001)(11)=12(11)=0\langle +|Z|+\rangle = \frac{1}{2}(1, 1)\begin{pmatrix}1 & 0 \\ 0 & -1\end{pmatrix}\begin{pmatrix}1 \\ 1\end{pmatrix} = \frac{1}{2}(1 - 1) = 0.

Introducing the Density Matrix

Now consider a more general situation: the system is in state ψi|\psi_i\rangle with probability pip_i (i=1,2,,Ni = 1, 2, \ldots, N), where ipi=1\sum_i p_i = 1. When we measure an observable MM, the probability of obtaining mjm_j is:

P(mj)=ipiψiPjψiP(m_j) = \sum_i p_i \langle\psi_i|P_j|\psi_i\rangle

The expectation value is:

M=ipiψiMψi\langle M \rangle = \sum_i p_i \langle\psi_i|M|\psi_i\rangle

Direct computation requires taking inner products for each possible state and then taking the weighted average—a cumbersome procedure. The density matrix provides a unified, compact description.

Definition: The density matrix or density operator of a system is defined as:

ρ=ipiψiψi\rho = \sum_i p_i |\psi_i\rangle\langle\psi_i|

This is a Hermitian, positive semidefinite operator (matrix) with unit trace.

Using the density matrix, probabilities and expectation values can be written as:

P(mj)=Tr(ρPj),M=Tr(ρM)P(m_j) = \text{Tr}(\rho P_j), \quad \langle M \rangle = \text{Tr}(\rho M)

where Tr(A)=kkAk\text{Tr}(A) = \sum_k \langle k|A|k\rangle is the trace of the matrix, independent of the choice of basis.

Pure States and Mixed States

  • Pure State: The system is definitely in some state ψ|\psi\rangle. In this case p=1p = 1, and the density matrix is: ρ=ψψ\rho = |\psi\rangle\langle\psi|

    A pure-state density matrix satisfies idempotence: ρ2=ψψψψ=ψψ=ρ\rho^2 = |\psi\rangle\langle\psi|\psi\rangle\langle\psi| = |\psi\rangle\langle\psi| = \rho.

    Hence: Tr(ρ2)=Tr(ρ)=1\text{Tr}(\rho^2) = \text{Tr}(\rho) = 1

  • Mixed State: The system is in multiple states with a non-trivial probability distribution. In this case: ρ2=i,jpipjψiψiψjψjρ\rho^2 = \sum_{i,j} p_i p_j |\psi_i\rangle\langle\psi_i|\psi_j\rangle\langle\psi_j| \neq \rho

    and: Tr(ρ2)=ipi2<1\text{Tr}(\rho^2) = \sum_i p_i^2 < 1

    (Because pi<1p_i < 1 and pi=1\sum p_i = 1, so pi2<(pi)2=1\sum p_i^2 < (\sum p_i)^2 = 1.)

Tr(ρ2)\text{Tr}(\rho^2) is called the purity and serves as the criterion distinguishing pure from mixed states: Tr(ρ2)=1\text{Tr}(\rho^2) = 1 if and only if ρ\rho describes a pure state; Tr(ρ2)<1\text{Tr}(\rho^2) < 1 for a mixed state. On the Bloch sphere (Section 2.5), purity corresponds to the square of the distance from the center: Tr(ρ2)=12(1+r2)\text{Tr}(\rho^2) = \frac{1}{2}(1 + r^2), where r2=x2+y2+z2r^2 = x^2 + y^2 + z^2.

Comparative Example: Pure vs Mixed State

Consider the following two single-qubit states:

State A (pure): ψ=+=0+12|\psi\rangle = |+\rangle = \frac{|0\rangle + |1\rangle}{\sqrt{2}}

Density matrix: ρA=++=12(11)(1,1)=12(1111)\rho_A = |+\rangle\langle +| = \frac{1}{2}\begin{pmatrix}1 \\ 1\end{pmatrix}(1, 1) = \frac{1}{2}\begin{pmatrix}1 & 1 \\ 1 & 1\end{pmatrix}

Purity: ρA2=14(2222)=12(1111)=ρA\rho_A^2 = \frac{1}{4}\begin{pmatrix}2 & 2 \\ 2 & 2\end{pmatrix} = \frac{1}{2}\begin{pmatrix}1 & 1 \\ 1 & 1\end{pmatrix} = \rho_A, hence Tr(ρA2)=Tr(ρA)=1\text{Tr}(\rho_A^2) = \text{Tr}(\rho_A) = 1. Bloch vector: (1,0,0)(1, 0, 0), on the sphere surface.

State B (mixed): The system has a 50% probability of being in 0|0\rangle and a 50% probability of being in 1|1\rangle. Note: this is not the superposition state +|+\rangle—it is “classical uncertainty”; we do not know whether the system is 0|0\rangle or 1|1\rangle, only their respective probabilities.

Density matrix: ρB=1200+1211=12(1000)+12(0001)=12(1001)=I2\rho_B = \frac{1}{2}|0\rangle\langle 0| + \frac{1}{2}|1\rangle\langle 1| = \frac{1}{2}\begin{pmatrix}1 & 0 \\ 0 & 0\end{pmatrix} + \frac{1}{2}\begin{pmatrix}0 & 0 \\ 0 & 1\end{pmatrix} = \frac{1}{2}\begin{pmatrix}1 & 0 \\ 0 & 1\end{pmatrix} = \frac{I}{2}

Purity: ρB2=I4\rho_B^2 = \frac{I}{4}, Tr(ρB2)=Tr(I4)=12<1\text{Tr}(\rho_B^2) = \text{Tr}(\frac{I}{4}) = \frac{1}{2} < 1. Bloch vector: (0,0,0)(0, 0, 0), at the center of the sphere.

These two states have entirely different physical meanings: ρA\rho_A describes the quantum superposition of 0|0\rangle and 1|1\rangle—measuring ZZ yields ±1\pm 1 with equal probability, but the system is in a definite coherent state. ρB\rho_B describes a classical mixture—measuring ZZ also yields ±1\pm 1 with equal probability, but there is no quantum coherence; the system “really is” 0|0\rangle or “really is” 1|1\rangle, we merely do not know which.

The crucial test: measure XX. For ρA\rho_A, X=Tr(ρAX)=1\langle X \rangle = \text{Tr}(\rho_A X) = 1 (because +|+\rangle is an eigenstate of XX with eigenvalue +1+1). For ρB\rho_B, X=Tr(I2X)=12Tr(X)=0\langle X \rangle = \text{Tr}(\frac{I}{2}X) = \frac{1}{2}\text{Tr}(X) = 0. The two yield completely different experimental predictions, despite having the same statistics under ZZ measurement.

Reduced Density Matrix

One of the most important applications of the density matrix is describing subsystems within a composite system. Consider a pure state ΨAB|\Psi\rangle_{AB} of a bipartite system; its density matrix is ρAB=ΨΨ\rho_{AB} = |\Psi\rangle\langle\Psi|.

When we focus only on particle AA (“ignoring” particle BB), the physics of AA is described by the reduced density matrix:

ρA=TrB(ρAB)\rho_A = \text{Tr}_B(\rho_{AB})

where TrB\text{Tr}_B denotes the partial trace over subsystem BB: if {bj}\{|b_j\rangle\} is an orthonormal basis for BB, then:

ρA=jbjρABbj\rho_A = \sum_j \langle b_j|\rho_{AB}|b_j\rangle

The result of the partial trace is an operator acting only on the Hilbert space of AA.

Key fact: Even if ρAB\rho_{AB} is a pure state (the whole has a perfectly definite quantum state), ρA\rho_A is generally a mixed state! This seemingly paradoxical phenomenon is precisely the mathematical signature of quantum entanglement.

Example: Consider the Bell state Φ+=00+112|\Phi^+\rangle = \frac{|00\rangle + |11\rangle}{\sqrt{2}}. The whole is a pure state, ρAB=Φ+Φ+\rho_{AB} = |\Phi^+\rangle\langle\Phi^+|.

Compute ρA\rho_A:

ρA=TrB(ρAB)=0BρAB0B+1BρAB1B\rho_A = \text{Tr}_B(\rho_{AB}) = \langle 0_B|\rho_{AB}|0_B\rangle + \langle 1_B|\rho_{AB}|1_B\rangle

Computing the first term: 0BΦ+=120B(0A0B+1A1B)=120A\langle 0_B|\Phi^+\rangle = \frac{1}{\sqrt{2}}\langle 0_B|(|0_A 0_B\rangle + |1_A 1_B\rangle) = \frac{1}{\sqrt{2}}|0_A\rangle

So 0BρAB0B=120A0A\langle 0_B|\rho_{AB}|0_B\rangle = \frac{1}{2}|0_A\rangle\langle 0_A|.

Similarly, 1BρAB1B=121A1A\langle 1_B|\rho_{AB}|1_B\rangle = \frac{1}{2}|1_A\rangle\langle 1_A|.

Therefore:

ρA=1200+1211=I2\rho_A = \frac{1}{2}|0\rangle\langle 0| + \frac{1}{2}|1\rangle\langle 1| = \frac{I}{2}

This is the maximally mixed state! Tr(ρA2)=1/2<1\text{Tr}(\rho_A^2) = 1/2 < 1.

Physical interpretation: Subsystem AA, “looked at in isolation” from the entangled state Φ+|\Phi^+\rangle, yields a description identical to a classical mixture “50% probability of 0|0\rangle, 50% probability of 1|1\rangle.” Subsystem AA contains no definite information about its own state—all the information resides in the correlations between AA and BB.

This property is the foundation of quantum cryptography’s security: if someone attempts to eavesdrop on quantum communication (i.e., perform a measurement on subsystem BB), they inevitably disrupt the entangled state, leaving detectable traces.

Summary: Projective measurement is described by the spectral decomposition of a Hermitian operator; the projection operators Pi=mimiP_i = |m_i\rangle\langle m_i| project the state onto the corresponding eigenspace; the measurement probability is P(mi)=ψPiψP(m_i) = \langle\psi|P_i|\psi\rangle (citing Sections 1.4 and 1.5). The expectation value is M=ψMψ\langle M\rangle = \langle\psi|M|\psi\rangle. The density matrix ρ=ipiψiψi\rho = \sum_i p_i|\psi_i\rangle\langle\psi_i| provides a unified description of pure and mixed states: pure states satisfy ρ2=ρ\rho^2 = \rho and Tr(ρ2)=1\text{Tr}(\rho^2) = 1; mixed states satisfy Tr(ρ2)<1\text{Tr}(\rho^2) < 1. The reduced density matrix ρA=TrB(ρAB)\rho_A = \text{Tr}_B(\rho_{AB}) describes the state of a subsystem within a composite system; the reduced density matrix of an entangled pure state is generally mixed. This is the key to understanding the nature of quantum entanglement: a subsystem of an entangled state, viewed in isolation, appears as a completely random classical mixture; all quantum information resides in the correlations.

Connection to Quantum Computing: The density matrix is the standard tool for analyzing open quantum systems and quantum noise. Real quantum computers are not perfectly isolated—they inevitably couple to the environment, leading to decoherence. The decoherence process turns pure states into mixed states and is described by non-unitary evolution of the density matrix. Quantum error-correcting codes encode logical quantum information into entangled states of multiple physical qubits, so that local noise causes only small, correctable perturbations. The entanglement entropy S=Tr(ρAlogρA)S = -\text{Tr}(\rho_A \log \rho_A) of the reduced density matrix quantifies the degree of entanglement between subsystems and is an important metric for evaluating a quantum circuit’s ability to generate entanglement resources. In the theoretical analysis of quantum machine learning, quantum thermodynamics, and quantum communication, the density matrix is an indispensable mathematical tool.


Chapter Summary

This chapter has “physically instantiated” the abstract mathematical tools of Chapter 1, one by one:

  • Section 2.1 demonstrated the failure of classical physics in the microscopic world, with the double-slit experiment revealing the essence of quantum probability ψA+ψB2ψA2+ψB2|\psi_A + \psi_B|^2 \neq |\psi_A|^2 + |\psi_B|^2—complex interference.
  • Section 2.2 systematically established the five postulates of quantum mechanics, each directly corresponding to the mathematical structures of Sections 1.2–1.6: Hilbert space, unitary operators, Hermitian operators and spectral decomposition, projection operators and probability, and tensor products.
  • Section 2.3, through the stationary problem of the Schrödinger equation, demonstrated the physical meaning of the eigenvalue equation Hψ=EψH|\psi\rangle = E|\psi\rangle and gave a complete solution for the infinite square well.
  • Section 2.4 connected the Pauli matrices (Section 1.3) to the spin-1/2 system, with the Stern–Gerlach experiment providing the physical foundation for quantization.
  • Section 2.5’s Bloch sphere provided a geometric picture for the qubit, with unitary operations corresponding to rotations on the sphere.
  • Section 2.6’s density matrix unified the description of pure and mixed states, and the reduced density matrix revealed the profound nature of entanglement.

These six sections constitute the complete physical foundation of quantum computing. The reader has now mastered: from mathematical structure to physical interpretation, from a single qubit to composite systems, from ideal pure states to realistic mixed states—the entire theoretical framework. The next chapter will, on this foundation, formally construct the computational model of quantum computing.

Appendix

Part 2: Feynman Workbook — Physical Foundations of Quantum Mechanics

Method: The Feynman Technique — if you cannot explain a concept in simple language, you have not yet truly understood it.

Difficulty Legend: ⭐ Computation/Verification · ⭐⭐ Derivation/Proof · 🗣️ Feynman Explanation · 💭 Insight Challenge

This workbook accompanies Sections 2.1–2.6 of the main tutorial.


2.1 From Classical to Quantum: Motivation & History

⭐ Temperature Dependence of the Ultraviolet Catastrophe

Black-body radiation experiments measured peak wavelengths at three temperatures: λmax(1)=966nm\lambda_{\max}^{(1)} = 966\,\text{nm} at T1=3000KT_1 = 3000\,\text{K}; λmax(2)=724nm\lambda_{\max}^{(2)} = 724\,\text{nm} at T2=4000KT_2 = 4000\,\text{K}; λmax(3)=580nm\lambda_{\max}^{(3)} = 580\,\text{nm} at T3=5000KT_3 = 5000\,\text{K}. Verify whether these three data points satisfy Wien’s displacement law λmaxT=constant\lambda_{\max} T = \text{constant}, and compute the experimental value of the constant. Planck’s formula reduces to Wien’s formula in the high-frequency limit—verify that when hνkBTh\nu \gg k_B T, u(ν,T)8πhν3c3ehν/kBTu(\nu, T) \approx \frac{8\pi h\nu^3}{c^3} e^{-h\nu/k_B T}.


⭐⭐ Conditions for the Disappearance of Interference Fringes

In the double-slit experiment, the wave functions from the two slits are ψL(x)=PLeiϕL\psi_L(x) = \sqrt{P_L}\, e^{i\phi_L} and ψR(x)=PReiϕR\psi_R(x) = \sqrt{P_R}\, e^{i\phi_R}, where PL+PR=1P_L + P_R = 1. Derive the expression for the quantum probability Pquantum=ψL+ψR2P_{\text{quantum}} = |\psi_L + \psi_R|^2 and prove that the magnitude of the interference term is bounded by 2PLPR2\sqrt{P_L P_R}. How does the interference term behave when PLPRP_L \ll P_R (one slit nearly closed)? Does your derivation imply that the interference fringes vanish in the limit PL0P_L \to 0?


🗣️ Feynman: Explaining Quantum Probability to Your Grandmother

Your grandmother has heard that “quantum mechanics says an electron can be in two places at once.” She frowns and says, “That’s absurd—how can one thing be in two places at once? Are you joking with me?”

Requirement: Without using any mathematical formulas, without using terms like “probability amplitude” or “superposition state,” using only everyday analogies and plain language, explain to your grandmother what the double-slit experiment is actually saying. Help her understand:

  • Why the question “which slit did the electron go through?” might itself be a wrong question
  • Why measurement affects the outcome
  • What is fundamentally different about quantum probability versus classical probability

💭 Insight: Einstein’s “God Does Not Play Dice”

In a 1926 letter to Born, Einstein wrote: “Quantum mechanics is certainly imposing. But an inner voice tells me that it is not yet the real thing. The theory says a lot, but does not really bring us any closer to the secret of the Old One. I, at any rate, am convinced that He does not throw dice.”

Einstein never accepted the probabilistic interpretation of quantum mechanics throughout his life. Yet his paper on the photoelectric effect (1905) is itself among the most powerful pieces of evidence for the quantum concept.

Reflect: What exactly was Einstein objecting to? Was it probability per se, or the assertion that “probability is the final answer”? If Einstein had worked in the era of the Many-Worlds Interpretation, would he still have said “God does not play dice”? Or does the “unacceptability” of the probabilistic interpretation stem from our innate classical-deterministic intuition?


2.2 Quantum Mechanical Postulates

⭐ Measurement Probabilities in Three Different Bases

The quantum state is ψ=130+231\lvert\psi\rangle = \frac{1}{\sqrt{3}}\lvert0\rangle + \sqrt{\frac{2}{3}}\lvert1\rangle. Compute the probability of obtaining the first result when measuring in each of the following three bases:

(a) ZZ basis {0,1}\{\lvert0\rangle, \lvert1\rangle\}: the probability of obtaining 0\lvert0\rangle.

(b) XX basis {+,}\{\lvert+\rangle, \lvert-\rangle\}: the probability of obtaining +\lvert+\rangle. (Hint: expand ψ\lvert\psi\rangle in terms of +,\lvert+\rangle, \lvert-\rangle.)

(c) YY basis {+i,i}\{\lvert+i\rangle, \lvert-i\rangle\}: the probability of obtaining +i\lvert+i\rangle.

Verify: does the sum of the three probabilities exceed 1? If so, is this a contradiction?


⭐⭐ Derivation of the Expectation Value Formula for Observables

Prove: for an observable (Hermitian operator) AA and an arbitrary normalized state ψ\lvert\psi\rangle, the expectation value of the measurement outcome satisfies A=ψAψ\langle A \rangle = \langle\psi|A|\psi\rangle. Your derivation should include the following steps:

(1) Use the spectral decomposition A=iaiPiA = \sum_i a_i P_i, where Pi=aiaiP_i = \lvert a_i\rangle\langle a_i\rvert.

(2) Write P(ai)P(a_i) from the Born rule.

(3) Write the definition of the expectation value A=iaiP(ai)\langle A \rangle = \sum_i a_i P(a_i).

(4) Combine the above steps to complete the proof.

Bonus challenge: If AA has degenerate eigenvalues, how must the proof be modified?


🗣️ Feynman: The “Rules of the Game” of Quantum Mechanics

Imagine you are a board game designer, creating a new game for your friends. Your friends have never encountered quantum mechanics, but you need the “quantum rules” in the game to be both accurate and easy to understand.

Requirement: Using board game analogies (game board, pieces, dice, cards, etc.), explain the five postulates of quantum mechanics to your friends. Each postulate corresponds to one rule of the game. Precise mathematical formulation is not required; instead, use game mechanics as analogies:

  • Postulate 1 (State Space): positions on the board
  • Postulate 2 (Unitary Evolution): movement rules
  • Postulate 3 (Observables): scoring criteria
  • Postulate 4 (Measurement): flipping a card / revealing a result
  • Postulate 5 (Composite Systems): multiplayer cooperative mode

After hearing your explanation, your friends should say, “Oh, so quantum mechanics isn’t so mysterious after all!”


💭 Insight: Is the Measurement Problem a “Problem”?

Postulate 4 says: measurement causes wave function collapse. But Postulate 2 says: the evolution of a closed system is described by a unitary operator. Collapse is not unitary evolution.

This creates a fundamental tension: the physical process of the measurement apparatus itself ought to be described by quantum mechanics (after all, the measurement apparatus is made of atoms too!). But if we also include the measurement apparatus in the system, then the whole “system + apparatus” should obey unitary evolution—where does collapse come from, then?

This question is called the “measurement problem,” and has troubled physicists since the birth of quantum mechanics. For a hundred years, no one has truly “solved” it—different interpretations have only been proposed to sidestep or dissolve it.

Reflect:

  • Is “collapse” a real physical process, or a psychological description of how we update our information?
  • If you think collapse is real, what triggers it? How large or complex must a system be to trigger collapse?
  • If you think collapse is not real, why do measurement outcomes appear definite?
  • Is there a third possibility—that collapse is neither real nor illusory, but that something is amiss with the way we ask the question?

2.3 Wave Functions and the Schrödinger Equation

⭐ The First Excited State of the Infinite Square Well

In the one-dimensional infinite square well [0,L][0, L], the wave function of the first excited state (n=2n=2) is ψ2(x)=2/Lsin(2πx/L)\psi_2(x) = \sqrt{2/L}\sin(2\pi x/L).

(a) Verify that ψ2\psi_2 satisfies the normalization condition.

(b) Compute the position expectation value x\langle x \rangle in this state.

(c) Compute the expectation value of the position squared x2\langle x^2 \rangle in this state. (Hint: sin2θ=(1cos2θ)/2\sin^2\theta = (1-\cos 2\theta)/2 and 0Lx2cos(4πx/L)dx\int_0^L x^2\cos(4\pi x/L)\,dx can be evaluated using integration by parts.)

(d) Using (b) and (c), compute the position uncertainty Δx=x2x2\Delta x = \sqrt{\langle x^2 \rangle - \langle x \rangle^2}.


⭐⭐ Momentum-Space Wave Function

The ground-state wave function of the infinite square well is ψ1(x)=2/Lsin(πx/L)\psi_1(x) = \sqrt{2/L}\sin(\pi x/L) (for 0<x<L0 < x < L), and zero outside.

The momentum-space wave function is defined as ϕ(p)=12π+ψ(x)eipx/dx\phi(p) = \frac{1}{\sqrt{2\pi\hbar}}\int_{-\infty}^{+\infty} \psi(x) e^{-ipx/\hbar}\,dx.

Compute ϕ(p)\phi(p) and prove that: ϕ(p)=1πLπ/L(π/L)2(p/)2[eipL/cos(pL/)ipπ/Lsin(pL/)]\phi(p) = \frac{1}{\sqrt{\pi L\hbar}} \frac{\pi\hbar/L}{(\pi\hbar/L)^2 - (p/\hbar)^2} \left[ e^{-ipL/\hbar} \cos(pL/\hbar) - \frac{ip\hbar}{\pi\hbar/L}\sin(pL/\hbar) \right]

(Hint: write sin\sin in exponential form, combine the integrals, and obtain two Fourier-like integrals.)

From the expression for ϕ(p)\phi(p), can you discern the main features of the momentum distribution? How many peaks does it have? Where are the peaks located?


🗣️ Feynman: The “Intuition” Behind the Schrödinger Equation

The Schrödinger equation itψ=Hψi\hbar\partial_t\psi = H\psi is the fundamental dynamical equation of quantum mechanics, analogous to F=maF = ma in classical mechanics. But its mathematical form—containing the imaginary unit ii, partial derivatives, and the Hamiltonian—often makes beginners feel that “this equation fell from the sky.”

Requirement: Without using any mathematical formulas (you may write itψ=Hψi\hbar\partial_t\psi = H\psi as a reference only), explain to a friend who has studied high-school physics:

  1. What is the Schrödinger equation actually saying? What is its “physical picture”?
  2. What does the ii in the equation signify? (Why is it not a real equation?)
  3. Why is the Schrödinger equation “quantum”? How does it differ intuitively from the classical wave equation?
  4. If F=maF = ma describes “the trajectory of a particle,” what does the Schrödinger equation describe?

You may use any analogies (water flow, sound waves, the propagation of probability, etc.), but must explain the limitations of each analogy.


💭 Insight: What Does the Wave Function Actually “Is”?

When Schrödinger first proposed the wave function, he thought ψ\psi described “some kind of matter wave of the electron”—the electron spreads through space like a cloud, and ψ2|\psi|^2 is the charge density. But Born soon pointed out: ψ2|\psi|^2 is a probability density, not a matter density. Schrödinger never agreed with this interpretation throughout his life.

This controversy has not fully subsided to this day: is the wave function an objectively real physical field (ontic), or a description of our state of knowledge (epistemic)?

Consider the following scenarios:

Scenario A: An electron is in a position superposition state ψ=(x1+x2)/2\lvert\psi\rangle = (\lvert x_1\rangle + \lvert x_2\rangle)/\sqrt{2}. Measurement finds the electron at x1x_1. According to the standard interpretation, “the wave function collapsed”—the probability at x2x_2 instantly became zero.

Scenario B: Two particles in an entangled pair Φ+=(00+11)/2\lvert\Phi^+\rangle = (\lvert00\rangle + \lvert11\rangle)/\sqrt{2} are separated by 1 light-year. Measuring particle A yields 00, and particle B’s wave function instantaneously collapses to 0\lvert0\rangle.

Does the “instantaneous collapse” in these two scenarios constitute superluminal signal transmission? If not, why not? What is the relationship between the “reality” of the wave function and nonlocality? Can something “unreal” exhibit nonlocal correlations?


2.4 Two-Level Systems and Spin

⭐ Spin Measurement in an Arbitrary Direction

A spin-1/2 system is in the state ψ=340+121\lvert\psi\rangle = \sqrt{\frac{3}{4}}\lvert0\rangle + \frac{1}{2}\lvert1\rangle. Consider the spin operator along the direction n=(sinθcosϕ,sinθsinϕ,cosθ)\vec{n} = (\sin\theta\cos\phi, \sin\theta\sin\phi, \cos\theta): Sn=2(sinθcosϕX+sinθsinϕY+cosθZ)S_{\vec{n}} = \frac{\hbar}{2}(\sin\theta\cos\phi\,X + \sin\theta\sin\phi\,Y + \cos\theta\,Z).

(1) For θ=π/4\theta = \pi/4, ϕ=0\phi = 0 (i.e., at 4545^\circ to the zz-axis in the xx-zz plane), compute the matrix representation of SnS_{\vec{n}}.

(2) Find the two eigenstates of SnS_{\vec{n}} (they may be expressed parametrically), and use them to compute the probability of obtaining +/2+\hbar/2 when measuring along this direction.

(3) When θ=0\theta = 0, to which simpler measurement does this probability reduce?


⭐⭐ Explicit Form of the Rotation Operator

For any unit vector n=(nx,ny,nz)\vec{n} = (n_x, n_y, n_z) and any angle α\alpha, prove that the rotation operator Rn(α)=eiα(nσ)/2R_{\vec{n}}(\alpha) = e^{-i\alpha(\vec{n}\cdot\vec{\sigma})/2} can be expanded as:

Rn(α)=cosα2Iisinα2(nxX+nyY+nzZ)R_{\vec{n}}(\alpha) = \cos\frac{\alpha}{2}\,I - i\sin\frac{\alpha}{2}\,(n_x X + n_y Y + n_z Z)

The key step in the derivation is to use the property of Pauli matrices that (nxX+nyY+nzZ)2=I(n_x X + n_y Y + n_z Z)^2 = I (proving this is itself part of the derivation).

Using this formula, compute: (1) Rx(π)R_x(\pi) — a rotation of π\pi about the xx-axis. What should this yield? (2) Ry(π/2)R_y(\pi/2) — a rotation of π/2\pi/2 about the yy-axis. Acting on 0\lvert0\rangle, what does this yield?

Verify the relationship Rx(π)=iXR_x(\pi) = -iX (up to the global phase i-i) with the XX gate.


🗣️ Feynman: The “Spatial Orientation Paradox” of Spin

Imagine this: you have a small magnetic needle that can point in any direction—north, northeast, east, and so on. That is the intuition of the classical world.

But if you have an electron, its “spin” measured along the zz-direction can only point “up” or “down.” The strangest part: if you first measure that it is indeed “up” along the zz-direction, then measure it along the xx-direction—“up” and “down” each appear with 50% probability! You go back and measure the zz-direction again, and it is 50-50 again! But you clearly remember that just a moment ago it was “up.”

Requirement: Explain to a high-school student who has never studied quantum mechanics:

  1. Why does the zz-direction spin information become “lost” after measuring the xx-direction?
  2. Why can’t we simultaneously know the values of the spin along the zz-direction and the xx-direction? How is this different from classical physics?
  3. What is the essential difference between spin and the “angular momentum” of classical rotation?
  4. Using the Stern–Gerlach apparatus as an analogy—why does a beam of silver atoms passing through an inhomogeneous magnetic field split into two beams rather than forming a continuous distribution?

You may draw diagrams to assist (text descriptions suffice), but do not use matrices or algebra.


💭 Insight: Spin — Quantum Mechanics’ “Purest Miracle”

Spin has no classical counterpart. It is not the electron “spinning” (if it were truly spinning, the electron’s surface speed would far exceed the speed of light). Spin is a purely intrinsic quantum degree of freedom; its existence and properties arise directly from the requirements of relativistic quantum mechanics (the Dirac equation)—not as an artificial assumption, but as a mathematical necessity.

When Dirac attempted to combine the Schrödinger equation with special relativity in 1928, he found that the solutions naturally contained a four-component spinor, of which two components correspond to the electron, two to the positron, and the electron’s two components correspond exactly to spin up and spin down. Spin was not “added in” — it “emerged.”

Reflect:

  1. If spin is a product of relativistic quantum mechanics, where does spin come from in non-relativistic quantum mechanics (the Schrödinger equation)? How was it “forcefully inserted”?
  2. The spin-statistics theorem states: particles with half-integer spin are fermions (obeying the Pauli exclusion principle), and particles with integer spin are bosons. This theorem is a result of relativistic quantum field theory; in the non-relativistic framework, it is merely an empirical assumption. Why is there such a deep intrinsic connection between spin and statistics?
  3. Pauli matrices satisfy σiσj=δijI+iεijkσk\sigma_i\sigma_j = \delta_{ij}I + i\varepsilon_{ijk}\sigma_k. This algebraic structure has a profound mathematical connection with quaternions—is this just a coincidence, or does it hint at some more fundamental structure of spacetime?

2.5 The Bloch Sphere

⭐ Geometric Positions of Four States on the Bloch Sphere

Given the following four quantum states, compute their Bloch coordinates (x,y,z)(x, y, z) and picture their positions in your mind:

(1) A=320+121\lvert A\rangle = \frac{\sqrt{3}}{2}\lvert0\rangle + \frac{1}{2}\lvert1\rangle

(2) B=120+eiπ/321\lvert B\rangle = \frac{1}{\sqrt{2}}\lvert0\rangle + \frac{e^{i\pi/3}}{\sqrt{2}}\lvert1\rangle

(3) C=150+251\lvert C\rangle = \frac{1}{\sqrt{5}}\lvert0\rangle + \frac{2}{\sqrt{5}}\lvert1\rangle

(4) D=110(0+3i1)\lvert D\rangle = \frac{1}{\sqrt{10}}( \lvert0\rangle + 3i\lvert1\rangle)

For each state, indicate whether it is in the northern hemisphere, southern hemisphere, on the equator, or at a pole. Which two states are closest on the sphere? Quantify using the spherical distance (central angle).


⭐⭐ Composition of Arbitrary Rotations

On the Bloch sphere, first rotate by angle α\alpha about the zz-axis (Rz(α)=eiαZ/2R_z(\alpha) = e^{-i\alpha Z/2}), then rotate by angle β\beta about the xx-axis (Rx(β)=eiβX/2R_x(\beta) = e^{-i\beta X/2}), and finally rotate by angle γ\gamma about the zz-axis (Rz(γ)=eiγZ/2R_z(\gamma) = e^{-i\gamma Z/2}). This sequence Rz(γ)Rx(β)Rz(α)R_z(\gamma)R_x(\beta)R_z(\alpha) is called the Z–Y decomposition or Euler angle decomposition.

Prove: any single-qubit unitary operation USU(2)U \in SU(2) can be expressed in the above form; i.e., there exist real parameters α,β,γ\alpha, \beta, \gamma such that: U=eiγZ/2eiβX/2eiαZ/2U = e^{-i\gamma Z/2} e^{-i\beta X/2} e^{-i\alpha Z/2}

Hint: Use the rotation formula from Section 2.4 to expand the three factors as matrices, then multiply and set the result equal to a general 2×22\times2 unitary matrix (abba)\begin{pmatrix} a & b \\ -b^* & a^* \end{pmatrix} (satisfying det=1\det = 1, a2+b2=1|a|^2 + |b|^2 = 1), and solve for α,β,γ\alpha, \beta, \gamma in terms of a,ba, b.

This result is extremely important in quantum computing: any single-qubit gate can be implemented in hardware via three parameterized instructions.


🗣️ Feynman: The Bloch Sphere as a “Quantum Compass”

The Bloch sphere is the most intuitive geometric tool for understanding qubits. But its mathematical definition—ψ=cos(θ/2)0+eiϕsin(θ/2)1\lvert\psi\rangle = \cos(\theta/2)\lvert0\rangle + e^{i\phi}\sin(\theta/2)\lvert1\rangle mapped to coordinates (sinθcosϕ,sinθsinϕ,cosθ)(\sin\theta\cos\phi, \sin\theta\sin\phi, \cos\theta)—looks like gibberish to a beginner.

Requirement: Explain the Bloch sphere to a high-school student who has only studied spherical coordinates. Specifically:

  1. Why is the Bloch sphere a sphere? How can two complex parameters be represented by a point on a spherical surface?
  2. Why are the north and south poles 0\lvert0\rangle and 1\lvert1\rangle? Why not the equator?
  3. Why is +\lvert+\rangle on the positive xx-axis? Why does the relative phase ϕ\phi determine the longitude?
  4. Why do “rotations” on the Bloch sphere correspond to quantum gate operations?
  5. Why are mixed states inside the sphere? What physical meaning does the center of the sphere carry?

Use a “compass” or “globe” as your primary analogy. If you have the bandwidth, also explain why a 9090^\circ rotation (like the Hadamard gate) appears as a π\pi rotation on the Bloch sphere rather than π/2\pi/2—note the θ/2\theta/2 in the parameterization!


💭 Insight: Does the Bloch Sphere Hide Something?

The Bloch sphere is a beautiful geometric representation of the qubit state space, but it may also “oversimplify” certain things. The state space of an nn-qubit system is a (2n1)(2^n-1)-dimensional complex projective space (CP2n1^{2^n-1}), far from being a simple Cartesian product of nn Bloch spheres.

Consider the following questions:

  1. For two qubits, why can’t the state space simply be described as “two Bloch spheres”? How would the Bell state Φ+=(00+11)/2\lvert\Phi^+\rangle = (\lvert00\rangle + \lvert11\rangle)/\sqrt{2} be described within the “two Bloch spheres” framework?

  2. What does the fact that “entangled states have no counterpart on the Bloch sphere” tell us about the deep relationship between geometry and information? Does this suggest that quantum information requires a new geometric language?

  3. Mixed states of a single qubit lie inside the Bloch sphere—a ball of radius rr. This “ball” can be viewed as a manifold described by 3 real parameters. But for nn qubits, the geometric structure of mixed states is far from simple. Does there exist an elegant generalization of the Bloch representation for many-body systems?

  4. If we lived “inside” the Bloch sphere (i.e., we ourselves are also quantum systems), could we “see” the boundary of the sphere? If not, are there quantum degrees of freedom that we cannot probe?


2.6 Measurement Theory and Density Matrices

⭐ Density Matrix for a Two-Component Probabilistic Mixture

A quantum system is in state +\lvert+\rangle with probability pp and in state \lvert-\rangle with probability 1p1-p, where p[0,1]p \in [0, 1].

(a) Write the density matrix ρ(p)\rho(p) in explicit matrix form (in terms of pp).

(b) Compute Tr[ρ(p)]\operatorname{Tr}[\rho(p)] and Tr[ρ(p)2]\operatorname{Tr}[\rho(p)^2].

(c) For which values of pp is ρ\rho a pure state? For which value of pp is the mixedness maximal?

(d) Compare ρ(p)\rho(p) with the density matrix for “probability pp of 0\lvert0\rangle, probability 1p1-p of 1\lvert1\rangle.” When are they identical? When do they differ? (Hint: compare their measurement predictions in the XX and ZZ bases.)


⭐⭐ Entanglement Criterion via the Reduced Density Matrix

Consider two entangled states: Ψ1=12(00+01)\lvert\Psi_1\rangle = \frac{1}{\sqrt{2}}(\lvert00\rangle + \lvert01\rangle) Ψ2=12(00+11)\lvert\Psi_2\rangle = \frac{1}{\sqrt{2}}(\lvert00\rangle + \lvert11\rangle)

(a) Determine whether Ψ1\lvert\Psi_1\rangle is separable (i.e., whether it can be written in the form ψAϕB\lvert\psi\rangle_A \otimes \lvert\phi\rangle_B).

(b) For each state, compute ρA=TrB(ΨΨ)\rho_A = \operatorname{Tr}_B(\lvert\Psi\rangle\langle\Psi\rvert).

(c) Compute the purity Tr(ρA2)\operatorname{Tr}(\rho_A^2) for both ρA\rho_A. Can purity being less than 1 serve as a criterion for entanglement? Compare your judgment from (a) with the purity calculation—are they consistent?

(d) If the reduced density matrix of a bipartite pure state is mixed (Tr(ρA2)<1\operatorname{Tr}(\rho_A^2) < 1), what does this “mean”? Conversely, if ρA\rho_A is pure, must ΨAB\lvert\Psi\rangle_{AB} be a separable state?


🗣️ Feynman: Probabilistic Mixture vs Quantum Superposition

Suppose you have a mystery box. Experimental physicist Anna says, “The qubit in the box is in a superposition of 50% 0\lvert0\rangle and 50% 1\lvert1\rangle.” Her colleague Bob says, “No, the box contains either 0\lvert0\rangle or 1\lvert1\rangle, we just don’t know which—each with 50% probability.”

Anna and Bob make exactly identical predictions for ZZ measurements: 50% probability of +1+1, 50% probability of 1-1. Under the ZZ basis, their predictions are indistinguishable.

Requirement: Explain to Anna and Bob:

  1. How can the two cases be distinguished under the XX basis? Why can XX measurement “see through” the difference between quantum superposition and classical mixture?
  2. Use everyday analogies to explain the essential difference between “quantum coherence” and “classical uncertainty.” (Hint: imagine two perfectly synchronized pendulums vs two independent, uncorrelated pendulums.)
  3. Why is this distinction crucial for quantum computing? If qubits were merely bits in the “classical probability” sense, could a quantum computer still function?
  4. How does the density matrix, as a unified mathematical object, describe these two drastically different physical situations simultaneously? When would two different “preparation procedures” yield exactly the same density matrix?

💭 Insight: Entanglement — From “Spooky” to “Useful”

In a 1935 paper, Schrödinger described entanglement (Verschränkung) as “the characteristic trait of quantum mechanics… not one trait, but rather the trait.” In the same paper, he also proposed the famous “Schrödinger’s cat” thought experiment to showcase the absurd consequences of entanglement. That same year, the EPR paper used entanglement to argue for the incompleteness of quantum mechanics.

But today, entanglement is the central resource of quantum information science—quantum communication, quantum cryptography, quantum error correction, and quantum metrology all depend on entanglement. A phenomenon that Einstein called “spooky action at a distance” has become an engineering resource.

Consider the following questions:

  1. Schrödinger’s cat amplifies microscopic entanglement to the macroscopic scale. At what scale do we cease to observe quantum effects? Does “decoherence” solve the measurement problem, or merely push it one step further back?

  2. If entanglement is nonlocal, is “information” also nonlocal? Does quantum teleportation really “transmit” information? If information cannot be transmitted faster than light, in what sense is the nonlocality of entanglement “nonlocal”?

  3. Quantifying entanglement is a central problem in quantum information theory. For bipartite pure states, the entanglement entropy S(ρA)=Tr(ρAlogρA)S(\rho_A) = -\operatorname{Tr}(\rho_A\log\rho_A) is the unique entanglement measure. But for mixed states and many-body systems, there exist multiple inequivalent entanglement measures (entanglement of formation, distillable entanglement, relative entropy of entanglement, etc.). Why is the seemingly straightforward question “how much entanglement does this system have” so complex?

  4. Quantum error correction encodes logical information into multiple entangled physical qubits, so that even if some qubits decohere, the information remains intact. Does this suggest that “entanglement” and “information” may be two sides of the same coin? Could it be that “information” is the fundamental concept of quantum mechanics, while “state” and “entanglement” are merely its derivative phenomena?


Reference Answers

Note: For 🗣️ and 💭 open-ended questions, the reference answers are not “standard answers,” but provide directions for thinking and key points. Your answer may differ from the ones here; as long as it is well-reasoned, it is a good answer.


2.1 From Classical to Quantum: Motivation & History

⭐ Reference Answer

Wien displacement constant calculation:

λmax(1)T1=966×109×3000=2.898×103m⋅K\lambda_{\max}^{(1)} T_1 = 966 \times 10^{-9} \times 3000 = 2.898 \times 10^{-3}\,\text{m·K}

λmax(2)T2=724×109×4000=2.896×103m⋅K\lambda_{\max}^{(2)} T_2 = 724 \times 10^{-9} \times 4000 = 2.896 \times 10^{-3}\,\text{m·K}

λmax(3)T3=580×109×5000=2.900×103m⋅K\lambda_{\max}^{(3)} T_3 = 580 \times 10^{-9} \times 5000 = 2.900 \times 10^{-3}\,\text{m·K}

The three data sets are consistent, with the constant being approximately 2.898×103m⋅K2.898 \times 10^{-3}\,\text{m·K}, in agreement with the standard value of Wien’s displacement constant, 2.898×103m⋅K2.898 \times 10^{-3}\,\text{m·K}. Planck’s formula in the high-frequency limit: u(ν,T)=8πhν3c31ehν/kBT1u(\nu,T) = \frac{8\pi h\nu^3}{c^3}\frac{1}{e^{h\nu/k_BT} - 1}. When hνkBTh\nu \gg k_B T, the denominator ehν/kBT1ehν/kBTe^{h\nu/k_BT} - 1 \approx e^{h\nu/k_BT}, so u8πhν3c3ehν/kBTu \approx \frac{8\pi h\nu^3}{c^3} e^{-h\nu/k_BT}, which is precisely Wien’s formula.


⭐⭐ Reference Answer

Pquantum=ψL+ψR2=PLeiϕL+PReiϕR2P_{\text{quantum}} = |\psi_L + \psi_R|^2 = |\sqrt{P_L}e^{i\phi_L} + \sqrt{P_R}e^{i\phi_R}|^2

=PL+PR+2PLPRcos(ϕLϕR)= P_L + P_R + 2\sqrt{P_L P_R}\cos(\phi_L - \phi_R)

The interference term is 2PLPRcosΔϕ2\sqrt{P_L P_R}\cos\Delta\phi, with maximum magnitude 2PLPR2\sqrt{P_L P_R}. By the AM–GM inequality, 2PLPRPL+PR=12\sqrt{P_L P_R} \leq P_L + P_R = 1, with the interference term magnitude maximized (at 1) when PL=PR=1/2P_L = P_R = 1/2. When PL0P_L \to 0, 2PLPR02\sqrt{P_L P_R} \to 0, and the interference fringes vanish—consistent with physical intuition: closing one slit naturally kills the interference. When PLPRP_L \ll P_R, the interference term is small but still present; verification requires extremely precise experiments.


🗣️ Reference Answer (Key Points)

Suggested core analogy — “Tossing Coins vs Ocean Waves”:

Explain to your grandmother: Imagine you’re at the seaside, and there are two very narrow entrances through which waves can enter a pool. If you open only one entrance, the waves form one kind of ripple pattern; if you open only the other, they form another. But if you open both entrances simultaneously, the two waves superimpose—in some places crest meets crest to form a large wave (increased probability); in other places crest meets trough and they cancel out (decreased probability). An electron (or photon) behaves more like a wave than a bullet—it “simultaneously” passes through both entrances and interferes with itself.

On the “electron splitting” misconception: It is not that the electron splits; rather, the electron’s “possibility” spreads like a water wave along both paths. When you go to “look” to see which path it took (by placing a detector), it is like touching the water surface with your hand—the ripples are disturbed by you, and the interference disappears.

On “why measurement destroys the result”: Imagine you are looking for a cat in a dark room—you must turn on the light to see it. But the very act of turning on the light startles the cat. In the microscopic world, the “disturbance” of the measuring apparatus on the measured system is non-negligible. You cannot “take a gentle peek” without affecting the electron.


💭 Reference Answer (Key Points)

What Einstein objected to was not probability per se (he made important contributions to statistical mechanics himself), but rather the claim that “probability is the final, irreducible answer.” His EPR paper attempted to prove: if quantum mechanics is complete, then there must exist “spooky action at a distance”—which he considered absurd, and therefore concluded by contradiction that quantum mechanics is incomplete.

Key insight: What Einstein could not accept was the declaration that “quantum mechanics is the ultimate theory.” He believed there must be a deeper, deterministic theory behind it (a hidden-variable theory). Interestingly, Bell’s inequality (1964) proved that no local hidden-variable theory can reproduce all the predictions of quantum mechanics. Experiments (Aspect 1982, and numerous high-precision experiments since) have sided with quantum mechanics.

Regarding the Many-Worlds Interpretation: Einstein would probably still not accept it. His taste was for “concise, elegant determinism,” and Many-Worlds, while eliminating collapse, introduces countless unobservable parallel universes—which he would likely find equally unacceptable.


2.2 Quantum Mechanical Postulates

⭐ Reference Answer

(a) ZZ basis: P(0)=0ψ2=1/32=1/3P(0) = |\langle 0|\psi\rangle|^2 = |1/\sqrt{3}|^2 = 1/3

(b) XX basis: +=(0+1)/2\lvert+\rangle = (\lvert0\rangle + \lvert1\rangle)/\sqrt{2}, +ψ=12(13+23)=121+23\langle+|\psi\rangle = \frac{1}{\sqrt{2}}(\frac{1}{\sqrt{3}} + \sqrt{\frac{2}{3}}) = \frac{1}{\sqrt{2}}\cdot\frac{1+ \sqrt{2}}{\sqrt{3}}, P(+)=(1+2)265.82860.971P(+) = \frac{(1+\sqrt{2})^2}{6} \approx \frac{5.828}{6} \approx 0.971

(c) YY basis: +i=(0+i1)/2\lvert+i\rangle = (\lvert0\rangle + i\lvert1\rangle)/\sqrt{2}, +iψ=12(13i23)\langle+i|\psi\rangle = \frac{1}{\sqrt{2}}(\frac{1}{\sqrt{3}} - i\sqrt{\frac{2}{3}}), P(+i)=12(13+23)=1/2P(+i) = \frac{1}{2}(\frac{1}{3} + \frac{2}{3}) = 1/2

The sum of the three probabilities is 1/3+0.971+0.51.804>11/3 + 0.971 + 0.5 \approx 1.804 > 1. This is not a contradiction, because they are probabilities in different measurement bases—each experiment can only use one basis; it is impossible to measure all three bases simultaneously.


⭐⭐ Reference Answer

(1) A=iaiPiA = \sum_i a_i P_i, where Pi=aiaiP_i = \lvert a_i\rangle\langle a_i\rvert

(2) Born rule: P(ai)=ψPiψP(a_i) = \langle\psi|P_i|\psi\rangle

(3) Definition of expectation value: A=iaiP(ai)=iaiψPiψ\langle A\rangle = \sum_i a_i P(a_i) = \sum_i a_i \langle\psi|P_i|\psi\rangle

(4) A=ψ(iaiPi)ψ=ψAψ\langle A\rangle = \langle\psi|\left(\sum_i a_i P_i\right)|\psi\rangle = \langle\psi|A|\psi\rangle

If AA has degeneracy (suppose aia_i corresponds to a gig_i-dimensional subspace, with projection operator Pi=k=1giai(k)ai(k)P_i = \sum_{k=1}^{g_i} \lvert a_i^{(k)}\rangle\langle a_i^{(k)}\rvert), then the form of the spectral decomposition in step (1) remains unchanged (PiP_i is the projection onto the entire degenerate subspace), the measurement probability is P(ai)=ψPiψP(a_i) = \langle\psi|P_i|\psi\rangle, and the proof proceeds exactly as before.


🗣️ Reference Answer (Key Points)

A suggested game design:

Game Name: “Quantum Quest”

Rule 1 (State Space): Your character can stand at any position on the board, but cannot be at two positions simultaneously—your “possibility,” however, can. Each “square” on the board represents a possible state. You cannot stand “between squares” (states must be points on a grid).

Rule 2 (Unitary Evolution): Each turn you must move your piece according to the instructions on a card. The movement rules are “reversible”—you can always go back to the previous step (if you remember how you moved). You cannot skip a move, nor move randomly (unless the card demands it). These are the quantum gate operations.

Rule 3 (Observables): Each game event has a “scoring rule” corresponding to a specific measurement scheme. Different scoring rules correspond to different “bases.”

Rule 4 (Measurement): When you flip over a card or check your score, the outcome becomes “definite.” But before you flip, the card is in a superposition of all possible values. The act of flipping—measurement—turns uncertainty into definiteness.

Rule 5 (Composite Systems): Two players can form an “alliance.” The alliance’s state space is much larger than a simple combination of the two individuals’ states—because the two can become “entangled” and share information. This is why quantum computers need multiple qubits working in concert.


💭 Reference Answer (Key Points)

The “measurement problem” is an open question that has persisted for nearly a century. Here are several major stances:

Stance 1 — Collapse is real (Standard Copenhagen): Measurement is an irreversible process, with collapse triggered by a “classical apparatus.” The problem is: where is the boundary between “classical” and “quantum”? Von Neumann believed collapse occurs upon “consciousness”—but most physicists reject this view.

Stance 2 — Collapse is not real (Many-Worlds Interpretation): There is no collapse, only branching. With each measurement, the universe splits into multiple parallel branches, each branch perceiving one definite outcome. The question: where does “probability” come from in this framework? If all possible outcomes actually occur, why do we perceive only one?

Stance 3 — Collapse is emergent (Decoherence Theory): Entanglement between the system and the environment causes “apparent” collapse—quantum information “leaks” into environmental degrees of freedom, and from the observer’s perspective, the system behaves as a classical mixture. But decoherence only explains “why it looks like collapse,” not why we experience only one definite outcome (the “preferred basis problem”).

Stance 4 — Collapse is objective (GRW Spontaneous Localization Models): Collapse is a real physical process, occurring spontaneously with extremely low probability (about 1016s110^{-16}\,\text{s}^{-1} for a single particle), but rapidly accumulating in macroscopic systems. This theory makes experimentally testable modified predictions, though experiments to date have not observed deviations from standard quantum mechanics.

Core insight: Each interpretation supplements or reinterprets Postulate 4 differently. They all reproduce the experimental predictions of standard quantum mechanics, but disagree on “what collapse actually is.” This tells us: quantum mechanics may not be a complete theory, but rather an “effective description” of a deeper theory.


2.3 Wave Functions and the Schrödinger Equation

⭐ Reference Answer

(a) 0Lψ2(x)2dx=2L0Lsin2(2πx/L)dx=2LL2=1\int_0^L |\psi_2(x)|^2 dx = \frac{2}{L}\int_0^L \sin^2(2\pi x/L)dx = \frac{2}{L}\cdot\frac{L}{2} = 1

(b) x=2L0Lxsin2(2πx/L)dx=2L0Lx12[1cos(4πx/L)]dx\langle x\rangle = \frac{2}{L}\int_0^L x\sin^2(2\pi x/L)dx = \frac{2}{L}\int_0^L x\cdot\frac{1}{2}[1 - \cos(4\pi x/L)]dx

=1L[L22]=L2= \frac{1}{L}[\frac{L^2}{2}] = \frac{L}{2} (the cos\cos term integral vanishes because 0Lxcos(4πx/L)dx=0\int_0^L x\cos(4\pi x/L)dx = 0)

(c) x2=2L0Lx2sin2(2πx/L)dx=1L0Lx2[1cos(4πx/L)]dx\langle x^2\rangle = \frac{2}{L}\int_0^L x^2\sin^2(2\pi x/L)dx = \frac{1}{L}\int_0^L x^2[1 - \cos(4\pi x/L)]dx

=1L[L330Lx2cos(4πx/L)dx]= \frac{1}{L}[\frac{L^3}{3} - \int_0^L x^2\cos(4\pi x/L)dx]

Using integration by parts: 0Lx2cos(kx)dx=2xk2cos(kx)+(x2k2k3)sin(kx)0L\int_0^L x^2\cos(kx)dx = \frac{2x}{k^2}\cos(kx) + (\frac{x^2}{k} - \frac{2}{k^3})\sin(kx)\big|_0^L, with k=4π/Lk = 4\pi/L

At x=Lx = L: cos(kL)=cos(4π)=1\cos(kL) = \cos(4\pi) = 1, sin(kL)=0\sin(kL) = 0; at x=0x = 0: cos(0)=1\cos(0) = 1, sin(0)=0\sin(0) = 0

Substituting: 0Lx2cos(4πx/L)dx=2Lk2cos(kL)0=2L(4π/L)2=L38π2\int_0^L x^2\cos(4\pi x/L)dx = \frac{2L}{k^2}\cos(kL) - 0 = \frac{2L}{(4\pi/L)^2} = \frac{L^3}{8\pi^2}

Hence x2=1L[L33L38π2]=L2(1318π2)\langle x^2\rangle = \frac{1}{L}[\frac{L^3}{3} - \frac{L^3}{8\pi^2}] = L^2(\frac{1}{3} - \frac{1}{8\pi^2})

(d) For the ground state: Δx=L2(1318π2)(L2)2=L1318π214=L11218π20.181L\Delta x = \sqrt{L^2(\frac{1}{3} - \frac{1}{8\pi^2}) - (\frac{L}{2})^2} = L\sqrt{\frac{1}{3} - \frac{1}{8\pi^2} - \frac{1}{4}} = L\sqrt{\frac{1}{12} - \frac{1}{8\pi^2}} \approx 0.181L


⭐⭐ Reference Answer

Hint: sin(πx/L)=(eiπx/Leiπx/L)/(2i)\sin(\pi x/L) = (e^{i\pi x/L} - e^{-i\pi x/L})/(2i)

ϕ(p)=12π2L12i0L(eiπx/Leiπx/L)eipx/dx\phi(p) = \frac{1}{\sqrt{2\pi\hbar}} \sqrt{\frac{2}{L}} \frac{1}{2i} \int_0^L (e^{i\pi x/L} - e^{-i\pi x/L}) e^{-ipx/\hbar} dx

This is a Gaussian-type integral: 0Leiαxdx=(eiαL1)/(iα)\int_0^L e^{i\alpha x} dx = (e^{i\alpha L} - 1)/(i\alpha), where α=±π/Lp/\alpha = \pm \pi/L - p/\hbar.

After combining:

ϕ(p)=12iπLei(pL/π)1p/π/Lei(pL/+π)1p/+π/L\phi(p) = \frac{1}{2i\sqrt{\pi L\hbar}} \frac{e^{i(pL/\hbar - \pi)} - 1}{p/\hbar - \pi/L} - \frac{e^{-i(pL/\hbar + \pi)} - 1}{p/\hbar + \pi/L}

After simplification, one obtains the expression given in the problem. The momentum distribution ϕ(p)2|\phi(p)|^2 has a main peak near p=π/Lp = \pi\hbar/L and a symmetric secondary peak near p=π/Lp = -\pi\hbar/L, with rapidly decaying oscillatory tails on both sides. This corresponds to the momentum uncertainty Δp>0\Delta p > 0 of the ground-state wave function, satisfying the uncertainty principle ΔxΔp/2\Delta x \Delta p \geq \hbar/2.


🗣️ Reference Answer (Key Points)

Core approach for explaining to a friend:

  1. What the Schrödinger equation is saying: It describes how a “probability wave” propagates through time. Imagine a raindrop falling on a calm water surface—you see ripples spreading outward in circles. The Schrödinger equation describes how “the probability distribution of where a particle might be” evolves like a water wave.

  2. The meaning of ii: The imaginary unit ii in the equation is a crucial ingredient. If the equation were a real version, the wave function would grow or decay exponentially (unstable). But ii makes the evolution a “rotation”—just as the complex number eiωte^{i\omega t} traces a circle in the complex plane—keeping the probability distribution stable (normalization preserved). This ensures the total probability is always 1.

  3. Difference from classical waves: The classical wave equation (e.g., the sound wave equation) is real and describes “real vibrations” (compressions and rarefactions of air molecules). The Schrödinger equation describes not “real vibrations” but “vibrations of possibility.” The wave function itself is not a directly measurable physical quantity—we can only measure its squared modulus.

  4. Difference in what is described: F=maF = ma tells you a definite trajectory—given initial position and velocity, you can compute the particle’s exact position at any time. The Schrödinger equation does not tell you a definite trajectory—it tells you “the probability for each position.” This is the essential difference: the fundamental shift from determinism to probabilism.

Key analogy: One can think of the Schrödinger equation as a “weather forecast model”—it cannot tell you for certain whether it will rain at your doorstep at 3 PM tomorrow, but it can tell you the probability of rain. The difference is that probability in quantum mechanics is not due to a lack of information, but because the world is fundamentally probabilistic at its base level.


💭 Reference Answer (Key Points)

The question of the reality of the wave function touches the deepest philosophical layer of quantum mechanics.

Ontic vs Epistemic Debate:

  • Ontic (real): The wave function is a field that truly exists in the physical world. Collapse is a real physical process (nonlocal, instantaneous). The problem: if the wave function is real, why can’t we measure it directly? Why can we only measure ψ2|\psi|^2?
  • Epistemic (knowledge-based): The wave function is merely an encoding of our state of knowledge about the system. Collapse is merely an update of information—just as you update your beliefs upon learning something new. The problem: if the wave function is merely knowledge, why can the “knowledge” about two entangled particles exhibit nonlocal correlations?

Scenario A Analysis: Collapse instantaneously turns the probability at x2x_2 to zero. According to the ontic view, the wave function field at x2x_2 genuinely vanished. According to the epistemic view, no “thing” physically propagated—upon learning that the electron is at x1x_1, the probability at x2x_2 simply becomes zero (a conditional probability update).

Scenario B Analysis: This is the core of EPR correlations. The “instantaneous” collapse cannot be used to transmit information—because the observer of particle B cannot, on their own, know that collapse has occurred unless they receive a classical communication from the observer of particle A (limited to the speed of light). So there is no violation of special relativity.

Core conclusion: The nonlocality in quantum mechanics is a “nonlocality of correlations,” not a “nonlocality of causation.” Bell’s theorem shows: any theory that perfectly reproduces quantum correlations must be either “nonlocal” or “non-real” (or both). This is known as the “Bell triangle” dilemma—you must choose to abandon locality, realism, or free will. Most physicists choose to abandon realism.


2.4 Two-Level Systems and Spin

⭐ Reference Answer

(1) n=(sinπ/4cos0,sinπ/4sin0,cosπ/4)=(22,0,22)\vec{n} = (\sin\pi/4\cos0, \sin\pi/4\sin0, \cos\pi/4) = (\frac{\sqrt{2}}{2}, 0, \frac{\sqrt{2}}{2})

Sn=2(22X+22Z)=22(X+Z)S_{\vec{n}} = \frac{\hbar}{2}(\frac{\sqrt{2}}{2}X + \frac{\sqrt{2}}{2}Z) = \frac{\hbar}{2\sqrt{2}}(X + Z)

X+Z=(1111)X + Z = \begin{pmatrix} 1 & 1 \\ 1 & -1 \end{pmatrix}, so Sn=22(1111)S_{\vec{n}} = \frac{\hbar}{2\sqrt{2}}\begin{pmatrix} 1 & 1 \\ 1 & -1 \end{pmatrix}

(2) Eigenvalue equation: 22(1111)v=λv\frac{\hbar}{2\sqrt{2}}\begin{pmatrix} 1 & 1 \\ 1 & -1 \end{pmatrix}\lvert v\rangle = \lambda \lvert v\rangle, where λ=±2\lambda = \pm\frac{\hbar}{2}

Eigenstate corresponding to +/2+\hbar/2: v+=cos(π/8)0+sin(π/8)1\lvert v_+\rangle = \cos(\pi/8)\lvert0\rangle + \sin(\pi/8)\lvert1\rangle (note: θ/2=π/8\theta/2 = \pi/8 here, because the “north pole” of SnS_{\vec{n}} corresponds to 4545^\circ)

P(+/2)=v+ψ2P(+\hbar/2) = |\langle v_+|\psi\rangle|^2

v+ψ=cos(π/8)3/4+sin(π/8)120.92390.8660+0.38270.50.800+0.1910.991\langle v_+|\psi\rangle = \cos(\pi/8)\sqrt{3/4} + \sin(\pi/8)\cdot\frac{1}{2} \approx 0.9239\cdot0.8660 + 0.3827\cdot0.5 \approx 0.800 + 0.191 \approx 0.991

P(+/2)0.982P(+\hbar/2) \approx 0.982

(3) When θ=0\theta = 0, n=(0,0,1)\vec{n} = (0, 0, 1), Sn=2ZS_{\vec{n}} = \frac{\hbar}{2}Z, which reduces to the zz-direction spin measurement. P(+/2)=0ψ2=3/4P(+\hbar/2) = |\langle 0|\psi\rangle|^2 = 3/4.


⭐⭐ Reference Answer

Key step: For any unit vector n\vec{n}, (nσ)2=(nxX+nyY+nzZ)2(\vec{n}\cdot\vec{\sigma})^2 = (n_x X + n_y Y + n_z Z)^2

Expand: nx2X2+ny2Y2+nz2Z2+nxny(XY+YX)+nynz(YZ+ZY)+nznx(ZX+XZ)n_x^2 X^2 + n_y^2 Y^2 + n_z^2 Z^2 + n_x n_y (XY + YX) + n_y n_z (YZ + ZY) + n_z n_x (ZX + XZ)

Using X2=Y2=Z2=IX^2 = Y^2 = Z^2 = I and the anticommutation relations {X,Y}=XY+YX=0\{X,Y\} = XY + YX = 0 (and likewise for YZYZ and ZXZX):

=(nx2+ny2+nz2)I=I= (n_x^2 + n_y^2 + n_z^2)I = I (because n\vec{n} is a unit vector)

Hence eiα(nσ)/2=k=01k!(iα/2)k(nσ)ke^{-i\alpha(\vec{n}\cdot\vec{\sigma})/2} = \sum_{k=0}^\infty \frac{1}{k!}(-i\alpha/2)^k (\vec{n}\cdot\vec{\sigma})^k

Even powers: (nσ)2k=I(\vec{n}\cdot\vec{\sigma})^{2k} = I, odd powers: (nσ)2k+1=(nσ)(\vec{n}\cdot\vec{\sigma})^{2k+1} = (\vec{n}\cdot\vec{\sigma})

=k=0(iα/2)2k(2k)!I+k=0(iα/2)2k+1(2k+1)!(nσ)= \sum_{k=0}^\infty \frac{(-i\alpha/2)^{2k}}{(2k)!} I + \sum_{k=0}^\infty \frac{(-i\alpha/2)^{2k+1}}{(2k+1)!} (\vec{n}\cdot\vec{\sigma})

=cos(α/2)Iisin(α/2)(nσ)= \cos(\alpha/2) I - i\sin(\alpha/2)(\vec{n}\cdot\vec{\sigma})

QED.

Application (1): Rx(π)=cos(π/2)Iisin(π/2)X=iXR_x(\pi) = \cos(\pi/2)I - i\sin(\pi/2)X = -iX, which acting on quantum states is equivalent to the XX gate (up to a global phase).

(2) Ry(π/2)=cos(π/4)Iisin(π/4)Y=12(IiY)R_y(\pi/2) = \cos(\pi/4)I - i\sin(\pi/4)Y = \frac{1}{\sqrt{2}}(I - iY)

Ry(π/2)0=12(0iY0)=12(0i(i1))=12(0+1)=+R_y(\pi/2)\lvert0\rangle = \frac{1}{\sqrt{2}}(\lvert0\rangle - iY\lvert0\rangle) = \frac{1}{\sqrt{2}}(\lvert0\rangle - i(i\lvert1\rangle)) = \frac{1}{\sqrt{2}}(\lvert0\rangle + \lvert1\rangle) = \lvert+\rangle

A rotation of π/2\pi/2 about the yy-axis takes 0\lvert0\rangle (north pole) to the positive xx-axis (+\lvert+\rangle).


🗣️ Reference Answer (Key Points)

Core approach for explaining to a high-school student:

  1. Why zz-axis information is lost after measuring xx: Imagine you have a pen in your hand. You can measure its length (zz-direction information), or you can measure its thickness (xx-direction information). But these two measurements are “incompatible”—to measure the length precisely, you might need to stretch the pen, which changes its thickness. Spin is similar—zz-direction and xx-direction spin are “incompatible observables”; you cannot know both values precisely at the same time.

  2. Why you can’t simultaneously know both directions: This is not a technological limitation (insufficient measurement precision), but a principle-level restriction—the Heisenberg uncertainty principle. In classical physics, you can simultaneously know an object’s position and momentum (though measurement itself introduces error, you can in theory increase precision arbitrarily). In quantum mechanics, there exist hard minimum-uncertainty constraints between certain pairs of observables, independent of measurement technology.

  3. Spin vs classical rotation: Classical angular momentum is “continuous” (can take any direction, any magnitude), and the components along three directions can be known simultaneously. Spin has only one fixed magnitude (/2\hbar/2), can take only ±/2\pm\hbar/2 along any direction, and the three directional components cannot be simultaneously determined. Spin is “quantum”—not rotation in the classical sense.

  4. Stern–Gerlach experiment analogy: Imagine a spinning billiard ball. When it passes through an inhomogeneous magnetic field, if its spin direction aligns with the field, it deflects one way; if opposite, it deflects the other way. Classical physics expects: the billiard ball can spin at any angle (just as the Earth’s rotation axis can point in any direction), so the deflection should be a continuous distribution. But the experiment found only two spots—this means the electron’s spin, when facing a magnetic field, has only two possible orientations. It is like a compass that can only point “north” or “south”—unable to point “northeast.”


💭 Reference Answer (Key Points)

  1. Origin of spin in non-relativistic QM: Spin in the non-relativistic Schrödinger equation was “put in by hand”—Pauli in 1927 introduced a two-component wave function and Pauli matrices to “cobble together” spin. In the non-relativistic framework, spin is an additional assumption with no deep justification. Relativistic quantum mechanics is different: the Dirac equation “automatically” produces spin—it is a requirement of covariance, not an artificial addition.

  2. The deep connection of the spin-statistics theorem: In relativistic quantum field theory, the spin-statistics theorem is a rigorous mathematical result. It arises from the combination of causality and the positive-energy condition. If one tries to place two fermions in the same quantum state, the constructed field operators would violate causality (unacceptable forms of nonlocality would appear). This hints that: the connection between spin and exchange symmetry is a property of spacetime itself, not an accident.

  3. Pauli matrices and quaternions: This is not a coincidence. The quaternions {1,i,j,k}\{1, i, j, k\} satisfy i2=j2=k2=1i^2 = j^2 = k^2 = -1 and ij=k,ji=kij = k, ji = -k, etc. The Pauli matrices (multiplied by ii) give a 2×22\times2 complex matrix representation of the quaternions: iσx,iσy,iσzi\sigma_x, i\sigma_y, i\sigma_z are isomorphic to the quaternion units i,j,ki, j, k. This indicates that the mathematical structure of spin is more deeply rooted in the “algebraic theory of rotations”—quaternions were invented precisely to describe rotations in three-dimensional space. Spinors are the double-valued representation of the rotation group SO(3)—a 360360^\circ rotation does not bring you back to the original state (it takes 720720^\circ to return), and this counterintuitive topological property is the origin of many “strange” phenomena in quantum mechanics.


2.5 The Bloch Sphere

⭐ Reference Answer

(1) A\lvert A\rangle: cos(θ/2)=3/2\cos(\theta/2) = \sqrt{3}/2, sin(θ/2)=1/2\sin(\theta/2) = 1/2θ=π/3\theta = \pi/3, ϕ=0\phi = 0(x,y,z)=(sinπ/3,0,cosπ/3)=(3/2,0,1/2)(x,y,z) = (\sin\pi/3, 0, \cos\pi/3) = (\sqrt{3}/2, 0, 1/2)

Northern hemisphere, +x+x direction.

(2) B\lvert B\rangle: θ/2=π/4\theta/2 = \pi/4θ=π/2\theta = \pi/2, ϕ=π/3\phi = \pi/3(x,y,z)=(sinπ/2cosπ/3,sinπ/2sinπ/3,0)=(1/2,3/2,0)(x,y,z) = (\sin\pi/2\cos\pi/3, \sin\pi/2\sin\pi/3, 0) = (1/2, \sqrt{3}/2, 0)

On the equator.

(3) C\lvert C\rangle: cos(θ/2)=1/5\cos(\theta/2) = 1/\sqrt{5}, sin(θ/2)=2/5\sin(\theta/2) = 2/\sqrt{5}θ2arccos(1/5)21.1072.214\theta \approx 2\arccos(1/\sqrt{5}) \approx 2\cdot 1.107 \approx 2.214 rad → southern hemisphere (θ>π/2\theta > \pi/2), ϕ=0\phi = 0(x,y,z)(sin2.214,0,cos2.214)(0.8,0,0.6)(x,y,z) \approx (\sin 2.214, 0, \cos 2.214) \approx (0.8, 0, -0.6)

(4) D\lvert D\rangle: D=110(0+3i1)\lvert D\rangle = \frac{1}{\sqrt{10}}(\lvert0\rangle + 3i\lvert1\rangle), cos(θ/2)=1/10\cos(\theta/2) = 1/\sqrt{10}, sin(θ/2)=3/10\sin(\theta/2) = 3/\sqrt{10}θ2arccos(1/10)2.498\theta \approx 2\arccos(1/\sqrt{10}) \approx 2.498 rad, ϕ=π/2\phi = \pi/2(x,y,z)(sin2.498cosπ/2,sin2.498sinπ/2,cos2.498)(0,0.949,0.316)(x,y,z) \approx (\sin2.498\cos\pi/2, \sin2.498\sin\pi/2, \cos2.498) \approx (0, 0.949, -0.316)

Southern hemisphere, +y+y direction.

Spherical distance: A\lvert A\rangle and B\lvert B\rangle are relatively close in both θ\theta and ϕ\phi—the central angle γ\gamma is given by cosγ=cosθAcosθB+sinθAsinθBcos(ϕAϕB)\cos\gamma = \cos\theta_A\cos\theta_B + \sin\theta_A\sin\theta_B\cos(\phi_A - \phi_B).


⭐⭐ Reference Answer

A general SU(2)SU(2) matrix is U=(abba)U = \begin{pmatrix} a & b \\ -b^* & a^* \end{pmatrix}, with a2+b2=1|a|^2 + |b|^2 = 1.

Rz(α)=eiαZ/2=(eiα/200eiα/2)R_z(\alpha) = e^{-i\alpha Z/2} = \begin{pmatrix} e^{-i\alpha/2} & 0 \\ 0 & e^{i\alpha/2} \end{pmatrix}

Rx(β)=eiβX/2=(cos(β/2)isin(β/2)isin(β/2)cos(β/2))R_x(\beta) = e^{-i\beta X/2} = \begin{pmatrix} \cos(\beta/2) & -i\sin(\beta/2) \\ -i\sin(\beta/2) & \cos(\beta/2) \end{pmatrix}

Rz(γ)Rx(β)Rz(α)=(eiγ/200eiγ/2)(cos(β/2)isin(β/2)isin(β/2)cos(β/2))(eiα/200eiα/2)R_z(\gamma) R_x(\beta) R_z(\alpha) = \begin{pmatrix} e^{-i\gamma/2} & 0 \\ 0 & e^{i\gamma/2} \end{pmatrix} \begin{pmatrix} \cos(\beta/2) & -i\sin(\beta/2) \\ -i\sin(\beta/2) & \cos(\beta/2) \end{pmatrix} \begin{pmatrix} e^{-i\alpha/2} & 0 \\ 0 & e^{i\alpha/2} \end{pmatrix}

=(ei(α+γ)/2cos(β/2)iei(αγ)/2sin(β/2)iei(γα)/2sin(β/2)ei(α+γ)/2cos(β/2))= \begin{pmatrix} e^{-i(\alpha+\gamma)/2}\cos(\beta/2) & -i e^{i(\alpha-\gamma)/2}\sin(\beta/2) \\ -i e^{i(\gamma-\alpha)/2}\sin(\beta/2) & e^{i(\alpha+\gamma)/2}\cos(\beta/2) \end{pmatrix}

Set a=ei(α+γ)/2cos(β/2)a = e^{-i(\alpha+\gamma)/2}\cos(\beta/2), b=iei(αγ)/2sin(β/2)b = -i e^{i(\alpha-\gamma)/2}\sin(\beta/2).

Given arbitrary a,ba, b (with a2+b2=1|a|^2 + |b|^2 = 1): β=2arccosa=2arcsinb\beta = 2\arccos|a| = 2\arcsin|b| α+γ=2arg(a)\alpha + \gamma = -2\arg(a) αγ=2arg(b)π\alpha - \gamma = -2\arg(b) - \pi (since i=eiπ/2-i = e^{-i\pi/2})

These two equations uniquely determine α\alpha and γ\gamma. Thus any SU(2)SU(2) matrix (hence any single-qubit unitary gate, up to a global phase) can be expressed as a Z–X–Z Euler angle decomposition. In superconducting qubits, this corresponds to the hardware implementation scheme of “virtual Z gate + microwave pulse X rotation.”


🗣️ Reference Answer (Key Points)

Explaining the Bloch sphere to a high-school student:

  1. Why it is a sphere: Two complex numbers α,β\alpha, \beta have 4 real parameters, but the normalization condition removes 1, the global phase removes 1, leaving 2 free parameters. The set of two parameters happens to correspond to all points on a spherical surface—just as a point on the Earth’s surface is determined by two parameters, latitude and longitude.

  2. Why the north and south poles are 0\lvert0\rangle and 1\lvert1\rangle: This is merely a convention. If you take a compass, “purely north” is 0\lvert0\rangle and “purely south” is 1\lvert1\rangle. Any other direction is a mixture of the two. Just as your position on the equator is “half north and half south”—but it is not a classical mixture of “north” and “south,” but a brand-new kind of “superposed direction.”

  3. Why relative phase determines longitude: Imagine a pendulum swinging in the xx-axis direction (+\lvert+\rangle), and another swinging in the yy-axis direction (+i\lvert+i\rangle)—their swing amplitudes are exactly the same, but their “phases” differ by 9090^\circ (or in time, by 1/4 of a period). On the Bloch sphere, this time difference is mapped to a longitude difference.

  4. Rotations correspond to quantum gates: The most beautiful aspect of the Bloch sphere: any quantum gate (described by a unitary matrix) corresponds to a rigid rotation of the sphere. A rotation about the xx-axis takes the north pole to the south pole—that is the XX gate (bit flip). A rotation about the zz-axis changes the phase—that is the ZZ gate (phase flip).

  5. Mixed states are inside the sphere: Pure states are points on the surface of the sphere—you know exactly which state the system is in. Mixed states are points inside the sphere—you don’t know the exact state, you only know the probability distribution. The center of the sphere corresponds to “complete ignorance”—the measurement outcome in any direction is completely random. From the center to the surface, your knowledge about the quantum state becomes increasingly precise.

On the θ/2\theta/2 explanation: When we rotate 180180^\circ on the sphere (from north pole to south pole), the corresponding quantum state parameter θ\theta goes through 180180^\circ, so in the parameterization θ/2\theta/2 goes from 00 to π/2\pi/2. A 360360^\circ sphere rotation (back to the starting point) corresponds to a 720720^\circ quantum state rotation! This is a spinor property—a spin-1/2 system rotated by 360360^\circ does not return to its original state, but acquires a 1-1 global phase. This sounds bizarre, but has been directly verified in neutron interference experiments.


💭 Reference Answer (Key Points)

  1. Multi-qubit systems cannot be reduced to multiple Bloch spheres: This is because of the existence of entangled states. If one attempts to describe the Bell state Φ+\lvert\Phi^+\rangle using “two Bloch spheres,” each subsystem’s reduced density matrix is I/2I/2 (the center of the sphere), completely losing the correlation information. The Cartesian product of the two centers contains all possibilities, but loses the information of which particular entangled pure state it is. Many-body quantum states require description in a high-dimensional space (CP2n1^{2^n-1}), whose geometric structure is far more complex than a sphere.

  2. The deep relationship between geometry and information: The fact that “the Bloch sphere loses entanglement information” suggests: our intuitive geometric understanding of quantum states fails in the face of entanglement. This motivates researchers to develop new geometric languages—such as entanglement entropy, tensor networks, and quantum information geometry—to capture the rich structure of many-body quantum states.

  3. The geometry of many-body mixed states: The density matrix of nn qubits constitutes a (4n1)(4^n-1)-dimensional convex space—the state space. The boundary of this space is not a smooth sphere, but a complex convex body with numerous “sharp points” and “flat regions.” The relationships between pure states (boundary points) and mixed states (interior points) are far more complex than in the two-dimensional case. The geometric classification of many-body quantum states is an active research direction in contemporary quantum information theory.

  4. Living inside the Bloch sphere: This question touches on “the boundary between observer and observed system.” If we ourselves are also quantum systems, the information accessible to us is constrained by the limits of quantum mechanics itself—this is the core subject of quantum epistemology. Are there “hidden” quantum degrees of freedom that no experiment can ever probe? This is the question that quantum metrology seeks to answer.


2.6 Measurement Theory and Density Matrices

⭐ Reference Answer

(a) ++=12(1111)\lvert+\rangle\langle+\rvert = \frac12\begin{pmatrix}1&1\\1&1\end{pmatrix}, =12(1111)\lvert-\rangle\langle-\rvert = \frac12\begin{pmatrix}1&-1\\-1&1\end{pmatrix}

ρ(p)=p12(1111)+(1p)12(1111)=12(12p12p11)\rho(p) = p\cdot\frac12\begin{pmatrix}1&1\\1&1\end{pmatrix} + (1-p)\cdot\frac12\begin{pmatrix}1&-1\\-1&1\end{pmatrix} = \frac12\begin{pmatrix}1&2p-1\\2p-1&1\end{pmatrix}

(b) Tr[ρ(p)]=12(1+1)=1\operatorname{Tr}[\rho(p)] = \frac12(1 + 1) = 1

ρ(p)2=14(12p12p11)2=14((1)2+(2p1)2(1)(2p1)+(2p1)(1)(2p1)(1)+(1)(2p1)(2p1)2+(1)2)\rho(p)^2 = \frac14\begin{pmatrix}1&2p-1\\2p-1&1\end{pmatrix}^2 = \frac14\begin{pmatrix}(1)^2+(2p-1)^2 & (1)(2p-1)+(2p-1)(1)\\(2p-1)(1)+(1)(2p-1) & (2p-1)^2+(1)^2\end{pmatrix}

=14(1+(2p1)22(2p1)2(2p1)1+(2p1)2)= \frac14\begin{pmatrix}1+(2p-1)^2 & 2(2p-1)\\2(2p-1) & 1+(2p-1)^2\end{pmatrix}

Tr[ρ(p)2]=14[1+(2p1)2+1+(2p1)2]=1+(2p1)22\operatorname{Tr}[\rho(p)^2] = \frac14[1+(2p-1)^2 + 1+(2p-1)^2] = \frac{1+(2p-1)^2}{2}

(c) Pure state: Tr(ρ2)=1    1+(2p1)2=2    (2p1)2=1    p=0\operatorname{Tr}(\rho^2) = 1 \implies 1 + (2p-1)^2 = 2 \implies (2p-1)^2 = 1 \implies p = 0 or p=1p = 1

Maximal mixedness: (2p1)2(2p-1)^2 minimized → p=1/2p = 1/2, at which point Tr(ρ2)=1/2\operatorname{Tr}(\rho^2) = 1/2

(d) Comparison with “probability pp of 0\lvert0\rangle, probability 1p1-p of 1\lvert1\rangle”: that density matrix is (p001p)\begin{pmatrix}p&0\\0&1-p\end{pmatrix}

They are equal if and only if 12=p\frac12 = p and 12=1p\frac12 = 1-p, i.e., p=1/2p = 1/2, when both equal I/2I/2.

For other values of pp, they differ: e.g., for p=1p=1, the former is ++\lvert+\rangle\langle+\rvert (pure state, XX measurement yields +1+1 with probability 1), while the latter is 00\lvert0\rangle\langle0\rvert (pure state, ZZ measurement yields +1+1 with probability 1, but XX measurement gives 50-50). The two behave completely differently in the XX basis.


⭐⭐ Reference Answer

(a) Ψ1=12(00+01)=012(0+1)=0+\lvert\Psi_1\rangle = \frac{1}{\sqrt{2}}(\lvert00\rangle + \lvert01\rangle) = \lvert0\rangle \otimes \frac{1}{\sqrt{2}}(\lvert0\rangle + \lvert1\rangle) = \lvert0\rangle \otimes \lvert+\rangle

This is a separable state (product state); there is no entanglement.

Ψ2=12(00+11)\lvert\Psi_2\rangle = \frac{1}{\sqrt{2}}(\lvert00\rangle + \lvert11\rangle) is the famous Bell state, inseparable (entangled), as there is no factorization ψAϕB\lvert\psi\rangle_A\otimes\lvert\phi\rangle_B.

(b) For Ψ1\lvert\Psi_1\rangle:

ρA(1)=TrB(Ψ1Ψ1)=TrB[(00)(++)]=00Tr(++)=00\rho_A^{(1)} = \operatorname{Tr}_B(\lvert\Psi_1\rangle\langle\Psi_1\rvert) = \operatorname{Tr}_B[(\lvert0\rangle\langle0\rvert) \otimes (\lvert+\rangle\langle+\rvert)] = \lvert0\rangle\langle0\rvert \cdot \operatorname{Tr}(\lvert+\rangle\langle+\rvert) = \lvert0\rangle\langle0\rvert

For Ψ2\lvert\Psi_2\rangle: from the main tutorial in Section 2.6, ρA(2)=I/2\rho_A^{(2)} = I/2

(c) Tr[(ρA(1))2]=Tr(00)=1\operatorname{Tr}[(\rho_A^{(1)})^2] = \operatorname{Tr}(\lvert0\rangle\langle0\rvert) = 1 (pure state)

Tr[(ρA(2))2]=Tr(I/4)=1/2\operatorname{Tr}[(\rho_A^{(2)})^2] = \operatorname{Tr}(I/4) = 1/2 (mixed state)

For Ψ1\lvert\Psi_1\rangle (separable state), ρA\rho_A is pure; for Ψ2\lvert\Psi_2\rangle (entangled state), ρA\rho_A is mixed. In this specific example, “the reduced density matrix is mixed” can serve as a criterion for entanglement. The two are consistent.

(d) For bipartite pure states: entangled state ⟺ ρA\rho_A is mixed (Tr(ρA2)<1\operatorname{Tr}(\rho_A^2) < 1). Conversely, if ρA\rho_A is pure, then ΨAB\lvert\Psi\rangle_{AB} must be a separable state (product state).

However, for mixed states (when the overall state is mixed), this criterion is no longer necessary and sufficient—there exist “separable mixed states” for which ρA\rho_A is also mixed. Distinguishing entangled mixed states from separable mixed states is an NP-hard problem, highlighting the complex geometry of quantum entanglement in the mixed-state case.


🗣️ Reference Answer (Key Points)

Core approach for explaining to Anna and Bob:

  1. How to distinguish in the XX basis: Anna’s superposition state +=(0+1)/2\lvert+\rangle = (\lvert0\rangle + \lvert1\rangle)/\sqrt{2} is an eigenstate of XX—measuring XX always yields +1+1. Bob’s classical mixture under the XX basis gives 50% probability each—it can never yield the same result every time. This is the key experimental difference: the statistical behavior differs in different bases.

  2. Everyday analogy — pendulums: Imagine two pendulums swinging in perfect synchrony—their phase relationship is fixed (this is “coherence”). This is the analogue of quantum superposition. Now imagine two independent pendulums swinging randomly—there is no fixed phase relationship between them (this is “classical mixture”). A coherent system can produce interference effects (constructive and destructive); an incoherent system only produces simple probabilistic addition.

  3. Importance for quantum computing: If qubits were merely bits in the “classical probability” sense, a quantum computer could not possibly be faster than a classical computer. Because classical probability superposition produces no interference—without interference, there is no quantum algorithmic speedup. The “secret weapon” of quantum computers is quantum coherence—the ability for the complex amplitudes of different “computational paths” to interfere, amplifying correct answers and suppressing incorrect ones.

  4. Unified description via the density matrix: The density matrix ρ=ipiψiψi\rho = \sum_i p_i\lvert\psi_i\rangle\langle\psi_i\rvert describes, within a single mathematical object, both quantum superposition (coherence among the ψi\lvert\psi_i\rangle) and classical uncertainty (the probabilities pip_i). Different preparation procedures can yield the same ρ\rho (e.g., uniformly mixing 0,1\lvert0\rangle, \lvert1\rangle and uniformly mixing +,\lvert+\rangle, \lvert-\rangle both give I/2I/2)—meaning that no experiment can distinguish between these two preparation schemes; for the physical world, they are “the same.”


💭 Reference Answer (Key Points)

  1. Schrödinger’s cat and decoherence: Decoherence theory explains: the inevitable entanglement of macroscopic objects with their environment causes quantum coherence to be transferred to environmental degrees of freedom at extremely fast rates (on the order of 102010^{-20} seconds), so that from the observer’s perspective, the system behaves as a classical mixture. But decoherence does not resolve the “preferred basis problem”—why do we experience one outcome, dead|{\text{dead}}\rangle or alive|{\text{alive}}\rangle, rather than a mixture of the two? So the “measurement problem” has not been “solved,” merely “pushed back”—from “when does collapse occur?” to “when does branching occur?”

  2. Entanglement nonlocality and information transmission: Quantum teleportation does indeed transmit a quantum state, but this transmission requires a classical communication channel as well—Alice must tell Bob her measurement result via a classical channel (speed ≤ cc). Without this classical information, Bob’s particle looks completely random. So the “information” transmission speed does not exceed the speed of light.

The nonlocality of entanglement is manifest in: the correlations between the two particles are nonlocal—Alice’s choice of measurement affects the statistics of Bob’s measurement outcomes (as can be seen when the results are later compared). But looking at Bob’s measurement outcomes alone, no “signal” is present. This is “nonlocality of correlations” ≠ “nonlocality of causation.”

  1. The complexity of entanglement measures: For mixed states and many-body systems, there exist multiple inequivalent entanglement measures because they capture different aspects of entanglement:
  • Entanglement of Formation: How much pure-state entanglement is needed to prepare this mixed state?
  • Distillable Entanglement: How much pure-state entanglement can be extracted from this mixed state?
  • For pure states, the two are equal; for mixed states, they are generally not equal—the “input” and “output” amounts of entanglement can differ, a manifestation of “entanglement loss” in quantum information processing.
  1. Entanglement and information may be two sides of the same coin: This is a deep theme of quantum information theory. The fact of quantum error correction—that information can be encoded in entanglement to resist decoherence—suggests that entanglement is not the enemy of information, but its protector. From the “It from Qubit” perspective, the laws of physics (including spacetime and gravity) may all be emergent phenomena of quantum information processing. The striking similarity between entanglement entropy and black hole entropy (Bekenstein–Hawking entropy)—the area law—hints that this connection may lead to quantum gravity.

Open-Ended Questions

The following questions are open-ended with no standard answers. They are intended to stimulate thought, promote discussion, and help you build quantum mechanical intuition while being aware that many unresolved philosophical and physical questions remain at the depths of this discipline.

  1. Einstein insisted that “God does not play dice,” but the experiments on Bell’s inequality appear to side with Bohr. If Schrödinger had not proposed “Schrödinger’s cat” after the 1935 EPR paper, how would the public understanding and philosophical discussion of quantum mechanics have differed? Has the “cat” thought experiment played a disproportionately large role in shaping the popular impression of “quantum = weird”?

  2. Feynman once said: “I think I can safely say that nobody understands quantum mechanics.” Yet quantum mechanics is the most precisely experimentally tested theory in physics—some quantum electrodynamics predictions agree with experiment to the level of 101210^{-12}. If a theory is this precise, why do its interpretational problems not matter? Conversely, if a theory can be this precise without telling us “what the world actually is,” should physics be satisfied with that?

  3. The basic operations of a quantum computer (unitary evolution + measurement) follow the postulates of quantum mechanics exactly. But when we say “a quantum computer is faster than a classical computer,” we are essentially saying “there exist quantum algorithms with lower complexity for certain computational problems.” What is the source of this “speedup”? Is it entanglement? Interference? The complex-number nature of probability amplitudes? A combination of these factors? If you were to capture the essence of “quantum speedup” in a single concise physical picture, how would you describe it?

  4. The partial trace operation on the density matrix “discards” the information of subsystem B. In quantum computing, this means that when one of the entangled qubits interacts with the environment and decoheres, we permanently lose some quantum information. But physical laws (such as CPT symmetry) preserve information at a fundamental level. So is the “information loss” in decoherence a genuine loss, or has the information merely been transferred from accessible degrees of freedom to inaccessible environmental degrees of freedom? How does the answer to this question relate to the black hole information paradox?

  5. If quantum mechanics is the “correct” fundamental theory, why is classical physics so effective? That is, why do macroscopic objects composed of N1023N \approx 10^{23} particles obey Newtonian mechanics rather than the Schrödinger equation? Is decoherence the sole answer? Could there exist a “phase transition” boundary from quantum to classical—above some critical scale or critical amount of entanglement, quantum behavior “spontaneously” freezes into classical behavior? Can this boundary be experimentally tested?