Sound, Magnets, Speakers and Phone Calls
9.0 What this chapter gives you#
- You will be able to say what sound physically is, with numbers.
- You will be able to draw a speaker and name every part in it.
- You will be able to explain why a magnet and a coil of wire move air.
- You will be able to explain a microphone as a speaker run backwards.
- You will be able to turn a wave into numbers and back, and say why 44,100 samples per second was chosen.
- You will be able to calculate the size of a song from first principles.
- You will be able to say what a codec throws away and why you do not hear it.
- You will be able to trace a phone call from 1876 copper to a modern internet call, and say what changed at each step.
- You will be able to explain why voice tolerates lost data but hates delay, and why file transfer is the exact opposite.
- You will be able to explain noise cancelling without any magic.
9.1 What sound actually is#
PLAIN9.1.1 in simple words#
- Sound is air being pushed and pulled.
- When something vibrates, it squashes the air next to it, and that squashed patch pushes the next patch, and so on outward.
- A squashed patch is a compression. A stretched patch, with less air in it, is a rarefaction.
- So sound is a travelling pattern of “more pressure, less pressure”.
- Nothing travels from the drum to your ear except the pattern. The air particles jiggle back and forth in place.
- Cycles per second is the frequency, in hertz (Hz). It sets the pitch.
- How hard the squeeze is, is the amplitude. It sets the loudness.
- A young ear hears roughly 20 Hz to 20,000 Hz. The top falls with age.
- Sound moves at about 343 metres per second, which is slow. That is why you see lightning long before you hear thunder.
PLAIN9.1.2 a picture in your head#
- Picture a line of people packed shoulder to shoulder in a corridor.
- Push the first person. They bump the next, who bumps the next.
- A bump travels the whole corridor, but nobody walked anywhere. Each person leaned forward and came back.
- Push rhythmically twice a second and bumps leave twice a second: 2 Hz.
- Push harder for amplitude, faster for frequency. The gap between bumps is the wavelength.
Where this comparison breaks: real air particles are not in a line and are not touching. They already fly about randomly at hundreds of metres per second. Sound is a small organized pressure change riding on that permanent chaos. And the corridor carries the wave one way, while real sound spreads in a sphere.
PLAIN9.1.3 a worked example#
- Take the note A above middle C, 440 Hz by convention.
- Wavelength = speed divided by frequency. 343 / 440 = 0.78 m.
- At the bottom of hearing, 343 / 20 = 17.15 m, longer than a bus.
- At the top, 343 / 20,000 = 0.017 m, about 17 mm.
- Across your hearing range, wave sizes differ by a factor of 1,000.
- That single fact explains most of loudspeaker design. Hold on to it.
| Frequency | Wavelength in air | About the size of |
|---|---|---|
| 20 Hz | 17.15 m | A five-storey building |
| 100 Hz | 3.43 m | A small room |
| 1,000 Hz | 34.3 cm | A forearm |
| 10,000 Hz | 3.43 cm | A thumb |
| 20,000 Hz | 1.72 cm | A fingernail |
PLAIN9.1.4 what is really happening inside#
- Loudness is measured in decibels (dB), on a scale that is odd on purpose.
- The ear responds to ratios, not to absolute energy. Whisper to shout is about a million times in pressure, and taking the logarithm turns that into 120.
- Every extra 6 dB doubles the pressure. Every extra 10 dB multiplies power by ten, and sounds roughly twice as loud.
- Zero decibels is not silence. It is a chosen reference near the quietest sound a good young ear can detect.
- Now the ear. The outer ear funnels pressure down the canal to the eardrum, and three tiny bones pass that movement on and strengthen it.
- The last bone pushes fluid inside a spiral tube, the cochlea.
- Inside is a strip that is stiff at one end and floppy at the other. High notes shake the stiff end, low notes the floppy end, so the cochlea sorts frequency by physical shape alone.
- Hair cells on that strip bend as it moves and fire a nerve signal. Pressure has become nerve pulses.
TECHNICAL9.1.5 the engineer’s version#
- Sound in air is a longitudinal wave: particle motion is parallel to travel, unlike light, which is transverse.
- Speed in dry air: c = 331.3 + 0.606 T m/s, T in degrees Celsius, so 343.4 m/s at 20 C. It depends on temperature, not pressure.
- Sound Pressure Level: SPL = 20 log10 (p / p0), with p0 = 20 micropascals RMS. That reference is a standard, set near threshold at 1 kHz. Intensity level uses 10 log10 (I / I0) with I0 = 1e-12 W/m^2: 10 for power, 20 for pressure.
- Characteristic acoustic impedance of air at 20 C is about 413 Pa*s/m.
- Equal-loudness contours, standardized in ISO 226 (2003 revision), peak near 3 to 4 kHz, matching the quarter-wave resonance of the ear canal. Age-related high-frequency loss is called presbycusis.
- The cochlea is about 35 mm long uncoiled, with roughly 3,500 inner hair cells as sensors and about 12,000 outer hair cells acting as an active amplifier. The frequency-to-place mapping is called tonotopy.
| Sound | SPL | Pressure (RMS) |
|---|---|---|
| Threshold of hearing | 0 dB | 20 uPa |
| Normal speech at 1 m | 60 dB | 20 mPa |
| Busy road | 80 dB | 200 mPa |
| Rock concert, front | 110 dB | 6.3 Pa |
| Threshold of pain | 130 dB | 63 Pa |
The honest version: “20 Hz to 20 kHz” is a convention, not a wall. Sensitivity fades gradually at both ends. Below about 20 Hz you feel pressure rather than hear a pitch, and the upper figure describes young ears in a laboratory.
WORDS9.1.6 remember these#
- Compression — a squashed patch of air — pressure above ambient.
- Rarefaction — a stretched patch of air — pressure below ambient.
- Frequency — how fast it wobbles — cycles per second, unit hertz.
- Amplitude — how big the wobble is — peak deviation from ambient pressure.
- Wavelength — the length of one wobble — lambda = c / f, in metres.
- Decibel — a loudness step — 20 log10 of a pressure ratio against 20 uPa.
- Cochlea — the spiral in your ear — the organ performing a mechanical frequency analysis on the basilar membrane.
9.2 The speaker, in full#
PLAIN9.2.1 in simple words#
- A speaker turns an electrical wiggle into an air wiggle, using four things.
- A permanent magnet, a lump of magnetized metal that always pulls.
- A voice coil, thin wire wound into a tube, sitting in the magnet’s gap.
- A cone of stiff paper or plastic, glued to that coil.
- A suspension of rubber and cloth, holding it centred so it can only move in and out.
- Send electricity through the coil and the coil becomes a magnet itself.
- Now there are two magnets. They push apart or pull together depending on which way the current flows, so reversing the current reverses the force.
- The coil is glued to the cone, so the cone moves out, then in, squashing the air in front of it and then stretching it. That is a pressure wave.
- If the signal wiggles 440 times a second, so does the cone, and you hear a 440 Hz note.
- The speaker knows nothing about music. It copies the shape of an electrical signal into the shape of a pressure wave.
PLAIN9.2.2 a picture in your head#
- Picture standing in a swimming pool holding a large flat tray.
- Push it forward and you shove a wall of water away. Pull it back and water rushes in to fill the space.
- Do that rhythmically and you make waves. Faster gives closer waves, further gives taller ones.
- The tray is the cone, the water is the air, your arms are coil and magnet.
- A big slow wave needs a big tray and a long movement, while tiny fast ripples want a small light paddle. That is why a speaker box has a big driver for bass and a small one for treble.
Where this comparison breaks: water is nearly incompressible, while the air wave exists precisely because air can be squashed. And your arms overpower the water, whereas a real speaker turns only about 0.5 to 2 percent of its electrical power into sound. The rest becomes heat in the coil.
PLAIN9.2.3 a worked example#
- A bookshelf speaker is marked 8 ohms and 87 dB.
- The 87 dB means: put in 1 watt, stand 1 metre away, measure 87 dB.
- One watt into 8 ohms needs 2.83 volts, because power = volts squared divided by resistance, and 2.83 squared is 8.
- Since +10 dB needs ten times the power, 10 W gives 97 dB and 100 W gives 107 dB.
- So going from 87 to 107 dB needs one hundred times the amplifier power.
- Swap in a 90 dB speaker and it reaches 97 dB on 5 W instead of 10 W. A sensitive speaker buys you more than a big amplifier does.
power for a target level, 8 ohm speaker
sensitivity 87 dB at 1 W / 1 m, target 100 dB at 1 m
difference = 100 - 87 = 13 dB
power ratio = 10 ^ (13 / 10) = 19.95
power needed = 19.95 W
voltage = sqrt(19.95 * 8) = 12.6 V RMS
current = 12.6 / 8 = 1.58 A RMS
PLAIN9.2.4 what is really happening inside#
- Name the parts from the back forwards. The magnet is a ring or slug of ferrite or neodymium.
- Steel shaped into a top plate and a pole piece squeezes its field into a narrow ring-shaped gap, where the field points straight outward.
- The voice coil is wound on a stiff tube, the former, and sits there.
- The spider, a corrugated cloth disc, keeps the coil centred, and the surround, a rubber ring, holds the cone edge. Together they are the suspension and act as a return spring.
- The cone is glued to the former, a dust cap covers the middle, and the assembly bolts to a basket.
- Physics law one, the Lorentz force: a wire carrying current in a magnetic field feels a sideways push, equal to field times current times wire length.
- Because the field points outward all round the gap, and the wire runs round the gap, that push is straight forwards or backwards. Nothing is wasted.
- Double the current and you double the force, so the cone faithfully copies the signal.
- Physics law two, Faraday’s law: a wire moving in a magnetic field generates a voltage. A moving cone therefore opposes its own drive, called back-EMF, and that is why a speaker also works, badly, as a microphone.
- Loudness depends on the volume of air moved: cone area times travel. For the same loudness, dropping an octave needs four times the travel.
- A 20 cm cone moving 5 mm shifts plenty of air but is too heavy to reverse 15,000 times a second. A 25 mm dome is light enough but shifts almost no air, so it makes no bass.
- Hence two or three drivers, with a crossover splitting the signal.
- Finally the box. When the cone pushes forward it pulls backward on the air behind it, producing an exactly opposite wave.
- At high frequencies the wave is short and the cone body blocks the path. At low frequencies the wave is metres long, wraps round, and cancels.
- A naked bass driver held in your hand makes almost no bass. The box exists to stop the back wave reaching the front.
cross-section of one driver (side view)
surround
|
+-------v---------------------+
\ / <- cone
\ dust cap /
\ ___ /
\ | | /
+====|===|==========+ <- spider
| | |
| | | <- voice coil on former
+----| | |----+
| N |_|_| N | <- gap, field points
| [pole ] | outward all round
| [piece] |
+-------------+ <- permanent magnet
TECHNICAL9.2.5 the engineer’s version#
- Force on the coil: F = B * l * I, with B the gap flux density in tesla, l the conductor length in the gap in metres, I the current in amperes. The product Bl, in tesla metres, is the driver’s motor strength. Back-EMF is V = Bl * u, with u the cone velocity in metres per second.
- Nominal impedance of 4, 6, 8 or 16 ohms is a convention, not a measured constant. An “8 ohm” driver measures about 5.5 to 6.5 ohms DC (Re) and swings from roughly 4 ohms to over 40 ohms across the band. The peak sits at free-air resonance Fs, where the coil moves fastest and generates the most back-EMF.
- Sensitivity is stated as dB SPL at 1 m for 2.83 V RMS, which is 1 W only into 8 ohms. Into 4 ohms that is 2 W, so 4 ohm figures flatter the driver by 3 dB.
- Typical values: 84 to 88 dB for a sealed bookshelf, 90 to 96 dB for a floorstander, 100 to 110 dB for a horn-loaded professional cabinet. Electroacoustic efficiency is 0.5 to 2 percent for a direct radiator.
- Low-frequency behaviour is described by the Thiele/Small parameters, developed by Neville Thiele in 1961 and extended by Richard Small in the early 1970s: Fs, Qts, Vas, Sd, Xmax, Re, Le and Mms. Volume displacement Vd = Sd * Xmax, and for constant SPL required displacement rises as 1/f^2.
- Crossovers are LC filters, passive after the amplifier or active before it, with slopes of 6, 12, 18 or 24 dB per octave. Fourth-order Linkwitz-Riley is the common modern choice, and a two-way typically crosses at 1.8 to 3 kHz.
- Enclosures: sealed (acoustic suspension, 12 dB per octave rolloff, best transient behaviour), bass reflex (a Helmholtz resonator formed by port and box, 24 dB per octave, about 3 dB more output near tuning), and transmission line.
- Why a phone cannot do bass, with numbers: Sd about 1 cm^2, Xmax under 0.3 mm, back volume under 1 cm^3, so Vd is around 0.03 cm^3. Producing 80 dB SPL at 50 Hz at half a metre needs tens of cm^3 of displacement. The phone is short by three orders of magnitude. It is a volume-of-air problem, so no amount of processing fixes it.
- Other driver types:
- Piezoelectric: a ceramic disc that bends under voltage. No magnet or coil, cheap, high impedance, poor bass. Buzzers and alarms.
- Balanced armature: a magnetized reed balanced between two poles, driven by a coil, coupled by a rod to a tiny diaphragm. Efficient and small, standard in hearing aids since the 1940s and in in-ear monitors. Narrow bandwidth each, so in-ears stack two to twelve with crossovers.
- Planar magnetic: a thin film with a printed conductor between magnet arrays, so force spreads over the whole diaphragm and it flexes less.
- Electrostatic: a charged film between perforated stators driven at high voltage, needing a bias of one to six kilovolts. Very low moving mass and distortion, poor bass without a large panel.
| Driver type | Moving part | Typical use |
|---|---|---|
| Dynamic (coil) | Cone plus coil | Almost everything |
| Balanced armature | Reed plus rod | In-ear monitors |
| Piezo | Ceramic disc | Buzzers, alarms |
| Planar magnetic | Printed film | Studio headphones |
| Electrostatic | Charged film | High-end panels |
The honest version: “the amplifier drives the speaker” is half the story. The speaker drives back, through back-EMF and its reactive impedance. A real amplifier must absorb returning energy, which is why output impedance must be low and why cable resistance measurably shifts response with some loads.
WORDS9.2.6 remember these#
- Voice coil — the wire that gets pushed — the conductor in the magnetic gap, described by Re, Le and Bl.
- Bl — motor strength — flux density times conductor length, in tesla metres.
- Excursion — how far the cone travels — Xmax, peak linear displacement in mm.
- Sensitivity — how loud for a given input — dB SPL at 1 m for 2.83 V RMS.
- Crossover — the circuit splitting bass from treble — a filter network with a stated order and slope in dB per octave.
- Bass reflex — a box with a hole — a Helmholtz resonator extending output.
- Back-EMF — the speaker generating as it moves — the motional voltage Bl*u.
- Thiele/Small parameters — the driver’s spec sheet — the lumped-parameter model of a driver near and below resonance.
9.3 The microphone as a speaker in reverse#
PLAIN9.3.1 in simple words#
- A speaker turns electricity into movement, and movement into pressure. A microphone does the same three things in the opposite order.
- Pressure arrives, pushes a thin sheet, and the sheet moves.
- The moving sheet changes something electrical, and that change is the signal.
- Way one: move a coil near a magnet. Moving wire in a field makes voltage. That is a dynamic microphone.
- Way two: make the sheet one plate of a capacitor, which is two plates with a gap. Change the gap and the electrical behaviour changes. That is a condenser microphone.
- Way three: build the whole thing on a chip the size of a grain of rice. That is a MEMS microphone, and it is the one in your phone.
PLAIN9.3.2 a picture in your head#
- Think of a trampoline with a heavy rope ring tied underneath.
- Somebody drops a ball. The trampoline dips and springs back, and the ring goes down and up with it.
- Now imagine the ring passes through an invisible field that makes a small voltage whenever the ring moves.
- Big bounce, big voltage. Fast wobble, fast wobbling voltage. The trampoline is the diaphragm and the ring is the coil: a dynamic microphone.
Where this comparison breaks: a real diaphragm moves nanometres for ordinary speech, the scale of a hundred atoms. And a trampoline keeps bouncing, whereas a good diaphragm is damped so it stops the instant the sound stops, or it would add its own ringing to everything.
PLAIN9.3.3 a worked example#
- Take a condenser microphone: one fixed plate, one moving diaphragm.
- Say the gap is 20 micrometres and the plate area is 3 cm^2. That is roughly 13 picofarads, a very small capacitor.
- Put a fixed charge on it and never let the charge change.
- Voltage across a capacitor = charge divided by capacitance.
- Sound pushes the diaphragm in by 20 nanometres, so the gap shrinks by one part in a thousand and capacitance rises by about the same fraction.
- Charge is fixed, so voltage falls by about one part in a thousand. Charged to 48 volts, that is a change of about 48 millivolts. That is the signal.
- Notice the trick. We did not generate power from the sound. We used the sound to modulate a power supply that was already there.
PLAIN9.3.4 what is really happening inside#
- The dynamic microphone is a small speaker used backwards: a light diaphragm with a coil glued to it, hanging in a magnet gap.
- Sound moves the coil through the field and Faraday’s law gives a voltage. It needs no power supply, because the sound itself does the work.
- Because the sound must move the coil’s mass, dynamics are less sensitive and less detailed high up, but rugged and happy with very loud sources.
- The ribbon microphone replaces the coil with a corrugated strip of very thin aluminium which is itself the diaphragm. Very low mass, fragile, and naturally equally sensitive front and back.
- The condenser needs an external voltage to charge its plates, plus an amplifier right beside the capsule, because the signal is tiny and the source impedance is enormous. That voltage is usually 48 volts sent down the same cable as the audio, called phantom power.
- The electret is a condenser whose charge is permanently baked into a plastic film during manufacture, so no polarizing supply is needed. Electrets made microphones cheap enough for every telephone and laptop for forty years.
- The MEMS microphone is that idea built by chip-making processes: a diaphragm a fraction of a millimetre across etched from silicon, with a perforated silicon backplate a few micrometres behind it.
- Beside it sits a small chip that buffers the signal and often digitizes it immediately. The package is about 3 x 4 x 1 mm with a sound hole.
- Because they are made photographically, thousands come out identical. That is what makes microphone arrays and beamforming possible in a phone.
TECHNICAL9.3.5 the engineer’s version#
- Dynamic output follows V = Bl * u, the same expression as speaker back-EMF. Sensitivity is typically 1 to 3 mV/Pa. The Shure SM58, introduced in 1966, is specified at 1.85 mV/Pa at 1 kHz into 150 ohms.
- Condenser capsules are 10 to 60 pF. Capacitive reactance at 20 Hz for 20 pF is about 400 megohms, so the impedance converter, a FET or valve stage, must sit within millimetres of the capsule.
- Phantom power is standardized as P48 in IEC 61938: 48 V nominal fed through two 6.81 kilohm resistors, one per leg of a balanced pair.
- The electret was invented at Bell Labs by Gerhard Sessler and James West in 1962, and the foil-electret condenser became the most manufactured microphone type in history.
- Polar pattern is a property of capsule venting, not of transducer type: omnidirectional if only the front is open, figure-of-eight if front and back are equally open, cardioid if the rear path is acoustically delayed.
- Digital MEMS parts output PDM, a one-bit stream clocked at 1 to 3.072 MHz and decimated in the host codec, or I2S/TDM directly. PDM lets a microphone sit far from the processor on two wires without picking up noise.
- Observation:
arecord -llists Linux capture devices, andsystem_profiler SPAudioDataTypelists macOS inputs.
| Parameter | Typical MEMS | Note |
|---|---|---|
| Sensitivity | -38 dBV/Pa | About 12.6 mV/Pa |
| Signal-to-noise ratio | 64 to 70 dB(A) | Higher is better |
| Acoustic overload point | 120 to 133 dB SPL | Where 10 % THD hits |
| Supply current | 100 to 250 uA | Always-on capable |
The honest version: “a microphone is a speaker in reverse” is exactly true for dynamic microphones and false for condenser and MEMS types. In a condenser the sound does not supply the signal energy; it modulates a bias already present. That is why a condenser needs power and a dynamic does not.
WORDS9.3.6 remember these#
- Diaphragm — the thin sheet the sound pushes — the moving membrane.
- Transducer — anything converting one energy to another — here acoustic to electrical, or the reverse.
- Phantom power — voltage sent up the microphone cable — P48 per IEC 61938.
- Electret — a permanently charged film — a dielectric with frozen-in polarization, removing the need for a bias supply.
- MEMS — a microphone built like a chip — micro-electro-mechanical system with a silicon diaphragm and an integrated ASIC.
- Acoustic overload point — where it distorts — the SPL at which total harmonic distortion reaches 10 percent.
9.4 From wave to numbers#
PLAIN9.4.1 in simple words#
- A computer cannot store a wiggle. It can only store numbers.
- So we measure the wiggle over and over, very fast, and write the numbers down.
- Each measurement is a sample: a snapshot of the voltage at one instant.
- How many snapshots per second is the sample rate. A music CD takes 44,100 per second, of each of two channels.
- How precisely each snapshot is written is the bit depth. A CD writes each one as a 16-bit number, one of 65,536 levels.
- Measure often, measure precisely, write it down. That is the whole idea.
- Playback means reading the numbers and rebuilding the voltage.
PLAIN9.4.2 a picture in your head#
- Imagine a heart-rate monitor drawing a line on moving graph paper.
- You cannot keep the paper. You may only write the height of the line at fixed moments: 42, 58, 71, 65, 44, 30, 41.
- Later somebody plots your numbers as dots.
- If the dots are close enough together, exactly one smooth curve fits them, and it is the original line.
- If they are too far apart, many curves fit equally well, and the information is gone forever.
Where this comparison breaks: on paper you would join dots with straight lines and get a jagged shape. Real reconstruction does not. The mathematics says that if you sampled fast enough there is exactly one band-limited curve through those points, and a filter finds it. The output is a smooth wave, not a staircase.
PLAIN9.4.3 a worked example#
sampling a wave: samples over one cycle
+1.0 | . - * - .
| * *
| * *
0.0 |--*-------------------*-------------------*--
| * *
| * *
-1.0 | . - * - .
| | | | | | | | | | | | | | |
t0 t15
the "*" marks are the values we keep. everything
between them is rebuilt by the reconstruction filter.
- Now watch what goes wrong when we sample too slowly.
- Suppose the real signal is 30,000 Hz but we take 44,100 samples a second.
- The readings are indistinguishable from a 14,100 Hz signal, because 44,100 minus 30,000 is 14,100.
- A tone you could barely hear reappears as a loud whistle in the middle of your hearing range. That is aliasing, and it cannot be removed afterwards, because both signals produced the same numbers.
- So we put a filter before the converter that removes everything above half the sample rate. That is the anti-aliasing filter, and it is not optional.
- You have seen aliasing in films: a wheel spinning forward appears to turn slowly backwards, because 24 frames a second cannot track the spokes.
PLAIN9.4.4 what is really happening inside#
- There is a theorem behind all this, worth stating properly.
- If a signal contains no frequency above F, sampling at any rate above 2F captures it perfectly, with nothing lost.
- That is the Nyquist-Shannon sampling theorem. Harry Nyquist stated the underlying result at Bell Labs in 1928, and Claude Shannon published the modern form in 1949. “Perfectly” means perfectly, not approximately.
- So to record up to 20,000 Hz we need more than 40,000 samples a second.
- Why 44,100 and not 40,000? A practical accident. Before disks were big enough, early digital audio was stored on video tape using a PCM adaptor which packed samples into fake video lines.
- 44,100 fits both video standards: 60 fields x 245 lines x 3 samples, and 50 fields x 294 lines x 3 samples, both give 44,100.
- The extra 4,100 Hz gives the anti-aliasing filter room to roll off between 20 kHz and 22.05 kHz.
- Now bit depth. Each sample is rounded to the nearest available level, and that quantization error appears as faint hiss.
- Each extra bit halves the error, which is 6 dB quieter, so 16 bits gives about 96 dB of range and 24 bits about 144 dB.
- One subtlety: for very quiet signals the error stops being random hiss and correlates with the music, sounding like gritty distortion.
- The fix is dither: add a tiny amount of deliberate noise before rounding. It sounds wrong and it works, turning ugly distortion into benign hiss, and letting you hear signals quieter than one single level.
TECHNICAL9.4.5 the engineer’s version#
- Nyquist-Shannon: a signal band-limited to B hertz is completely determined by samples at fs > 2B, and fs/2 is the Nyquist frequency. Energy above fs/2 folds back to |f - fs| and cannot be separated afterwards, so anti-alias filtering before the sample-and-hold is mandatory.
- Signal-to-quantization-noise ratio for an ideal uniform quantizer with a full-scale sine: SQNR = 6.02 N + 1.76 dB. For N = 16 that is 98.09 dB. The commonly quoted 96 dB is 6.02 x 16 without the 1.76 term.
- Dither is usually triangular PDF noise at 2 LSB peak-to-peak, fully decorrelating the error at the cost of 4.77 dB more noise. Noise-shaped dither moves that noise above 15 kHz where the ear is least sensitive.
- Two converter architectures dominate.
- Successive approximation (SAR): a binary search. The held sample is compared against half full scale, then a quarter, and so on, one bit per clock, so N bits take N comparisons and latency is one conversion.
- Sigma-delta: samples at 64 to 256 times the target rate with a coarse, often one-bit, quantizer inside a feedback loop that shapes quantization noise up out of the audio band, then decimates digitally. This trades speed for resolution and is why 24-bit converters are cheap. Nearly every audio ADC and DAC since the mid-1990s is sigma-delta.
- Oversampling also lets the analogue anti-alias filter be gentle, because the first alias lands 64 times higher. The steep filtering happens in the digital decimator, where it can be made exact and phase-linear.
| Sample rate | Nyquist limit | Where it is used |
|---|---|---|
| 8 kHz | 4 kHz | Telephone, G.711 |
| 16 kHz | 8 kHz | Wideband voice, AMR-WB |
| 44.1 kHz | 22.05 kHz | CD, most music files |
| 48 kHz | 24 kHz | Video, broadcast, most OS |
| 96 kHz | 48 kHz | Studio production |
| 192 kHz | 96 kHz | Studio, largely marketing |
Experts disagree here. A 2007 AES paper by E. Brad Meyer and David R. Moran found listeners could not distinguish high-resolution playback from the same signal passed through a 16-bit 44.1 kHz loop. A 2016 AES meta-analysis by Joshua Reiss pooled many studies and found a small but statistically significant ability to discriminate, especially with training. Any effect is small, and higher rates are most clearly useful during production, where repeated processing accumulates error.
WORDS9.4.6 remember these#
- Sample — one measurement of the wave — one quantized amplitude at one instant.
- Sample rate — how often we measure — samples per second, in hertz.
- Bit depth — how precisely we measure — bits per sample, giving 2^N levels.
- Nyquist frequency — the highest note recordable — half the sample rate.
- Aliasing — a high note pretending to be a low one — spectral folding of content above fs/2 back into the baseband.
- Dither — helpful noise added on purpose — low-level TPDF noise decorrelating quantization error from the signal.
- Sigma-delta — the converter in everything — oversampling noise-shaping converter plus digital decimation filter.
9.5 From numbers back to wave#
PLAIN9.5.1 in simple words#
- Playback runs the recording chain backwards.
- A chip reads the numbers in order, at exactly the right speed, and sets an output voltage to match each one. That is the DAC, the digital-to-analogue converter.
- Its raw output is a series of steps, because it holds each value until the next arrives.
- A filter smooths those steps back into the original wave. That is the reconstruction filter.
- The smoothed wave is correct but far too weak to move a speaker.
- So it goes into an amplifier, which uses the small wave as a template and builds a big copy, taking the energy from the power supply.
PLAIN9.5.2 a picture in your head#
- Imagine a precise but tiny artist painting a perfect miniature.
- It is correct in every detail but one centimetre across, so nobody at the back of the hall can see it.
- A second worker copies it onto a wall, stroke for stroke, a thousand times bigger, using buckets from a store room.
- The tiny artist is the DAC, the wall painter is the amplifier, the store room is the power supply.
- The painter adds nothing new, and a mistake in the miniature becomes a huge mistake on the wall. The paint comes from the store room, not the miniature.
Where this comparison breaks: a real amplifier is not a passive copier. It is a feedback system constantly comparing its output against its input and correcting the difference far faster than the music changes. It also adds a little of its own noise and distortion.
PLAIN9.5.3 a worked example#
- A laptop headphone socket puts out at most about 1 volt RMS, with a source impedance of perhaps 10 ohms.
- Headphone A is a 32 ohm in-ear rated 105 dB per milliwatt. Power = volts squared divided by ohms, so at 0.8 V that is 0.64 / 32 = 20 milliwatts. It will be painfully loud.
- Headphone B is a 250 ohm studio model rated 96 dB per milliwatt. At 1 V, power = 1 / 250 = 4 milliwatts, which is +6 dB over one milliwatt, giving about 102 dB. Loud enough, but with little headroom.
- Now the source impedance problem. A 10 ohm output feeding a 32 ohm headphone is a divider, losing about a quarter of the voltage.
- Worse, headphone impedance varies with frequency, so a high source impedance changes the tone. The design rule is to keep source impedance under one eighth of the load.
- That is why 250 and 600 ohm headphones want a real amplifier, and why 16 ohm in-ears sound tonally different from a high-impedance output.
PLAIN9.5.4 what is really happening inside#
- The DAC receives a stream of numbers plus a clock saying when each is due.
- The clock matters enormously. If its timing wobbles, the wave is rebuilt at slightly wrong moments. That is jitter.
- Most audio DACs oversample rather than building the exact voltage in one step.
- A digital filter invents extra samples between the real ones, raising the internal rate by 8, 16, 64 or 128 times.
- Then a very fast, very coarse converter produces an output at that high rate, with its errors pushed up into ultrasonic frequencies.
- Then a gentle analogue filter removes the ultrasonic mess. It is the sigma-delta idea run in reverse.
- The output is typically 1 to 2 volts RMS at a few milliamps: line level, meant for another circuit, not for a speaker.
- The amplifier keeps that voltage shape but supplies amperes.
- A class AB amplifier uses transistors that are always partly on, wasting power as heat, at roughly 50 to 70 percent efficiency.
- A class D amplifier switches fully on and fully off hundreds of thousands of times a second, varying how long it stays on, and a filter averages that into the wanted wave. Efficiency reaches 85 to 95 percent.
- Class D is why a phone, a soundbar and a car stereo can be loud without a heatsink the size of a brick.
TECHNICAL9.5.5 the engineer’s version#
- The raw DAC output is a zero-order hold, imposing a sinc amplitude droop of sin(x)/x with x = pi*f/fs, about -3.2 dB at 20 kHz with fs = 44.1 kHz. Oversampling DACs correct it in the digital interpolation filter.
- Reconstruction is usually a linear-phase FIR filter at 8x to 128x oversampling followed by a second or third order analogue filter. The choice between linear phase (pre-ringing) and minimum phase (phase shift) is selectable on many current DAC chips and is a matter of taste.
- Clock jitter budget: below roughly 100 picoseconds RMS for 16-bit performance at 20 kHz, and about 1 picosecond for 24-bit. That is why asynchronous USB audio, where the DAC owns the clock, replaced adaptive USB audio.
- Headphone impedances in use: 16 to 32 ohms for in-ears and headsets, 32 to 80 ohms for consumer over-ears, 250 and 600 ohms in the Beyerdynamic DT770 and DT880 families, 300 ohms for the Sennheiser HD 600 and HD 650. The high-impedance variants were designed for professional equipment with high output voltage. Damping factor is load impedance divided by source impedance, and the “rule of eighth” convention says keep it above 8.
- What a product sold as “a DAC” actually contains: a USB, S/PDIF or Bluetooth receiver, a clock, the converter silicon, an analogue output stage, a volume control, usually a headphone amplifier, and a power supply. The converter chip is often a few dollars of the bill of materials, and most audible difference between such products comes from the analogue output stage, the headphone amplifier and the volume control, not the chip named on the box.
- Observation:
cat /proc/asound/card0/pcm0p/sub0/hw_paramson Linux shows the rate, format and buffer actually in use while audio plays. On macOS, Audio MIDI Setup shows the current device rate and bit depth.
| Class | How it works | Efficiency |
|---|---|---|
| A | Always fully conducting | 20 to 30 percent |
| AB | Mostly on, brief overlap | 50 to 70 percent |
| D | Switching at 300 kHz plus | 85 to 95 percent |
| G / H | AB with stepped rails | 60 to 80 percent |
WORDS9.5.6 remember these#
- DAC — the chip turning numbers into voltage — digital-to-analogue converter, usually oversampling sigma-delta.
- Zero-order hold — holding each value until the next — the staircase output causing sinc amplitude droop.
- Reconstruction filter — the smoother — the low-pass filter recovering the unique band-limited waveform.
- Jitter — timing wobble — deviation in sample clock edge timing, in picoseconds RMS.
- Line level — a weak but correct signal — nominally 2 V RMS consumer or 1.228 V RMS (+4 dBu) professional.
- Class D — the efficient switching amplifier — pulse-width modulated output stage with an LC output filter.
- Damping factor — how firmly the amplifier grips the driver — load impedance divided by amplifier output impedance.
9.6 Storing audio#
PLAIN9.6.1 in simple words#
- The simplest way to store sound is to write every sample in order, with nothing clever done to it. That is PCM, pulse-code modulation.
- A WAV file is mostly raw PCM with a small label saying how fast, how many bits and how many channels.
- Raw PCM is honest and huge, so we compress it, in two different ways.
- Lossless compression packs the data tighter and gives back every original number exactly. FLAC and ALAC do this, roughly halving the size.
- Lossy compression throws information away permanently, choosing parts you are unlikely to notice. MP3, AAC and Opus do this, cutting by ten or more.
- A codec is the method of coding and decoding. A container is the file format holding the coded data plus the title, artist and length.
- The same codec can live in different containers, and one container can hold different codecs. That is the most confusing thing about audio files.
PLAIN9.6.2 a picture in your head#
- Imagine posting a long letter. Raw PCM is writing every word out in full.
- Lossless compression is shorthand. The reader expands it back to the same letter word for word. Nothing lost, thinner envelope.
- Lossy compression is writing a summary. The reader gets the meaning and never knows which words you dropped.
- And you cannot un-summarize. Converting a summary back into a full letter just produces a fatter summary.
- That is why converting an MP3 to FLAC restores nothing. It stores the damaged version losslessly.
Where this comparison breaks: a written summary drops whole ideas, which you would notice. A good audio codec drops only sounds physically masked by louder sounds happening at the same moment, which your ear could not have detected. It is closer to not printing the parts of a photograph that are behind a wall.
PLAIN9.6.3 a worked example#
- Step one: 44,100 samples per second.
- Step two: 16 bits each, so 44,100 x 16 = 705,600 bits per second per channel.
- Step three: stereo doubles it, giving 1,411,200 bits per second. That is the famous 1.411 Mbps.
- Step four: divide by 8 for bytes, giving 176,400 bytes per second.
- Step five: a three-minute song is 180 seconds, so 176,400 x 180 = 31,752,000 bytes, about 31.8 MB, shown as 30.3 MiB by a file manager.
- The same song as a 128 kbps MP3 is 128,000 / 8 x 180 = 2,880,000 bytes, about 2.9 MB, roughly eleven times smaller.
- As FLAC at a typical 60 percent, about 19 MB. As 64 kbps Opus, about 1.4 MB.
CD audio bitrate, from first principles
44100 samples/s
x 16 bits/sample = 705600 bits/s per channel
x 2 channels = 1411200 bits/s = 1.4112 Mbit/s
/ 8 = 176400 bytes/s
x 180 s (a 3 minute song) = 31752000 bytes
= 31.75 MB (30.28 MiB)
a full 74 minute CD:
176400 x 4440 s = 783216000 bytes = 783 MB
PLAIN9.6.4 what is really happening inside#
- Lossless codecs work in two stages. First, prediction: guess the next sample from the previous few using a small polynomial or adaptive filter. Music is smooth, so the guess is usually close.
- Second, store only the error between guess and truth. Those errors are small numbers clustered near zero, and they compress very well with a code that gives short bit patterns to common values.
- Nothing is discarded, so decoding reproduces the original bit for bit.
- Lossy codecs add a third idea: a model of your ear.
- Fact one, the threshold of hearing: very quiet sounds are inaudible, and the level at which that happens depends strongly on frequency.
- Fact two, simultaneous masking: a loud tone makes nearby quieter tones inaudible. A loud 1 kHz tone can hide a tone 40 dB quieter at 1.1 kHz.
- Fact three, temporal masking: for a few milliseconds before, and up to about 200 milliseconds after a loud sound, quieter sounds near it are inaudible.
- So the encoder splits the signal into frequency bands, works out how much noise each band can hide, and spends just enough bits per band to keep its own quantization noise under that hidden threshold.
- The bits go where the ear is looking, and everywhere else gets almost nothing.
- That is why 128 kbps MP3 is not “lower quality audio”. It is full-range audio with its errors deliberately parked underneath louder sounds.
- When the bitrate is too low the encoder runs out of bits, errors climb above the masking threshold, and you hear it: swishy cymbals, a metallic ring on voices, and a hard cut-off of high frequencies.
TECHNICAL9.6.5 the engineer’s version#
- WAV is a container built on the Microsoft and IBM RIFF chunk format from 1991: a
fmtchunk declares rate, channels and bits, adatachunk holds interleaved PCM. AIFF is Apple’s 1988 equivalent on the Electronic Arts IFF format, big-endian where WAV is little-endian. - MP3 is formally MPEG-1 Audio Layer III, from the Fraunhofer Institute for Integrated Circuits in Erlangen and the University of Erlangen-Nuremberg, with Karlheinz Brandenburg completing the underlying doctoral work in 1989. ISO/IEC 11172-3 reached committee draft in 1991, was finalized in 1992 and published in 1993. MPEG-2 Layer III, ISO/IEC 13818-3, added lower sample rates and was published in 1995.
- The
.mp3extension was chosen by the Fraunhofer team on 14 July 1995; before that the files used.bit. The last US patent covering MP3 expired on 16 April 2017. - MP3 supports 32, 44.1 and 48 kHz under MPEG-1 and 16, 22.05 and 24 kHz under MPEG-2, with 14 fixed bitrates from 32 to 320 kbit/s. The 128 kbit/s convention is an 11:1 ratio against 1,411.2 kbit/s CD audio. Internally it uses a 32-band polyphase filter followed by an MDCT of 18 or 6 lines, non-uniform quantization, Huffman coding, and a bit reservoir.
- AAC was standardized as MPEG-2 Part 7, ISO/IEC 13818-7, in 1997, then as MPEG-4 Part 3, ISO/IEC 14496-3, in 1999, developed jointly by AT&T Labs, Dolby, Fraunhofer IIS and Sony, with Nokia joining as co-licensor in 2002. It uses a pure MDCT filter bank with 2048 or 256 sample windows.
- Opus was standardized by the IETF as RFC 6716 on 10 September 2012, merging SILK, a linear-prediction speech coder from Skype, with CELT, a low-latency MDCT coder from Xiph.Org. It spans 6 to 510 kbit/s with frames of 2.5 to 60 ms and a default algorithmic delay of 26.5 ms, reducible to 5 ms. It is the default codec for WebRTC.
- FLAC reached version 1.0 on 20 July 2001, written by Josh Coalson, and typically compresses to 50 to 70 percent of the original. In December 2024 it was formally specified as IETF RFC 9639. ALAC is Apple’s equivalent, introduced in 2004 and open sourced in 2011.
- Codec versus container, precisely: a codec defines a bitstream syntax and a decoding process, while a container defines how coded streams, metadata, timestamps and index tables are packed into a file. MP4 and M4A are the same container, the ISO base media file format, and may hold AAC or ALAC. Ogg and WebM commonly hold Opus. The
.mp3case is unusual: a bare stream with no real container, which is why its metadata is a bolt-on called ID3.
| Format | Type | Typical bitrate | Best use |
|---|---|---|---|
| PCM in WAV | None | 1,411 kbps | Editing, mastering |
| FLAC | Lossless | 700 to 1,000 kbps | Archiving music |
| ALAC | Lossless | 700 to 1,000 kbps | Apple archiving |
| MP3 | Lossy | 128 to 320 kbps | Legacy compatibility |
| AAC | Lossy | 96 to 256 kbps | Streaming, video, iOS |
| Opus | Lossy | 6 to 128 kbps | Calls, WebRTC, chat |
# what does this file actually contain?
ffprobe -hide_banner song.m4a
# container = mov,mp4,m4a codec = aac 48000 Hz stereo
# convert losslessly, keeping every sample
ffmpeg -i track.wav -c:a flac -compression_level 8 track.flac
# encode speech efficiently
opusenc --bitrate 32 --speech voice.wav voice.opus
The honest version: “MP3 removes frequencies above 16 kHz” is a symptom, not the method. Low-bitrate encoders often apply a lowpass because high frequencies are expensive and least missed, but that is a bit-allocation decision by one encoder, not part of the standard. Implementation detail: LAME, the dominant open-source MP3 encoder, and Fraunhofer’s own encoder make different choices at the same bitrate, and LAME’s variable-bitrate presets beat fixed 128 kbps.
WORDS9.6.6 remember these#
- PCM — the raw numbers — pulse-code modulation, uncompressed linear samples.
- Codec — the coding method — the coder/decoder pair defining a bitstream.
- Container — the box the data sits in — a file format multiplexing streams, metadata and timing.
- Lossless — exactly what you put in — bit-exact reconstruction of the PCM.
- Lossy — some detail gone for good — perceptually irrelevant information discarded permanently.
- Masking — a loud sound hides a quiet one — simultaneous and temporal elevation of the hearing threshold.
- MDCT — the maths inside every modern codec — modified discrete cosine transform with 50 percent overlapping windows.
9.7 How recorded sound was stored before computers#
PLAIN9.7.1 in simple words#
- Before computers, sound was stored as a physical shape or a magnetic pattern.
- The phonograph stored it as a wiggly groove cut into a spinning cylinder.
- The groove is literally the waveform. You can see the sound under a microscope.
- Magnetic tape stored it as magnetized specks on a plastic ribbon.
- The vinyl record used a groove again, on a flat pressed disc.
- The compact disc stored numbers instead, as microscopic dents read by a laser.
- Only the last is digital. The first three are direct physical copies of the wave.
PLAIN9.7.2 a picture in your head#
- Picture drawing a wobbly line along a very long road with chalk.
- Later, walk the same road dragging your finger along the chalk line.
- Your finger is forced to wobble exactly as the line wobbles.
- Connect the finger to a drum skin, and the skin wobbles, and the air wobbles, and you hear the original sound.
- That is a record player. No code, no numbers, no cleverness. The groove is a drawing of the pressure wave, and the needle is pushed by it.
Where this comparison breaks: a record groove wobbles sideways, and a stereo groove wobbles both walls independently at 45 degrees. Records are also not cut flat: bass is reduced and treble boosted when cutting, then reversed on playback, so grooves stay narrow and surface noise falls. That curve is the RIAA equalization standard from 1954.
PLAIN9.7.3 a worked example#
- Take a 33 and a third rpm LP, 12 inches across.
- It turns once every 1.8 seconds, so the outer groove passes the needle at about 50 cm per second and the inner groove at only about 21 cm per second.
- That is why the last track on a side often sounds slightly worse. Less groove length is available per second of music.
- A side holds about 25 minutes, and the microgroove standard allows a stylus tip radius of 25 micrometres.
- A CD instead runs at constant linear velocity, 1.2 to 1.4 m/s, so it spins faster in the middle and slower at the edge, and every part of the surface carries data at the same density.
PLAIN9.7.4 what is really happening inside#
- Magnetic tape is plastic ribbon coated with billions of needle-shaped iron oxide particles.
- Each particle is a tiny permanent magnet holding whichever direction you last pushed it into.
- The record head is an electromagnet with a very narrow gap, and audio current through it makes a field at that gap.
- Tape passes the gap, and particles leaving the field freeze in whatever magnetization the field last gave them: strongly aligned for a loud moment, weakly aligned for a quiet one.
- On playback the moving magnetized tape passing another gap induces a voltage in a coil, by Faraday’s law, the same law as the microphone.
- Hold this for Chapter 12: a hard disk is the same idea with the tape replaced by a rigid spinning platter, a head flying nanometres above it, and the magnetization standing perpendicular to the surface rather than along it. The physics of “particles keep the direction you last pushed them” is unchanged.
- The compact disc is different, because it stores numbers rather than a wave.
- A spiral track of pits is pressed into plastic and coated with reflective aluminium. A laser shines up at it and a photodiode watches the reflection.
- A pit is about a quarter wavelength deep, so light from a pit edge cancels light from the flat land beside it and the brightness drops sharply.
- The honest version: a pit is not a 1 and a land is not a 0. Every transition from pit to land or land to pit is a 1, and no transition is a 0. That coding is called NRZI, and it exists because an edge is far more reliably detected than an absolute level.
TECHNICAL9.7.5 the engineer’s version#
- Edouard-Leon Scott de Martinville patented the phonautograph on 25 March 1857. It traced sound onto soot-blackened paper but could not play back. The oldest recovered recording from it, of “Au Clair de la Lune”, dates from 9 April 1860 and was recovered digitally in 2008.
- Thomas Edison announced the phonograph on 21 November 1877 and demonstrated it on 29 November 1877, recording on tinfoil wrapped round a cylinder.
- Emile Berliner moved to flat discs, beginning commercial gramophone production in 1892 with five-inch records and moving to shellac in 1895.
- Valdemar Poulsen demonstrated magnetic recording on steel wire at the end of the 1890s. Fritz Pfleumer patented magnetic tape in Germany in 1928, AEG built it into the Magnetophon with BASF tape in the mid-1930s, and American engineers, notably Jack Mullin, brought captured machines home after the Second World War, after which Ampex productionized them.
- Columbia Records introduced the 33 and a third rpm microgroove LP at a press conference on 21 June 1948. Stereo LPs followed from 1957.
- Compact Disc Digital Audio, the Red Book, was published by Philips and Sony in 1982 and adopted as IEC 60908 in 1987. The first commercial CD release was on 17 August 1982, and the first fifty titles went on sale in Japan on 1 October 1982.
- EFM, devised by Kees Schouhamer Immink, maps each 8-bit byte to a 14-bit pattern chosen so runs of zeros are between 2 and 10 long, keeping pit lengths within what the optics resolve while leaving enough transitions for clock recovery. Three merging bits join the patterns.
- CIRC is Sony’s cross-interleaved Reed-Solomon code. Interleaving spreads consecutive samples across the disc so a scratch destroys one symbol from many codewords rather than many symbols from one, which is exactly what Reed-Solomon repairs. It fully corrects bursts of about 4,000 bits, roughly 2.5 mm of track.
| CD specification | Value |
|---|---|
| Disc diameter | 120 mm |
| Laser wavelength | 780 nm, infrared |
| Track pitch | 1.6 micrometres |
| Maximum playing time | 74 min 33 s |
| Channel coding | EFM, eight-to-fourteen |
| Error correction | CIRC, Reed-Solomon |
WORDS9.7.6 remember these#
- Groove — the wiggly channel in a record — a mechanically modulated spiral carrying the waveform laterally.
- Stylus — the needle — the tip tracing the groove, typically 25 um radius on microgroove records.
- Coercivity — how hard a magnet is to flip — the reverse field needed to demagnetize a recorded particle.
- Pit and land — the dents and the flat between them — features whose transitions encode 1 bits under NRZI.
- EFM — how bytes become pit patterns — eight-to-fourteen modulation with a run-length constraint of 2 to 10.
- CIRC — the CD’s scratch repair — cross-interleaved Reed-Solomon coding.
- RIAA curve — the record equalization standard — the 1954 pre-emphasis and playback de-emphasis specification.
9.8 Audio inside a computer#
PLAIN9.8.1 in simple words#
- Inside a computer, sound is a stream of numbers moving to a deadline.
- A chip called the audio codec holds the converters, in both directions.
- Your program writes numbers into a shared area of memory called a buffer, and the hardware reads from it at a fixed, unforgiving rate.
- If your program is late even once, the hardware runs out of numbers and plays silence or repeats old data. That gap is heard as a click.
- So the whole design problem of computer audio is: never be late, ever.
- A bigger buffer gives more safety margin but more delay before you hear anything. That delay is latency.
PLAIN9.8.2 a picture in your head#
- Think of a conveyor belt feeding a machine that must never stop.
- You stand at one end putting parts on. The machine takes one every second, forever.
- A long belt lets you wander off for a minute, but a part you place now is not used for a long time.
- A short belt means what you place is used almost immediately, which is what you want when playing a guitar through the computer, but blink at the wrong moment and the machine starves and jams.
- Length buys safety, shortness buys responsiveness, and you cannot have both.
Where this comparison breaks: the computer’s belt is circular. It is a ring buffer, written at one point and read at another chasing it round. Nothing is added to an end, and if the reader catches the writer you get the click.
PLAIN9.8.3 a worked example#
- Buffering latency = frames in the buffer divided by the sample rate.
- At 48,000 samples a second, a 256-frame buffer is 256 / 48000 = 5.33 milliseconds.
- But there are normally at least two such periods, one playing and one being filled, so real output latency is around 10.7 ms.
- Add the input side for live monitoring and you are near 21 ms.
- Sound travels 343 metres a second, so 21 ms is the delay of standing 7.2 metres from an amplifier. Musicians notice about 10 ms and hate 25 ms.
- Drop to a 64-frame buffer and each period is 1.33 ms. Beautiful, and it will crackle on a busy machine.
| Buffer size | At 48 kHz | Typical use |
|---|---|---|
| 32 frames | 0.67 ms | Live guitar, low latency |
| 128 frames | 2.67 ms | Studio recording |
| 512 frames | 10.7 ms | Mixing, general use |
| 2048 frames | 42.7 ms | Video playback, games |
PLAIN9.8.4 what is really happening inside#
- The codec chip contains ADCs, DACs, a mixer, headphone and speaker amplifiers, and often a microphone bias supply.
- It talks to the processor over a small serial bus: I2S or TDM for the audio samples, I2C or SoundWire for control commands.
- A DMA engine, which is hardware that moves memory without troubling the CPU, walks round the ring buffer handing samples to the codec on a timer.
- When it finishes a period it raises an interrupt, the driver wakes the audio thread, and that thread must fill the next period before the engine arrives.
- If it does not, that is an underrun, called an xrun in ALSA.
- Why does an underrun click rather than just go quiet? Because the waveform jumps instantly to zero, and a vertical step contains energy at every frequency at once. Your ear hears a broadband snap.
- Sample rate conversion is the other everyday complication: your file is 44.1 kHz and your device runs at 48 kHz, so somebody must resample.
- The ratio is 160 to 147, so good resampling needs a long filter and costs CPU, while bad resampling adds audible aliasing.
- Above the driver sits a mixer, because many applications want the one device at once. It resamples everything to a common rate, applies per-stream volume, sums, and hands a single stream to the hardware.
TECHNICAL9.8.5 the engineer’s version#
- Historical anchors: the Creative Labs Sound Blaster arrived in 1989 and set the PC convention for a decade. Intel’s AC’97 specification came in 1997 and split the digital controller from the analogue codec. Intel High Definition Audio, codenamed Azalia, replaced it in 2004 and remains the PC standard, usually paired with a Realtek ALC-series codec.
- WASAPI has two modes. Shared mode goes through the Windows audio engine, is resampled to a fixed device format and mixes with other applications. Exclusive mode hands the device to one application bit-perfectly. ASIO bypasses the Windows engine entirely, which is why it survives in professional work.
- On Linux, ALSA is the kernel driver interface. PulseAudio, from 2004, added per-application volume and network transparency but had poor low-latency behaviour, while JACK served professional audio. PipeWire replaced both with one graph-based server, becoming the default in Fedora 34 in April 2021 and in Ubuntu 22.10 in October 2022.
- Core Audio expresses everything as float32 at the device rate, and is the only one of the three major platforms where the low-latency path is also the ordinary path.
- Rule of thumb for round-trip latency: buffer periods, plus converter group delay (sigma-delta decimation and interpolation filters add roughly 0.5 to 1.5 ms each way), plus any resampling delay. Manufacturers quote only the buffer figure, so measured round-trip is always worse than the control panel.
| Stack | Platform | Note |
|---|---|---|
| Core Audio | macOS and iOS | Since Mac OS X 10.0, 2001 |
| WASAPI | Windows | Since Vista, 2007 |
| ASIO | Windows | Steinberg, 1997, pro use |
| ALSA | Linux kernel | Kernel driver layer |
| PipeWire | Linux userspace | Fedora 34 default, 2021 |
# Linux: what hardware exists, and what it is doing now
aplay -l
cat /proc/asound/card0/pcm0p/sub0/hw_params
pw-top # live quantum, rate and xrun counts
# macOS
system_profiler SPAudioDataType
WORDS9.8.6 remember these#
- Buffer — the waiting area for samples — a ring buffer read by DMA on a fixed timer.
- Latency — the delay before you hear it — frames divided by sample rate, times the number of periods, plus converter delay.
- Underrun / xrun — the moment audio runs out — the DMA pointer overtaking the application’s write pointer.
- Codec chip — the audio front end on the board — integrated ADC, DAC, mixer and amplifiers on an I2S/I2C interface.
- Resampling — changing the sample rate — polyphase interpolation and decimation, 147:160 between 44.1 and 48 kHz.
- Exclusive mode — one application owns the device — bit-perfect path bypassing the system mixer, WASAPI exclusive or ASIO.
9.9 Phone calls, part 1 - the old way#
PLAIN9.9.1 in simple words#
- The original telephone is astonishingly simple.
- A microphone at one end changes the electric current in a wire, following the sound.
- That same current runs down the wire through a small speaker at the far end, which moves and makes the sound again.
- No code, no numbers, no computer. The wire carries a shrinking and growing electrical copy of the pressure wave.
- The pair of copper wires from your house to the exchange is the local loop.
- To connect any phone to any other, you go through an exchange, a building whose whole job is joining wires together.
- At first humans did the joining, plugging cords into sockets. Later machines did it, driven by the clicks your dial sent down the line.
- The crucial point: while you talked, a continuous electrical path existed from your house to theirs, and nobody else could use it. That is circuit switching.
PLAIN9.9.2 a picture in your head#
- Think of the old phone network as a railway that gives you a private track.
- To make a call, the railway builds a track from your station to theirs, and it stays yours for the whole call, even during silences.
- When you hang up, the track is dismantled and the pieces return to the pool.
- This is wonderful for quality. Nothing can get in your way, so the delay never changes and the sound never breaks up.
- It is terrible for efficiency. About half of a conversation is silence, and all that track sits idle.
Where this comparison breaks: after the 1960s the private track stopped being a continuous piece of copper. Your voice was digitized at the exchange and given a fixed repeating slot on a shared high-speed line. It behaves exactly like a private track, with guaranteed capacity and constant delay, but physically it shares one cable with dozens of other calls.
PLAIN9.9.3 a worked example#
- Here is how a phone call became exactly 64,000 bits a second.
- The telephone band is 300 Hz to 3,400 Hz, chosen to keep speech intelligible while using as little of the wire as possible.
- To sample up to about 4,000 Hz you need at least 8,000 samples a second, so 8 kHz was chosen, and each sample is stored in 8 bits.
- 8,000 x 8 = 64,000 bits per second. That is one voice channel, called a DS0.
- Why is 8 bits acceptable when a CD needs 16? Because of companding.
- Instead of 256 evenly spaced levels, the steps are fine near zero and coarse near the loudest values, following a logarithmic curve.
- Quiet passages get fine resolution where you would notice the noise, and loud passages get coarse steps where loudness hides them.
- An 8-bit companded sample carries roughly the useful range of a 12 to 13-bit linear sample.
one voice channel, from air to bits
speech -> 300-3400 Hz filter -> sample at 8 kHz
-> compand to 8 bits -> 64000 bit/s (one DS0)
T1 = 24 DS0 + 1 framing bit
= 193 bits every 125 us = 1.544 Mbit/s
E1 = 32 timeslots x 8 bits
= 256 bits every 125 us = 2.048 Mbit/s
(TS0 framing, TS16 signalling, 30 voice channels)
PLAIN9.9.4 what is really happening inside#
- The original transmitter was the carbon microphone: a chamber of carbon granules behind a diaphragm.
- Squeeze the granules and they touch better, so resistance falls. Release them and resistance rises.
- A battery pushes current through the granules, so sound modulates a current far stronger than the sound itself.
- That makes the carbon microphone an amplifier as well as a microphone, which is why telephones worked over kilometres of wire decades before the electronic amplifier existed. It is the single reason early telephony was commercially possible.
- The exchange supplies power down your own line, which is why a landline kept working when the house lost electricity.
- Automatic switching started with an undertaker’s grudge. Almon Brown Strowger believed an operator was diverting calls to a rival and patented an automatic exchange in 1891.
- His step-by-step switch used the dial pulses directly. The first digit’s pulses ratcheted a contact arm up one of ten rows, and the second digit’s pulses rotated it to one of ten contacts in that row. Ten by ten is a hundred lines, and switches were stacked in stages for longer numbers.
- That is why old dials sent pulses, and why “dialling” is still the word.
- Signalling, meaning the instructions about a call rather than the call itself, originally travelled in the same audio path as the voice, as tones.
- That was elegant and disastrous: anyone able to play the right tone into a handset could instruct the network. Out-of-band signalling replaced it, so call control now travels on its own separate data network.
a 1960s long-distance call, end to end
handset -> local loop (copper pair, -48 V from exchange)
-> local exchange (switch)
-> channel bank: filter, sample, compand -> DS0
-> multiplexed into a T1/E1 with other calls
-> trunk exchanges, path reserved end to end
-> demultiplex -> DAC -> local loop -> handset
the path is held for the duration of the call,
whether anyone is speaking or not.
TECHNICAL9.9.5 the engineer’s version#
- Alexander Graham Bell was granted US patent 174,465 on 7 March 1876, for “Improvement in Telegraphy”, covering transmission of vocal sounds by electrical undulations. The first intelligible sentence followed on 10 March
- Elisha Gray filed a caveat on 14 February 1876, and Antonio Meucci had filed caveat 3335 in 1871, which lapsed in 1874. The priority dispute has never fully died down.
- Edison filed for a carbon transmitter on 27 April 1877; US patent 222,390 was granted in 1879.
- Strowger’s automatic exchange patent dates from 1891, and the first installation opened in La Porte, Indiana, on 3 November 1892 with about 75 subscribers. Crossbar switches displaced step-by-step from the late 1930s, and stored-program electronic switching from the mid-1960s.
- Local loop electrical figures, which are conventions of the North American and most European networks rather than one global standard: -48 V DC central office battery, 20 to 50 mA loop current off-hook, ringing at about 90 V RMS at 20 Hz in North America and 25 Hz in much of Europe. Loops are typically 24 or 26 AWG copper, up to roughly 5 km without loading coils.
- G.711, released by the CCITT in 1972, defines the 8 kHz, 8-bit, 64 kbit/s voice channel. A-law compands a 13-bit signed linear sample to 8 bits and is used in most of the world; mu-law compands 14 bits to 8 and is used in North America and Japan. Algorithmic delay is 0.125 ms with no look-ahead.
- Multiplexing: the North American DS1 (T1) carries 24 DS0s plus one framing bit in a 193-bit frame every 125 microseconds, giving 1.544 Mbit/s. The CEPT E1 carries 32 timeslots of 8 bits in a 256-bit frame every 125 microseconds, giving 2.048 Mbit/s, with timeslot 0 for framing and timeslot 16 usually for signalling, leaving 30 voice channels.
- SS7, Signalling System No. 7, grew out of Bell System work in the 1970s and was standardized by the ITU-T as the Q.700-series recommendations in 1980. It moved call control onto a separate packet network. Links run at 56 or 64 kbit/s, with high-speed links at 1.5 or 2.0 Mbit/s. ISUP sets up and tears down voice circuits, and SIGTRAN carries SS7 over IP using SCTP.
- SS7 was designed for a closed club of national carriers and has essentially no authentication between operators. Since 2014, published research and reported incidents have demonstrated subscriber location tracking and SMS interception, which is the concrete reason SMS is now considered a weak second factor for authentication.
WORDS9.9.6 remember these#
- Local loop — the pair of wires to your house — the subscriber line from premises to the serving exchange.
- Circuit switching — a reserved path for the whole call — end-to-end resource allocation held for the call duration.
- Companding — squeezing loud and stretching quiet — logarithmic quantization per ITU-T G.711 A-law or mu-law.
- DS0 — one voice channel — 64 kbit/s, 8 kHz at 8 bits.
- TDM — sharing one cable by taking turns — time-division multiplexing into fixed 125 microsecond frames.
- SS7 — the network’s own instruction channel — out-of-band common channel signalling, ITU-T Q.700 series.
- Voice band — the slice of hearing a phone carries — 300 to 3,400 Hz.
9.10 Phone calls, part 2 - mobile and internet#
PLAIN9.10.1 in simple words#
- A mobile phone cannot send a wave down a wire, so it sends numbers by radio.
- Radio time is scarce and expensive, so those numbers must be squeezed hard. A digital landline used 64,000 bits a second per call, and early mobile used about 13,000.
- To get that small, mobile phones do not send the sound at all.
- They send a description of how your throat and mouth were shaped, updated fifty times a second, and the far end builds a fresh voice from that recipe.
- That is why mobile voice sounds correct but slightly synthetic, and why music down a phone sounds ruined. The recipe assumes a human voice.
- Internet calls, such as WhatsApp or a work meeting, work differently again.
- They chop the sound into small parcels of about twenty milliseconds, address each one, and throw it onto the ordinary internet.
- Nobody reserves anything. The parcels take their chances with everyone else’s video, downloads and games.
- Most arrive. Some arrive late. Some never arrive at all.
PLAIN9.10.2 a picture in your head#
- The old phone call was a private railway track, held for you alone.
- An internet call is posting numbered postcards, one every twenty milliseconds, and hoping.
- Postcards may arrive out of order, bunched up, or not at all.
- So the receiver keeps a small tray. Postcards land in it, get sorted, and are read out at a perfectly steady rhythm from the tray rather than as they arrive. That tray is the jitter buffer.
- If a postcard is missing when its turn comes, the receiver does not wait. It invents a plausible twenty milliseconds from what came before.
- Asking for a re-send would be pointless. By the time the replacement arrived, that moment of the conversation would be long past.
Where this comparison breaks: postcards are independent, but audio frames are not. Modern codecs predict each frame from the last, so a lost frame degrades the next few as well. Opus reduces this by optionally embedding a low-quality copy of the previous frame inside each new one, a form of forward error correction.
PLAIN9.10.3 a worked example#
- Cost out a plain internet call using G.711, the old 64 kbps telephone codec, over IPv4.
- We send one packet every 20 milliseconds, so 50 packets a second.
- 20 ms of G.711 is 8,000 samples a second times 0.02 s = 160 samples of one byte each, so 160 bytes of audio.
- Now the addressing: RTP header 12 bytes, UDP header 8 bytes, IPv4 header 20 bytes. That is 40 bytes of overhead.
- Total 200 bytes per packet, so 200 x 8 x 50 = 80,000 bits per second. A 64 kbps codec costs 80 kbps on the wire, and a quarter of the traffic is envelopes.
- Repeat with Opus at 24 kbps and 20 ms frames: 60 bytes of audio plus the same 40 bytes of headers = 100 bytes, so 40 kbps on the wire.
- The fixed 40 bytes dominates at low bitrates, which is why voice codecs settled on 20 ms frames. Shorter frames would multiply the header cost.
| Codec | Payload bitrate | On the wire, IPv4 |
|---|---|---|
| G.711 (64 kbps) | 64.0 kbps | 80.0 kbps |
| G.729 (8 kbps) | 8.0 kbps | 24.0 kbps |
| AMR-NB (12.2 kbps) | 12.2 kbps | 28.2 kbps |
| Opus (24 kbps) | 24.0 kbps | 40.0 kbps |
PLAIN9.10.4 what is really happening inside#
- A 2G GSM phone runs the GSM Full Rate codec: 13 kbit/s, working on frames of 160 samples at 8 kHz, which is 20 milliseconds.
- It is a vocoder. It fits a mathematical model of the vocal tract to your speech, then transmits the model’s settings and a description of the buzz or hiss driving it, rather than the waveform itself.
- That is why cellular voice sounds thin: band-limited to about 3.4 kHz, and then re-synthesized rather than reproduced.
- Setting up an internet call needs two separate things, and confusing them is the classic beginner error.
- Signalling is the conversation about the call: who you want, are they there, which codecs can we both do, are they ringing, have they answered. SIP does this.
- Media is the actual sound. RTP does this, over UDP, which never retransmits anything.
- Why UDP and not TCP? Because TCP guarantees delivery by retransmitting, and waiting for a retransmission is worse than the loss it fixes.
- So voice deliberately chooses a protocol that gives up on lost data, then repairs the hole locally.
- Packet loss concealment extends the last pitch period and fades it over a few tens of milliseconds. One or two lost packets are genuinely hard to hear.
- Echo is the other big problem. Your speaker plays my voice, your microphone picks it up, and I hear myself two hundred milliseconds later, which is far more disruptive than noise.
- Acoustic echo cancellation works because the device already knows exactly what it sent to its own speaker. It learns a filter describing the room path from speaker to microphone, predicts the echo, and subtracts it.
- That is why speakerphone in a bathroom is hard. The room path is long, changing and full of reflections, so the filter cannot keep up.
TECHNICAL9.10.5 the engineer’s version#
- AMR-NB frames are 160 samples at 8 kHz, 20 ms long, and the network can switch mode frame by frame to trade voice bits against error protection as radio conditions change.
- AMR-WB samples at 16 kHz with a 50 to 7,000 Hz passband and was developed by Nokia and VoiceAge. It is what carriers market as HD Voice. EVS, Enhanced Voice Services, extends this to superwideband and fullband.
- SIP is defined in RFC 3261 (2002), replacing RFC 2543 (1999). It is a text request/response protocol resembling HTTP, on port 5060 plain and 5061 over TLS. Core methods are INVITE, ACK, BYE, CANCEL, REGISTER and OPTIONS. Media parameters are negotiated by carrying SDP bodies in an offer/answer exchange.
- RTP is defined in RFC 3550 (2003), replacing RFC 1889 (1996). Its 12-byte header carries a 7-bit payload type, a 16-bit sequence number for loss and reorder detection, a 32-bit media timestamp, and a 32-bit synchronization source identifier. RTCP reports loss, jitter and round-trip time to the sender.
- Jitter buffers are adaptive: they measure arrival variation and resize themselves, typically holding 20 to 100 ms, trading added delay against the probability that a packet arrives too late to use.
- ITU-T Recommendation G.114 gives the delay budget: under 150 ms one-way mouth-to-ear is acceptable for essentially all applications, 150 to 400 ms is usable but degraded, and over 400 ms is unacceptable for conversation. Line echo cancellers are specified in ITU-T G.168.
- VoLTE carries voice as IP over an LTE bearer, controlled by IMS, the IP Multimedia Subsystem, rather than by a circuit-switched core. The first commercial launch was by MetroPCS in Dallas in August 2012. Its media plane uses AMR-NB as a minimum, AMR-WB for HD Voice, and EVS in newer deployments.
- The key architectural difference from an over-the-top application: VoLTE media rides a dedicated bearer with a guaranteed bit rate and a strict quality-of-service class. A WhatsApp call is best-effort traffic on the same data bearer as everything else, is end-to-end encrypted by the application, uses Opus, and crosses the operator’s network as ordinary internet data. Both are voice over IP. Only one has the network holding capacity open.
- Voice tolerates loss and hates delay because it is an isochronous stream consumed in real time: a 20 ms hole is concealed, but a 400 ms delay breaks turn-taking, which humans manage on a scale of about 200 ms. File transfer is the exact opposite. It is verified in full at the end, so one wrong byte ruins it, while an extra thirty seconds costs nothing. That is why one runs on UDP with concealment and the other on TCP with retransmission.
| Codec | Bitrate | Standardized |
|---|---|---|
| GSM Full Rate (RPE-LTP) | 13 kbit/s | Early 1990s |
| GSM Half Rate | 5.6 kbit/s | 1990s |
| GSM Enhanced Full Rate | 12.2 kbit/s | 1990s |
| AMR-NB, 8 modes | 4.75 to 12.2 kbit/s | 3GPP, Oct 1999 |
| AMR-WB / G.722.2 | 6.6 to 23.85 kbit/s | ITU-T, 2002 |
a SIP call setup, simplified
caller callee
|------ INVITE (with SDP offer) --->|
|<----- 100 Trying -----------------|
|<----- 180 Ringing ----------------|
|<----- 200 OK (with SDP answer) ---|
|------ ACK ----------------------->|
|==== RTP audio over UDP, both ways =|
|------ BYE ----------------------->|
|<----- 200 OK ---------------------|
WORDS9.10.6 remember these#
- Vocoder — a codec that rebuilds your voice rather than copying it — a parametric coder transmitting vocal tract model coefficients.
- RTP — the protocol carrying the sound — Real-time Transport Protocol, RFC 3550, over UDP.
- SIP — the protocol that rings the phone — Session Initiation Protocol, RFC 3261, with SDP offer/answer.
- Jitter buffer — the tray that smooths arrival times — an adaptive de-jitter queue trading latency for late-packet loss.
- PLC — filling in a lost moment — packet loss concealment by pitch-period extrapolation.
- AEC — stopping the caller hearing themselves — acoustic echo cancellation by adaptive filtering of the known playback signal.
- VoLTE — carrier calls over the data network — IMS-controlled voice on a dedicated LTE bearer with guaranteed bit rate.
- Best effort — no promises about delivery — the default internet service model with no reserved capacity.
9.11 Noise cancelling, spatial audio and the modern tricks#
PLAIN9.11.1 in simple words#
- There are two ways to make the world quieter in headphones.
- Passive means physically blocking sound: a tight seal, heavy cups, foam. It is just a wall.
- Active means fighting sound with sound.
- A microphone on the outside listens to the noise arriving, and a chip works out the exact opposite wave: where the noise pushes, the opposite pulls.
- The speaker plays that opposite. Push and pull arrive together and cancel.
- This works well for low rumbling noise such as an aircraft or a bus engine, and poorly for sharp high noise such as speech or a fork on a plate.
- Separately, phones use several microphones at once to work out which direction a sound came from and keep only the direction you speak from. That is beamforming.
- And headphones can fool you into hearing a sound as if it came from your left, above, or behind, by copying the way your own head would have changed it.
PLAIN9.11.2 a picture in your head#
- Imagine two people pushing a swing.
- If they push together at the right moments, the swing goes higher. That is adding two waves in phase.
- If the second pushes exactly when the first pulls, with equal strength, the swing does not move at all. That is cancellation.
- Noise cancelling is the second person, and the chip’s whole job is getting the timing exactly right.
- To push at exactly the wrong moment, you must know when that moment is before it happens. For a slow swing that is easy. For one wobbling twenty thousand times a second, the chip has microseconds to hear, calculate and play.
Where this comparison breaks: cancellation is only exact at one point in space. Move your head a few centimetres and the arrival times change, so what cancelled now partly adds. At low frequencies the wave is metres long, so a few centimetres hardly matter, which is the real reason active cancellation is a low-frequency technique.
PLAIN9.11.3 a worked example#
- Take a 100 Hz engine rumble. Its wavelength is 3.43 metres.
- To cancel it you must be right within a small fraction of that, say 10 cm, which is 3 percent of a wavelength, about 10 degrees of phase error.
- 10 cm of air is 0.29 milliseconds, comfortable for a chip running millions of operations a second.
- Now take a 5,000 Hz hiss. Its wavelength is 6.9 cm.
- The same 10 cm of positional error is now more than a whole wavelength. Your cancellation arrives at some arbitrary phase and may make it louder.
- That single calculation is why every noise cancelling headphone works well below about 1 kHz and relies on the ear cup seal above it.
- Beamforming numbers: two phone microphones 14 cm apart. Sound from directly ahead reaches both at the same instant, while sound from the side arrives 0.14 / 343 = 408 microseconds later at the far microphone. Delay one microphone by 408 microseconds and add, and side sounds partly cancel while front sounds double.
PLAIN9.11.4 what is really happening inside#
- Active noise cancellation comes in three arrangements.
- Feedforward: a microphone outside the cup hears the noise before it reaches your ear, so there is time to compute. It cannot tell what actually arrived at the eardrum.
- Feedback: a microphone inside the cup hears the result, including whatever the headphone did, and corrects the error. It has almost no time, so it works only at low frequencies.
- Hybrid: both. Nearly all current premium headphones do this.
- Now spatial audio. Your brain locates sounds with three main clues.
- Interaural time difference: sound from your left reaches the left ear first, by up to about 660 microseconds.
- Interaural level difference: your head shadows the far ear, mostly at high frequencies where the head is large compared with the wavelength.
- Spectral shaping: the folds of your outer ear boost and notch particular frequencies depending on whether a sound comes from above, in front or behind. That is what resolves up from down, which timing alone cannot.
- All three together, measured as a filter for every direction, are the head-related transfer function, or HRTF.
- Headphone spatial audio applies that filter to a sound, so your brain receives the clues it would have received from a real source.
- Head tracking makes it far more convincing, because when you turn your head the virtual source stays put, exactly as a real one would.
TECHNICAL9.11.5 the engineer’s version#
- Paul Lueg was granted US patent 2,043,416 in 1936 for cancelling sinusoidal tones in a duct by phase advance. Lawrence Fogel patented cockpit noise cancellation in the 1950s. Bose shipped an active noise cancelling aviation headset in 1989 and the consumer QuietComfort headphone in 2000.
- Practical performance for current premium over-ear models: roughly 20 to 30 dB of active attenuation below 300 Hz, tapering to near zero by 1.5 to 2 kHz, on top of 10 to 25 dB of passive attenuation above 1 kHz. Manufacturers quote the sum and rarely publish the curve.
- Feedforward controllers are typically fixed or slowly adaptive FIR filters. Feedback loops are analogue or very low-latency digital, and are limited by the Bode sensitivity integral, which guarantees that attenuation at one frequency must be paid for with amplification at another. That is why poorly designed cancellation adds a faint hiss or a pressure sensation around 2 to 4 kHz.
- Beamforming: delay-and-sum is the simplest form. Adaptive methods, notably minimum variance distortionless response, steer nulls at interfering sources instead. A two-microphone phone array gives modest directivity, while laptop and smart-speaker arrays of four to eight microphones do far better.
- Channel layouts, as a standard: 2.0 stereo; 5.1, meaning left, centre, right, surround left, surround right and a low-frequency effects channel limited to about 120 Hz; 7.1; and object-based systems such as Dolby Atmos, introduced in 2012, which transmit audio objects with position metadata and render them to whatever speakers exist.
- HRTF datasets used in research include the CIPIC database and the MIT KEMAR measurements. Rendering convolves source audio with the left and right HRTF impulse responses for the wanted direction, interpolating between measured points.
- Active research, not settled fact: individualized HRTFs. Generic HRTFs give weaker elevation cues and more front-back confusion for listeners whose ear geometry differs from the measurement dummy. Approaches under investigation include photogrammetry from phone cameras, anthropometric regression and machine-learned personalization. Vendor claims of personalized spatial audio from a short ear scan are an engineering approximation, not a solved problem.
WORDS9.11.6 remember these#
- Anti-phase — the exact opposite wave — a signal inverted by 180 degrees at the frequency of interest.
- Feedforward ANC — the outside microphone — control derived from the disturbance before it reaches the ear.
- Feedback ANC — the inside microphone — closed-loop correction of the residual error at the ear.
- Beamforming — listening in one direction — spatial filtering across a microphone array by delay-and-sum or adaptive weights.
- ITD and ILD — which ear hears it first and loudest — interaural time and level differences, the primary horizontal localization cues.
- HRTF — how your own ears colour a sound by direction — head-related transfer function, the direction-dependent filter of head, torso and pinna.
- Object audio — sounds with positions rather than channels — metadata-driven rendering to an arbitrary speaker layout.
9.98 Common wrong ideas#
- Wrong: higher bitrate always sounds better. Right: bitrate matters only until the codec has enough bits to hide its errors under the masking threshold. Beyond that, extra bits change nothing audible, and a good encoder at 96 kbps AAC can beat a poor encoder at 256 kbps MP3.
- Wrong: more watts means louder. Right: loudness follows speaker sensitivity far more than amplifier power. Doubling power adds only 3 dB, while a speaker 6 dB more sensitive is as loud on a quarter of the power.
- Wrong: lossless always sounds audibly better than lossy. Right: lossless is guaranteed bit-identical, which matters for archiving, editing and re-encoding. In blind listening at modern high bitrates, most listeners cannot reliably tell them apart. Keep lossless because it is future-proof, not because you can hear it.
- Wrong: the microphone records the sound. Right: the microphone only converts pressure into a voltage. Nothing is recorded until something samples that voltage and stores numbers. In a condenser or MEMS microphone the sound does not even supply the signal energy; it modulates a bias already present.
- Wrong: digital audio is a staircase, so it is less smooth than analogue. Right: the staircase is an intermediate artefact inside the converter. The reconstruction filter recovers the unique smooth band-limited wave passing through the samples, so there is no jaggedness in the output.
- Wrong: MP3 works by cutting off high frequencies. Right: that is one bit-allocation decision some encoders make. The mechanism is shaping quantization noise to hide beneath louder sounds across the whole spectrum.
- Wrong: noise cancelling headphones cancel all noise. Right: they work well below roughly 1 kHz, because cancellation must be accurate to a fraction of a wavelength. Higher frequencies are blocked by the physical seal instead.
- Wrong: a WhatsApp call and a mobile phone call are the same thing on the same network. Right: both are voice over IP, but a carrier VoLTE call rides a dedicated bearer with guaranteed bit rate under IMS control, while an application call is best-effort internet traffic with no reserved capacity.
- Wrong: a bigger speaker is always better. Right: a large cone can move enough air for bass but is too heavy to reproduce treble accurately, which is why real systems split the range between drivers sized for each job.
- Wrong: 44.1 kHz means the system can handle only 44,100 different sounds. Right: it means 44,100 measurements per second, which by the sampling theorem perfectly captures every frequency below 22,050 Hz.
9.99 Chapter summary in 20 lines#
- Sound is a travelling pattern of air pressure: compressions and rarefactions.
- Frequency in hertz sets pitch, amplitude sets loudness, and wavelength is the speed of sound, about 343 m/s at 20 C, divided by frequency.
- Decibels are logarithmic because hearing responds to ratios: +6 dB doubles pressure and +10 dB multiplies power by ten.
- A speaker is a coil in a magnet’s gap glued to a cone. Current makes force by F = Bl*I, the cone moves, the air moves, and you hear it.
- Faraday’s law runs the same machine backwards, which is why a moving cone generates voltage and why a dynamic microphone needs no power.
- Big cones move enough air for bass but are too heavy for treble, so systems split the band with crossovers, and enclosures stop the back wave cancelling the front.
- A phone speaker cannot make bass because it cannot displace enough air, by about three orders of magnitude. That is physics, not software.
- Condenser, electret and MEMS microphones modulate a bias voltage rather than generating the signal energy from the sound itself.
- Sampling captures a wave perfectly if the rate exceeds twice the highest frequency present, which is the Nyquist-Shannon theorem.
- Anti-aliasing filtering before the converter is mandatory, because folded frequencies can never be separated afterwards.
- Bit depth sets dynamic range at about 6 dB per bit, so 16 bits gives roughly 96 to 98 dB, and dither turns ugly quantization distortion into benign hiss.
- Almost every audio converter today is oversampling sigma-delta, in both directions.
- CD audio is 44,100 samples a second, 16 bits, two channels, which is 1.4112 Mbit/s and about 31.75 MB for a three-minute song.
- Lossless codecs such as FLAC predict and store the error, reaching 50 to 70 percent, while lossy codecs such as MP3, AAC and Opus discard what masking makes inaudible and reach a tenth or less.
- A codec is a coding method and a container is the file that holds it, which is why MP4 can hold either AAC or ALAC and why extensions mislead.
- Records store the waveform as a physical groove, tape stores it as aligned magnetic particles, and the CD stores numbers as pit-to-land transitions read by a 780 nm laser.
- Computer audio is a race against a DMA deadline: buffer size sets latency, and missing the deadline puts a step in the waveform heard as a click.
- The old telephone reserved a whole path per call and digitized voice as 8 kHz, 8-bit companded samples giving 64 kbit/s, multiplexed into 1.544 Mbit/s T1 or 2.048 Mbit/s E1 and controlled by SS7.
- Mobile voice sends vocal tract parameters instead of a waveform at around 13 kbit/s, which is why it sounds thin, while internet calls send RTP over UDP with jitter buffers, loss concealment and echo cancellation.
- Voice tolerates loss but hates delay, because a 20 ms hole can be concealed and a 400 ms delay destroys turn-taking, while file transfer is the opposite, which is why one uses UDP and the other TCP.