Walk the independent claim. Apple's grant US11546687B1, "Head-tracked spatial audio" (issued January 3, 2023; inventors including Symeon Delikaris Manias and Juha Merimaa), is a granted patent — the B1 kind code indicates a grant without prior publication. Its CPC tags H04R 1/326 and H04R 1/323 are directional-transducer classes, sitting beside spatial-rendering classes, and the claim text reveals a specific signal-processing recipe rather than a vague "audio follows your head" idea.

Claim 1 is a three-step method. First, generate a plurality of spatial filters that map the response of an audio capture device — one with multiple microphones — to a set of head-related transfer functions (HRTFs) for different positions of the capture device relative to those HRTFs. Second, determine a current set of spatial filters based on that plurality and on the user's head position. Third, convolve the microphone signals with the current set of filters to produce output binaural audio channels. The mechanism is therefore a precomputed filter bank indexed by head position: rather than re-deriving the soundfield from scratch each frame, the system stores filters that already fold the microphone array's directional response into the listener's HRTFs, then selects or blends among them as the head moves.

“Spatial filters are generated that map response of an audio capture device to head related transfer functions (HRTFs) for different positions of the audio capture device relative to the HRTFs.”— U.S. Patent No. 11,546,687 source

The dependent claims expose how the filter selection and the filters themselves are constructed. Claim 3 selects one or more of the stored filters as the current set based on head position; claim 4 interpolates among the selected filters — so head positions between the sampled grid points are handled by blending, not by snapping to the nearest stored filter, which is what keeps the spatial image smooth as the listener turns. Claim 9 makes each stored filter correspond to a different head position, and claim 10 defines head position by at least roll, pitch, or yaw. Claim 2 adds a notable efficiency property: the binaural channels are generated directly from the microphone signals "without creating an intermediate format" — no detour through ambisonics or another scene representation, which trims latency and computation. Claim 5 prunes the filter bank by discarding filters whose reproduction error exceeds a threshold and storing the rest on-device or on a connected device.

Claim 6 is the most revealing about the disclosed method. It represents the spatial filters as beamforming filters and builds them in three moves: generate virtual beamformed speakers from the directivity pattern of the microphones, generate virtual beamformed microphone pickups from the HRTFs, and compute beamforming filters that map the virtual speakers onto the virtual pickups. That is a concrete recipe for turning a microphone array's geometry plus a listener's HRTFs into a single convolution that renders binaural, head-locked audio. Claims 7 and 8 then close the loop to hardware: the head position is generated by sensors integrated with a headphone set or head-mounted display, and the output binaural channels drive that set's left and right speakers to produce the spatial effect. The system and electronic-device claims (11 and 17) recast the same pipeline, with claim 17 phrased around accessing a pre-existing filter bank rather than generating it.

So the element doing the work is not "head tracking" or "spatial audio" in the abstract; it is the precomputed, head-position-indexed beamforming-filter bank, selected or interpolated by measured head orientation and convolved with live microphone signals to yield binaural output without an intermediate scene format. Plain stereo moves with your head; this method does the opposite, anchoring a virtual source in the room because the renderer continuously swaps in the filter set matched to the head's current roll, pitch, and yaw.

What it reads on is the spatial-audio feature in Apple's earbuds and headphones — the experience marketed where audio appears to come from a fixed point regardless of head movement. The directional-transducer classification ties the claim to the physical audio hardware as well as the rendering math.

Scope discipline: the claim does not own spatial audio, and it does not own head tracking. It protects the recited coupling — a head-position-indexed bank of spatial/beamforming filters convolved with multi-microphone signals to produce binaural channels. A system that renders fixed stereo, or that does head-locked audio through a fundamentally different pipeline (say, a full ambisonic scene rotated in real time, which is the "intermediate format" claim 2 explicitly excludes), may operate outside the narrower claims. The defensible element is the filter-bank-plus-convolution architecture, not the user-facing effect alone.

Granted status makes US11546687B1 a live consideration in a contested space — Apple, Meta, Sonos, Dolby, and Bose all hold spatial-audio IP. For anyone building head-tracked audio into earbuds or a headset, this grant is part of the freedom-to-operate map specifically on the head-tracking-plus-rendering combination, and claim 6's beamforming construction is the limitation a competing implementation should be measured against. For a strategist, the patent marks Apple's claim not to the broad category but to the particular, low-latency method behind the experience it popularized.

The pruning and storage details in claim 5 deserve a second look, because they reveal that the claimed system is engineered for an on-device budget, not a studio. Generating filters for every possible head position would be both wasteful and inaccurate, so the method discards filters whose reproduction error exceeds a threshold and stores the survivors on the computing device or a connected one. That, combined with claim 4's interpolation between stored filters and claim 2's avoidance of an intermediate format, describes a pipeline tuned to run on the limited compute of earbuds or a headset paired to a phone: a compact, error-screened filter bank, blended on the fly, convolved directly with the live microphone feed. The three independent claims — method (1), audio processing system (11), and electronic device (17) — repeat the core selection-and-convolution loop, with claim 17 framed around a device that merely accesses a pre-built bank, extending protection to a product that loads filters generated elsewhere rather than computing them itself.