Life on the audio thread

Sống trên audio thread

Real-time audio on a phone is a scheduling problem with human consequences. The operating system calls your render function every N samples and expects a full buffer of PCM floats back before the hardware drains what it already has. At 48 kHz with a 512-sample block, the deadline lands every 10.67 ms. All night. A few million times per sleep session. Miss it once and the listener hears a click or a pop. Miss it repeatedly and Estua's sleep sound turns unusable at 3 AM, which becomes a one-star review, which becomes a real problem for a studio with five apps and no venture capital cushion.

Sonarish runs the same gauntlet in the opposite direction. Its microphone capture path receives samples from the audio HAL on a high-priority thread, and whatever happens inside that callback decides whether the noise meter reads truth or garbage. Two apps, one discipline. This post writes the discipline down.

The deadline, derived

Start with arithmetic. One sample period at 48 kHz lasts 1/48000 s ≈ 20.8 µs. A 512-sample block therefore spans 512/48000 ≈ 10.67 ms, and that interval is the ceiling for everything the app does per callback: every noise generator, every filter, every mix stage, plus whatever the OS itself burns on resampling and the Bluetooth stack before your floats reach a speaker. We aim to use well under half of it. Bench note: on an iPhone 12 and a Pixel 6a we logged callback wake-up timestamps through one 8-hour Estua session per device; median wake-up jitter stayed under 0.3 ms, but Bluetooth reconnects produced isolated wake-ups arriving 2 ms late. A render chain that needs 9 ms of the block has no armor against a late wake-up. One that needs 1 ms shrugs.

On iOS, Estua uses AVAudioEngine with a manual source node callback. On Android it uses AudioTrack with ENCODING_PCM_FLOAT in low-latency mode when the OEM supports it. Both platforms invoke the render function on a real-time priority thread that bypasses the normal UI run loop. The scheduler promises to wake that thread on time. What happens if the thread then waits on something slow is nobody's promise, and every rule below exists because of that gap.

Why malloc is the enemy

The audio thread must never wait on contended locks, never allocate from the heap, never perform I/O, and never call code that might do any of those things transitively. The problem is variance. A malloc on a fragmented heap might take 200 microseconds on a good day and 5 milliseconds on a bad day, and 5 ms is half the block gone before a single sample gets rendered. Bench note from the same Pixel 6a: after fragmenting a test heap with 1 million mixed-size allocations, we timed 100,000 further small mallocs. Median 0.1 µs. Worst single call 4.1 ms. That tail is the click.

Locks fail differently. A real-time thread that blocks on a mutex held by a background thread inherits that background thread's scheduling luck, and on a loaded phone the holder may not run again for several milliseconds. Priority inheritance rescues you on some kernel versions and silently does nothing on others. The transitive clause is the sneaky part: an innocent call into a string formatter, a Swift protocol witness, or a convenience API from the OS can allocate on your behalf three stack frames down. If we cannot read the code path to the bottom, it stays off the audio thread.

Rules we enforce in code review

These are review blockers, not guidelines. ktuyen files violations as bugs and atuan will not tag a release while one stays open.

  • No heap allocation in the callback. No Swift Array append, no Kotlin list growth, no std::vector push_back, no NSString formatting. Every buffer is preallocated at session start on the UI thread and reaches the callback as a pointer in its context struct.
  • No locks that can block. Data crosses between the audio thread and the UI thread through lock-free single-producer single-consumer ring buffers, in both directions.
  • No logging, file writes, or network calls, even guarded by debug flags. Debug builds ship to ktuyen's overnight soak tests, and one forgotten print statement that allocated a string has caused more clicks than any DSP bug we have written.
  • On iOS, no Objective-C message sends on the hot path. The callback stays in C or in Swift struct code the compiler can inline.
  • Parameters are computed off-thread. Biquad filter coefficients, LFO rates, and scene parameters are derived on the UI thread when the user changes a setting; the callback reads only const structs that were written once and are never mutated concurrently.

The last rule deserves its own section, because it is the one engineers get wrong in the most interesting ways.

Publishing parameters without locks

When you drag Estua's ocean scene toward heavier surf, the app has to hand new filter cutoffs to a thread it is forbidden to lock against. The pattern we use is a two-slot swap. Parameters live in two preallocated slots, an atomic integer names the slot the audio thread may read, and the UI thread always writes the spare slot before publishing it.

// two preallocated slots, one atomic index
struct Params { cutoffHz, gain, lfoRate[4] }   // plain floats only
slots     = [Params, Params]     // filled at session start
activeIdx = AtomicInt(0)         // slot the audio thread reads

// UI thread, when a setting changes:
spare = 1 - activeIdx.load(acquire)
slots[spare] = computeParams()   // biquad coeffs, LFO rates, scene
activeIdx.store(spare, release)  // publish

// audio callback, once per block:
p = slots[activeIdx.load(acquire)]   // copy by value, tens of bytes
renderBlock(out, 512, p)

The release/acquire pair guarantees the callback never observes a half-written slot, and the copy costs tens of bytes. Stretch the same idea around a power-of-2 array with separate read and write indices and you get the SPSC ring buffers that carry Sonarish's captured PCM out and Estua's metering data up to the UI. No mutex appears anywhere in either app's audio path. One honest caveat: two settings changes in very quick succession could in principle reuse a slot mid-read. The per-block copy is fast enough that we have never observed it, and a triple buffer would close the gap if we ever did.

How Estua fits inside the budget

Estua's synthesis chain, described end to end in our non-repeating audio post, generates pink noise with Paul Kellet's economical recursive filter: three one-pole state updates of the form b0 = 0.99765·b0 + w·0.0990460 plus a weighted sum, a handful of multiplies and adds per sample with no FFT per block. Slow amplitude movement comes from one-pole low-pass filters on the absolute value of the signal envelope, y[n] = y[n−1] + α·(x[n] − y[n−1]), running at the full sample rate with precomputed coefficients. Multiple uncorrelated LFOs modulate filter cutoff and gain at 0.02–0.15 Hz, so the output never repeats on timescales a human can track. The product reasoning for synthesis over loops is in our sleep and waves post.

All of it fits in roughly 3–8% of one CPU core on modern phones, comfortably inside the 10.67 ms block, leaving headroom for Bluetooth audio output and for the moments when the OS briefly starves a backgrounded app. Headroom is the feature. A chain that saturated the block would pass a bench test and fail on a random Tuesday night.

Buffer sizes OEMs actually hand you

Not every Android device wants 512 samples. Some OEM audio paths request 192 or 256 for their DSP route, and fighting the platform's preferred size costs you the low-latency path entirely. Estua adapts the requested buffer size at session start and keeps every piece of internal DSP state, filter memories and LFO phases included, continuous across the change, so a resize never produces a discontinuity click. Only the per-block bookkeeping changes. The synthesis state carries straight through.

On the capture side, Sonarish's callback does one job: copy incoming PCM into the ring buffer and return. The FFT engine consumes that ring on a background queue and never runs on the audio thread, which is how the app computes a 4096-point transform while the meter updates at 30 Hz without glitches. Transform details live in the FFT post. The architectural sentence worth memorizing is short: the audio thread copies samples and returns, and everything else happens elsewhere.

Debugging the clicks that still happen

Clicks still happen. When one appears we profile the audio thread directly: Instruments Time Profiler with the audio thread marked on iOS, Systrace with Trace.beginSection markers on Android. In our experience 99% of production clicks trace to debug-only code that leaked into the hot path — a String format for logging, a guard assertion that allocates, an accidental Swift optional unwrap that lands on a slow path. The DSP math is almost never the culprit. The scaffolding around it usually is.

So every piece of debug instrumentation sits behind compile-time flags that we verify are off in release builds, and the final gate is ktuyen's overnight Estua soak test: 8 hours of continuous playback in airplane mode on real hardware before atuan tags any release. A click that survives review, the static checks, and a full night of playback has earned a bug report with its name on it.

Audio real-time trên điện thoại là bài toán scheduling mà hậu quả rơi thẳng vào người nghe. Hệ điều hành gọi hàm render mỗi N sample và đòi đủ một buffer PCM float trước khi phần cứng phát hết chỗ đang có. Ở 48 kHz, block 512 sample, deadline gõ cửa mỗi 10,67 ms. Cả đêm. Vài triệu lần mỗi phiên ngủ. Trễ một lần, người nghe dính tiếng click hoặc pop. Trễ liên tục thì âm ngủ của Estua thành vô dụng lúc 3 giờ sáng, thành review một sao, thành chuyện nghiêm túc với một studio năm app không có đệm vốn đầu tư.

Sonarish chạy đúng đường đua đó theo chiều ngược lại. Đường capture micro nhận sample từ audio HAL trên thread ưu tiên cao; mọi thứ xảy ra trong callback quyết định đồng hồ đo ồn hiện đúng hay hiện rác. Hai app, một kỷ luật. Bài này chép kỷ luật đó ra giấy.

Suy ra deadline từ số học

Bắt đầu bằng phép chia. Một chu kỳ sample ở 48 kHz dài 1/48000 s ≈ 20,8 µs. Block 512 sample chiếm 512/48000 ≈ 10,67 ms. Đó là trần cho mọi thứ app làm trong một callback: mọi bộ tạo noise, mọi filter, mọi tầng mix, cộng cả phần OS tự đốt cho resample và stack Bluetooth trước khi float của bạn tới loa. Mình nhắm dùng dưới một nửa. Ghi chú bench: trên một iPhone 12 và một Pixel 6a, mình log timestamp lúc callback thức dậy suốt một phiên Estua 8 giờ mỗi máy; jitter trung vị dưới 0,3 ms, nhưng lúc Bluetooth reconnect có cú thức dậy trễ tới 2 ms. Chuỗi render ngốn 9 ms mỗi block không đỡ nổi cú trễ đó. Chuỗi ngốn 1 ms thì nhún vai cho qua.

Trên iOS, Estua dùng AVAudioEngine với callback source node thủ công. Trên Android là AudioTrack, ENCODING_PCM_FLOAT, chế độ low-latency khi OEM hỗ trợ. Cả hai nền tảng gọi hàm render trên thread ưu tiên real-time, bỏ qua hẳn run loop UI. Scheduler hứa đánh thức thread đúng giờ. Còn thread thức dậy xong đi chờ thứ gì đó chậm thì scheduler không hứa gì hết. Toàn bộ luật phía dưới sinh ra từ khoảng trống này.

Vì sao malloc là kẻ thù

Audio thread không được chờ lock đang tranh chấp, không cấp phát heap, không I/O, không gọi bất kỳ code nào có thể làm mấy việc đó gián tiếp. Vấn đề nằm ở variance. Malloc trên heap phân mảnh có thể mất 200 µs ngày đẹp trời và 5 ms ngày xấu, mà 5 ms là bay nửa block trước khi render được sample nào. Ghi chú bench trên đúng con Pixel 6a: sau khi băm nát heap thử nghiệm bằng 1 triệu lần cấp phát đủ cỡ, mình đo thêm 100.000 lần malloc nhỏ. Trung vị 0,1 µs. Cú tệ nhất 4,1 ms. Cái đuôi đó chính là tiếng click.

Lock hỏng theo kiểu khác. Thread real-time block trên mutex do một thread nền giữ sẽ thừa hưởng luôn vận may scheduling của thread nền đó; máy đang tải nặng thì kẻ giữ lock có khi vài millisecond nữa mới được chạy tiếp. Priority inheritance cứu bạn trên kernel này, im lặng bó tay trên kernel khác. Khoản "gián tiếp" mới hiểm: một cú gọi vô hại vào string formatter, một protocol witness của Swift, một API tiện tay của OS đều có thể cấp phát giùm bạn ở ba stack frame phía dưới. Không đọc được code path tới đáy thì để nó đứng ngoài audio thread.

Luật chặn thẳng trong code review

Đây là blocker review chứ không phải khuyến nghị. ktuyen file vi phạm thành bug và atuan không tag release khi còn bug mở.

  • Không cấp phát heap trong callback. Không Swift Array append, không Kotlin list grow, không std::vector push_back, không format NSString. Mọi buffer cấp phát sẵn lúc mở session trên UI thread, vào callback qua pointer trong context.
  • Không lock nào có thể block. Data qua lại giữa audio thread và UI thread bằng ring buffer lock-free single-producer single-consumer, cả hai chiều.
  • Không log, không ghi file, không network, kể cả nấp sau debug flag. Bản debug chạy soak test qua đêm của ktuyen; một câu print bỏ quên từng cấp phát string đã gây click nhiều hơn bất kỳ bug DSP nào team viết ra.
  • Trên iOS, không message send Objective-C trên hot path. Callback nằm trong C hoặc code struct Swift mà compiler inline được.
  • Tham số tính ngoài thread. Hệ số filter biquad, tốc độ LFO và tham số scene tính trên UI thread lúc người dùng đổi setting; callback chỉ đọc struct const ghi một lần, không ai mutate song song.

Luật cuối đáng một mục riêng, vì đó là chỗ kỹ sư hay làm sai theo kiểu thú vị nhất.

Công bố tham số mà không cần lock

Bạn kéo scene biển của Estua sang nhiều sóng hơn, app phải chuyển cutoff filter mới cho một thread mà nó bị cấm lock. Mẫu mình dùng là hoán đổi hai slot. Tham số nằm trong hai slot cấp phát sẵn, một số nguyên atomic chỉ định slot audio thread được đọc, còn UI thread luôn ghi vào slot trống trước rồi mới công bố.

// hai slot cấp phát sẵn, một chỉ số atomic
struct Params { cutoffHz, gain, lfoRate[4] }   // chỉ float trơn
slots     = [Params, Params]     // đổ đầy lúc mở session
activeIdx = AtomicInt(0)         // slot audio thread đọc

// UI thread, khi đổi setting:
spare = 1 - activeIdx.load(acquire)
slots[spare] = computeParams()   // hệ số biquad, tốc LFO, scene
activeIdx.store(spare, release)  // công bố

// audio callback, mỗi block một lần:
p = slots[activeIdx.load(acquire)]   // copy theo giá trị, vài chục byte
renderBlock(out, 512, p)

Cặp release/acquire bảo đảm callback không bao giờ thấy slot ghi dở; bản copy chỉ vài chục byte. Kéo giãn đúng ý tưởng này quanh một mảng cỡ lũy thừa 2 với hai chỉ số đọc ghi riêng là ra ring buffer SPSC: chở PCM capture của Sonarish ra ngoài và số liệu metering của Estua lên UI. Không mutex nào xuất hiện trong đường audio của cả hai app. Một điểm thành thật: đổi setting hai lần quá nhanh, về lý thuyết, có thể ghi đè slot đang được đọc. Copy mỗi block đủ nhanh nên mình chưa từng quan sát thấy; nếu thấy thì nâng lên triple buffer là xong.

Estua sống trong định mức thế nào

Chuỗi synthesis của Estua, mô tả đầy đủ trong bài âm thanh không lặp, tạo pink noise bằng filter đệ quy bản tiết kiệm của Paul Kellet: ba cập nhật trạng thái one-pole dạng b0 = 0.99765·b0 + w·0.0990460 cộng một tổng có trọng số, vài phép nhân cộng mỗi sample, không FFT nào mỗi block. Biên độ trôi chậm nhờ filter one-pole low-pass trên giá trị tuyệt đối của envelope, y[n] = y[n−1] + α·(x[n] − y[n−1]), chạy đủ sample rate với hệ số tính sẵn. Nhiều LFO không tương quan ở 0,02–0,15 Hz điều chế cutoff và gain, nên output không lặp trên thang thời gian tai người bám theo được. Vì sao chọn synthesis thay vì loop, đọc bài giấc ngủ và sóng.

Tất cả gói trong khoảng 3–8% một nhân CPU trên máy đời mới, nằm thoải mái trong block 10,67 ms, chừa headroom cho output audio Bluetooth và những lúc OS bóp app chạy nền. Headroom chính là tính năng. Chuỗi synthesis ăn kín block sẽ qua được bài bench rồi rớt vào một tối thứ ba ngẫu nhiên nào đó.

Cỡ buffer OEM thật sự đưa cho bạn

Không phải máy Android nào cũng muốn 512 sample. Vài đường audio OEM đòi 192 hoặc 256 cho DSP path của họ. Cãi lại cỡ ưa thích của nền tảng là mất luôn đường low-latency. Estua chỉnh cỡ buffer yêu cầu lúc mở session và giữ toàn bộ trạng thái DSP nội bộ, gồm cả bộ nhớ filter và phase LFO, liên tục xuyên qua lần đổi, nên resize không bao giờ sinh click gián đoạn. Chỉ phần sổ sách theo block thay đổi. Trạng thái synthesis chạy thẳng qua.

Phía capture, callback của Sonarish làm đúng một việc: copy PCM vào ring buffer rồi return. Engine FFT tiêu thụ ring đó trên background queue, không bao giờ chạy trên audio thread, nhờ vậy app tính transform 4096 điểm trong khi meter cập nhật 30 Hz không glitch. Chi tiết transform nằm ở bài FFT. Câu kiến trúc đáng thuộc lòng rất ngắn: audio thread copy sample rồi return, mọi thứ khác diễn ra chỗ khác.

Debug những cú click vẫn lọt

Click vẫn lọt. Khi nó xuất hiện, mình profile thẳng audio thread: Instruments Time Profiler đánh dấu audio thread trên iOS, Systrace với marker Trace.beginSection trên Android. Kinh nghiệm của team là 99% click production truy về code chỉ-debug rò vào hot path — một String format để log, một guard assertion có cấp phát, một cú unwrap optional Swift vô tình rơi vào slow path. Toán DSP gần như không bao giờ là thủ phạm. Giàn giáo quanh nó thì thường là.

Nên mọi instrumentation debug đều nấp sau compile-time flag được kiểm chứng là tắt ở build release. Cổng cuối là soak test Estua qua đêm của ktuyen: 8 giờ phát liên tục ở chế độ máy bay trên phần cứng thật trước khi atuan tag bất kỳ release nào. Click nào sống sót qua review, qua check tĩnh, qua trọn một đêm phát nhạc thì xứng đáng có bug report mang tên nó.