Sound that never repeats

Âm thanh không bao giờ lặp

Loop-based sleep apps fail for a perceptual reason, and better microphones cannot fix it. An 8-minute rain recording, however professionally captured, carries repeating micro-structure: the same cluster of droplet impacts, the same phase relationship between low rumble and high hiss, coming back every 480 s. By the third minute your auditory cortex has reclassified the signal from "environment" to "recording," and sleep grows lighter at exactly the moment you needed darkness and continuity. Estua goes the other way. It synthesizes ambient audio in real time at 48 kHz on the phone, so the output has no fixed period and no seam to detect. This post walks the whole chain, from the sample clock to the limiter.

The clock and the contract

Everything hangs off a master clock: 48 kHz, fixed blocks of 512 samples. Divide one by the other and you get the deadline, 512 / 48000 = 10.67 ms per callback. Miss it once and the OS ships whatever stale data sits in the output buffer, which a listener hears as a click. Miss it at 3 AM and the product has failed at its one job. So the audio thread runs under a contract we treat like physics: no heap allocation, no locks, no logging, no network calls. Sonarish's capture path obeys the same rules, and the full discipline (pre-allocated buffers, lock-free parameter snapshots) is written up in life on the audio thread.

Estua adds one property that makes the contract simpler and stranger at the same time. It is a pure generator. No file to decode, no stream to buffer, no disk read racing the deadline. Every one of those 512 samples is computed from scratch, every 10.67 ms, all night. The signal path has to stay cheap enough to leave headroom for Bluetooth encoding and whatever the OS decides to do behind our back, and deterministic enough that block N sounds like it belongs after block N−1 with no stored history beyond a few filter states.

From white to pink

The chain starts with white noise. One PRNG draw per sample, flat spectrum, equal energy per hertz. Cheap and wrong. Each octave spans twice the bandwidth of the octave below it, so a flat-per-hertz spectrum doubles its energy per octave as frequency climbs, and the ear reads that as harsh hiss. What sleep wants is a 1/f power spectrum, S(f) ∝ 1/f. Integrate that across any octave and you get ∫ df/f = ln 2, whether the octave starts at 100 Hz or at 4 kHz. Constant energy per octave, a slope of −3 dB per octave. That is pink noise, and it is roughly what rain on a roof and distant surf look like on a spectrum analyzer.

An exact 1/f filter would need infinite state, because every ordinary filter pole contributes −6 dB per octave and we need precisely half of that. Paul Kellet's economical recursive filter solves it by staggering three one-pole low-pass sections at different time constants and summing them, so the staircase of their responses leans along the 1/f line:

// per sample; white is a uniform PRNG draw in [-1, 1]
b0 = 0.99765 * b0 + white * 0.0990460
b1 = 0.96300 * b1 + white * 0.2965164
b2 = 0.57000 * b2 + white * 1.0526913
pink = (b0 + b1 + b2 + white * 0.1848) * outGain

Eight multiplies and six adds per sample, three floats of state, no transcendental functions. On our bench the staircase wobbles around the ideal slope by about ±0.5 dB from 20 Hz to 16 kHz, measured by averaging 300 FFT frames of the raw generator output. Nobody can hear that ripple against the noise itself. The iOS build sometimes swaps in an equivalent biquad stack for tighter shaping; the CPU difference is lost in measurement noise. Underneath sits an optional brown layer for ocean scenes: 1/f², a −6 dB per octave slope, generated by leaky integration of white noise, all rumble below roughly 200 Hz.

Slow oscillators against periodicity

Static pink noise beats a loop, but it is still static. Real environments breathe. The fix is modulation slow enough that it never registers as rhythm: a bank of low-frequency oscillators running between 0.02 and 0.15 Hz, periods of 6.7 to 50 s, each phase drawn uncorrelated from the session seed. One LFO leans on a filter cutoff. Another rides a layer's gain. A third nudges stereo width. The rates are chosen deliberately incommensurate, no period an integer multiple of another, so the joint state of the bank realigns only after the least common multiple of every period involved, which for our chosen ratios exceeds any plausible night of listening. Nothing repeats at 30 or 60 s intervals, exactly the timescales where loop-based apps betray themselves.

Amplitude envelopes come from the same toolbox. A one-pole low-pass on the absolute value of a signal tracks its envelope at the full 48 kHz rate using precomputed coefficients, so a slow swell in the ocean layer costs two multiplies per sample and zero branches. Everything stays on the audio thread, every operation is O(1) per sample, and nothing allocates.

Scenes, seeds, and the statistics of surf

Each scene (Rain, Ocean, Wind, a few others) layers its own processing on the pink foundation. Rain adds band-passed noise bursts whose inter-arrival times follow a Poisson process: draw u uniform in (0, 1], schedule the next droplet at t = −ln(u)/λ, and let the rate λ itself wander under LFO control so the shower thickens and thins. Real droplets on a roof arrive with exactly this statistic; they do not queue politely. Ocean mixes slow amplitude modulation on a low band between 80 and 400 Hz with mid-frequency hiss, which produces the breathing quality of surf without any recorded waveform to repeat. Wind is pink noise pushed through a resonant band-pass whose center frequency drifts, close to what a doorframe does to a gust. At the end of the chain a soft limiter clips at −1 dBFS, preventing digital full-scale spikes that would startle a sleeping listener at 3 AM.

Each night you choose a scene and a seed, or let the app pick randomly. The seed initializes the pseudorandom generators behind the noise samples and the LFO phase offsets, so the same seed reproduces the same soundscape for that session. If you liked last night's rain, you can have it back. We also optionally stir the seed with the calendar date, so that "Rain plus seed 42" is not bit-identical every night for years. Reproducibility within a session, variety across the weeks.

Why crossfading two loops is not enough

The standard industry workaround crossfades between two copies of a loop to hide the seam. It works, in the narrow sense that the audible click at the splice point disappears. The statistics do not move. Autocorrelation of looped rain still spikes at the loop period T and at every integer multiple of it, because the sample you hear at time t is literally the sample from time t − T. Human hearing picks up that kind of periodicity at remarkably subtle levels; evolution spent a long time rewarding brains that noticed patterned sounds in rustling grass. Real ocean surf approximates pink noise with occasional transient swash events whose timing follows no fixed schedule, and no crossfade manufactures that from an 8-minute source. Autocorrelation is the time-domain cousin of the power spectrum; the gentle introduction lives in FFT made readable.

We checked our own homework. Bench note: 30-minute captures of one commercial rain loop and of Estua's Rain scene, normalized autocorrelation computed offline in Python on the recorded output. The loop shows a correlation peak of 0.6 at its period. Estua stays below 0.05 at every lag beyond the filter memory of a few hundred milliseconds. The full synthesis stack costs 3–8% CPU on modern phones, and what that buys is hours of continuous playback with nothing for the auditory cortex to latch onto.

Platform plumbing

On iOS, Estua drives an AVAudioEngine manual source node callback under the playback session category, which keeps audio alive while the screen is locked. On Android it feeds an AudioTrack in PCM float encoding, requesting low-latency mode where the OEM supports it — latency matters less for a sleep app than for an instrument, but the low-latency path tends to be the best-tested one. Both platforms enforce the no-allocation rule in the hot path; parameters cross over from the UI thread as atomically swapped snapshot structs, never through locks. Bench note: across 6 phones (iPhone 12 through 15, a Pixel 6a, a Galaxy A54), an 8-hour overnight synthesis run with the screen off drained 11–19% of battery, logged by recording the battery level at the start and end of each run. The short version of all this lives on the Estua app page.

What Estua refuses to be

Estua contains no binaural beats, no AI-generated music, and no streaming wrapper around somebody else's server. There are no accounts and no analytics on listening habits; how you sleep is nobody's dataset, ours included. A timer fades the volume over 20 to 45 minutes on a curve slow enough that the fade never becomes an audible event of its own. A separate alarm uses a gentle synthesized tone drawn from a different scene, so the sound that wakes you is not the sound you slept to — waking a brain with the exact texture it spent all night filtering out works poorly. For the psychoacoustic background on 1/f noise and sleep, see sleep, waves, and why your brain likes surf. For why the whole thing runs locally, read why everything runs on your phone.

App ngủ chạy loop thua ngay ở tai người nghe, mic xịn đến đâu cũng không cứu được. Track mưa 8 phút, thu chuyên nghiệp cỡ nào, vẫn mang micro-structure lặp: cùng cụm giọt rơi, cùng quan hệ phase giữa rumble trầm với hiss cao, cứ 480 giây quay lại một lần. Nghe đến phút thứ ba, vỏ não thính giác đổi nhãn tín hiệu từ "môi trường" sang "bản thu". Giấc ngủ nông đi đúng lúc cần sâu và liền mạch nhất. Estua đi hướng ngược lại: synthesize ambient real-time ở 48 kHz ngay trên điện thoại, output không có chu kỳ cố định, không có seam nào cho tai bám. Bài này đi trọn chuỗi synthesis, từ sample clock đến limiter.

Master clock và bản hợp đồng

Mọi thứ treo trên một master clock: 48 kHz, block cố định 512 sample. Chia hai số cho nhau là ra deadline, 512 / 48000 = 10,67 ms mỗi callback. Trễ một nhịp, OS phát luôn dữ liệu cũ nằm sẵn trong buffer, tai nghe thành tiếng click. Trễ lúc 3 giờ sáng thì app coi như hỏng đúng nhiệm vụ duy nhất của nó. Vậy nên audio thread chạy theo bản hợp đồng mình coi như định luật vật lý: không cấp phát heap, không lock, không log, không gọi network. Capture path của Sonarish theo cùng luật. Toàn bộ kỷ luật này, từ buffer cấp phát trước đến snapshot tham số lock-free, nằm trong sống trên audio thread.

Estua có một điểm làm bản hợp đồng vừa gọn hơn vừa lạ hơn: nó là generator thuần. Không file để decode, không stream để buffer, không cú đọc disk nào đua với deadline. Cả 512 sample được tính từ con số không, mỗi 10,67 ms, suốt đêm. Chuỗi tín hiệu phải đủ rẻ để chừa headroom cho encode Bluetooth cùng mấy việc OS tự làm sau lưng, lại phải đủ deterministic để block N nghe liền mạch với block N−1 mà chỉ dựa vào vài biến state của filter.

Từ white sang pink

Chuỗi bắt đầu bằng white noise. Mỗi sample rút PRNG một lần, phổ phẳng, năng lượng đều theo từng hertz. Rẻ, nhưng sai chỗ dùng. Octave nào cũng rộng gấp đôi octave ngay dưới, nên phổ phẳng theo hertz nghĩa là năng lượng mỗi octave nhân đôi khi tần số leo lên, tai nghe thành hiss chói. Thứ giấc ngủ cần là phổ 1/f, S(f) ∝ 1/f: tích phân qua octave nào cũng ra ln 2, bất kể octave bắt đầu ở 100 Hz hay 4 kHz. Năng lượng mỗi octave bằng nhau, dốc −3 dB mỗi octave. Đó là pink noise. Mưa trên mái hay sóng biển xa nhìn trên spectrum analyzer cũng ra dạng gần thế.

Filter 1/f chính xác đòi state vô hạn, vì mỗi pole thường góp dốc −6 dB mỗi octave còn mình cần đúng một nửa. Mẹo của Paul Kellet: xếp so le ba khâu one-pole low-pass với hằng số thời gian khác nhau rồi cộng lại, bậc thang đáp ứng của chúng tựa dọc theo đường 1/f:

// mỗi sample; white là một lần rút PRNG uniform trong [-1, 1]
b0 = 0.99765 * b0 + white * 0.0990460
b1 = 0.96300 * b1 + white * 0.2965164
b2 = 0.57000 * b2 + white * 1.0526913
pink = (b0 + b1 + b2 + white * 0.1848) * outGain

Tám phép nhân với sáu phép cộng mỗi sample, ba float state, không hàm transcendental nào. Trên bench của mình, bậc thang lệch quanh dốc lý tưởng khoảng ±0,5 dB trong dải 20 Hz–16 kHz, đo bằng cách trung bình 300 frame FFT trên output thô của generator. Ripple cỡ đó không tai nào tách nổi khỏi chính lớp noise. Bản iOS thỉnh thoảng thay bằng stack biquad tương đương để shape chặt hơn, chênh lệch CPU chìm trong nhiễu đo. Dưới cùng là layer brown tùy chọn cho scene biển: 1/f², dốc −6 dB mỗi octave, tạo bằng tích phân rò white noise, toàn rumble dưới chừng 200 Hz.

Dao động chậm chống lại chu kỳ

Pink noise tĩnh đã hơn loop nhưng vẫn là tĩnh. Môi trường thật biết thở. Cách sửa: modulation chậm tới mức không bao giờ thành nhịp. Một dàn LFO chạy 0,02–0,15 Hz, chu kỳ 6,7–50 giây, phase từng con rút ngẫu nhiên không tương quan từ seed của session. Một LFO đè lên cutoff filter. Con khác lái gain của một layer. Con thứ ba đẩy nhẹ độ rộng stereo. Tần số cố tình chọn không chia chẵn cho nhau, không chu kỳ nào là bội nguyên của chu kỳ khác, nên trạng thái chung của cả dàn chỉ thẳng hàng lại sau bội chung nhỏ nhất của mọi chu kỳ — với tỷ lệ đã chọn thì dài hơn bất kỳ đêm nghe nào. Không gì lặp ở mốc 30 hay 60 giây, đúng vùng thời gian app loop hay lộ.

Envelope biên độ lấy từ cùng hộp đồ nghề. One-pole low-pass đặt lên giá trị tuyệt đối của tín hiệu sẽ bám envelope ở đủ 48 kHz với hệ số tính sẵn, nên một cú swell chậm của layer biển tốn hai phép nhân mỗi sample, không branch nào. Tất cả nằm trên audio thread, mỗi sample O(1), không gì allocate.

Scene, seed và thống kê sóng biển

Mỗi scene (Mưa, Biển, Gió, vài scene nữa) chồng phần xử lý riêng lên nền pink. Mưa thêm burst noise đã band-pass, khoảng thời gian giữa hai burst theo Poisson: rút u uniform trong (0, 1], hẹn giọt kế tiếp ở t = −ln(u)/λ, còn rate λ tự trôi theo LFO cho cơn mưa lúc dày lúc thưa. Giọt mưa thật trên mái rơi đúng kiểu thống kê này, chẳng giọt nào chịu xếp hàng. Biển trộn modulation biên độ chậm trên dải trầm 80–400 Hz với hiss tần trung, ra đúng nhịp thở của sóng mà không cần waveform thu sẵn nào để lặp. Gió là pink noise đẩy qua band-pass cộng hưởng có tần số trung tâm trôi dần, khá gần thứ khung cửa làm với một cơn gió giật. Cuối chuỗi, soft limiter clip ở −1 dBFS, chặn spike full-scale digital dựng người đang ngủ dậy lúc 3 giờ sáng.

Mỗi đêm bạn chọn scene với seed, hoặc để app tự random. Seed khởi tạo các bộ sinh giả ngẫu nhiên đứng sau sample noise cùng offset phase của LFO, nên cùng seed sẽ dựng lại đúng soundscape của session đó. Thích cơn mưa đêm qua thì gọi lại được. App còn tùy chọn trộn seed với ngày trên lịch, để "Mưa cộng seed 42" không giống hệt nhau đêm này qua năm nọ. Tái lập được trong một session, đa dạng qua nhiều tuần.

Vì sao crossfade hai loop vẫn không đủ

Chiêu quen trong ngành là crossfade giữa hai bản copy của loop để giấu mối nối. Có tác dụng thật, theo nghĩa hẹp: tiếng click ở điểm ghép biến mất. Thống kê thì đứng nguyên. Autocorrelation của mưa loop vẫn spike ở chu kỳ T cùng mọi bội nguyên của nó, vì sample bạn nghe ở thời điểm t chính là sample của thời điểm t − T. Tai người bắt periodicity ở mức tinh đến khó tin; tiến hóa thưởng rất hậu cho bộ não nào nhận ra âm có pattern trong đám cỏ xào xạc. Sóng biển thật xấp xỉ pink noise cộng những cú swash thoáng qua chẳng theo lịch nào; không crossfade nào nặn được thứ đó từ nguồn 8 phút. Autocorrelation là họ hàng miền thời gian của phổ công suất, bản dẫn nhập nhẹ nhàng nằm ở bài FFT dễ đọc.

Mình tự chấm bài mình. Ghi chú bench: thu 30 phút một rain loop thương mại và 30 phút scene Mưa của Estua, autocorrelation chuẩn hóa tính offline bằng Python trên bản thu. Loop cho peak tương quan 0,6 tại chu kỳ của nó. Estua nằm dưới 0,05 ở mọi lag ngoài trí nhớ filter cỡ vài trăm mili giây. Cả stack synthesis tốn 3–8% CPU trên máy đời mới; đổi lại là hàng giờ phát liên tục mà vỏ não thính giác không có gì để bám.

Đường ống nền tảng

Trên iOS, Estua lái callback source node thủ công của AVAudioEngine dưới session category playback, giữ audio sống khi màn hình khóa. Trên Android là AudioTrack với PCM float, xin low-latency mode ở máy nào OEM chịu hỗ trợ — độ trễ với app ngủ không quan trọng bằng với nhạc cụ, nhưng đường low-latency thường được test kỹ nhất. Hai nền tảng cùng ép luật không allocate trên hot path; tham số từ UI thread đi sang dạng snapshot struct swap nguyên tử, không bao giờ qua lock. Ghi chú bench: 6 máy (iPhone 12 đến 15, một Pixel 6a, một Galaxy A54), chạy synthesis 8 tiếng qua đêm với màn hình tắt, tụt 11–19% pin, ghi bằng mức pin lúc bắt đầu và lúc kết thúc mỗi lần chạy. Bản rút gọn của toàn bộ chuyện này nằm ở trang app Estua.

Những thứ Estua từ chối làm

Estua không có binaural beats, không nhạc AI sinh, không bọc stream quanh server của ai khác. Không tài khoản, không analytics thói quen nghe; bạn ngủ ra sao không phải dataset của bất kỳ ai, kể cả của mình. Timer fade volume trong 20–45 phút theo đường cong chậm tới mức bản thân cú fade không bao giờ thành một sự kiện nghe được. Báo thức tách riêng, dùng tone synthesize nhẹ lấy từ scene khác: âm đánh thức bạn không trùng âm bạn đã ngủ cùng, vì dựng một bộ não dậy bằng đúng texture nó đã lọc bỏ cả đêm là ý tồi. Nền psychoacoustic của noise 1/f với giấc ngủ nằm ở ngủ, sóng và vì sao não thích biển; lý do mọi thứ chạy local đọc ở vì sao mọi thứ chạy trên điện thoại.