Skip to content跳到正文
All works全部作品
Code代码

Wonjae Studio Web DAWWonjae Studio 浏览器 DAW

A browser digital audio workstation — recording, a real-time plugin rack, multitrack arranging and lossless WAV mastering, all client-side.一个跑在浏览器里的数字音频工作站——录音、实时插件机架、多轨编排和无损 WAV 母带,全部在客户端完成。

Stack技术栈
TypeScript · React · Web Audio API · Canvas · IndexedDB

A full recording and mixing environment that ran entirely in the browser — no server, no upload step, no native plugin host. It shipped as part of this site for a while, and was retired when the site became a portfolio rather than a tool. The code lives on in the repository's history.

What it did

Recording with pre-roll. A 2/4/6-second count-in played the backing track at 40% gain from startAt - preroll, then faded up at the downbeat. The stopwatch counted negative through the pre-roll so you could feel the bar before singing.

Zero-touch alignment. The hard part of recording vocals against a backing track in a browser is latency: monitoring delay, buffer size and device offset all shift the take by tens of milliseconds, and the amount varies per session. So instead of asking the user to nudge a waveform, the take was aligned automatically on stop:

// Cross-correlate RMS envelopes of the dry vocal and the backing track over the
// first 20 s, scanning a physically plausible latency window.
const WINDOW_MS = { min: -40, max: 240 };

function bestOffset(vocal: Float32Array, backing: Float32Array) {
  let best = { offsetMs: 0, score: -Infinity };
  for (let ms = WINDOW_MS.min; ms <= WINDOW_MS.max; ms += 1) {
    const score = dot(vocal, shift(backing, ms));
    if (score > best.score) best = { offsetMs: ms, score };
  }
  return best.offsetMs;
}

Correlating energy envelopes rather than raw samples is what makes this work: a sung consonant and a snare hit don't share a waveform, but they share a transient. The peak of the correlation is the point where the singer's accents line up with the beat.

A plugin rack. Parametric EQ with draggable curve nodes, compressor, gate, stereo imager, saturation, limiter, chorus, delay, and a convolution reverb. Plus an autotune built on autocorrelation pitch detection and time-domain overlap-add pitch shifting, with selectable root and scale.

Lossless export. Mixdown didn't re-record in real time. On export it spun up an OfflineAudioContext, rendered the aligned multitrack in memory in a few milliseconds, then packed the result as 16-bit 48 kHz stereo PCM WAV.

A 3D spectrogram. Real-time FFT bins projected onto an isometric grid as a scrolling terrain, drawn on a plain 2D canvas with a hand-rolled perspective transform.

Try the part that survived

The recorder is gone, but the synthesis and the spectrogram are not. Program a pattern and it runs on the same kind of graph the DAW used — oscillators and filtered noise scheduled against the audio clock, feeding a live FFT.

Beat
1...2...3...4...
Kick
Snare
Clap
Hat

16 steps = one bar of sixteenth notes. Synthesised, not sampled — the kick is a pitch-swept sine, snare and hat are filtered noise. The terrain is a real FFT of what you programmed.

The tracks elsewhere on this site cannot drive this display: their CDN sends no Access-Control-Allow-Origin header, so an AnalyserNode fed from those elements receives silence rather than samples. Locally synthesised audio has no such problem, which is the only reason the terrain here is real and not a decorative animation.

What I'd keep

The auto-alignment. Everything else in the rack has a better native equivalent, but "press stop and the take is already in time" is genuinely nicer than any desktop DAW I've used, and it only works because the browser knows both signals.

Why it's archived

It was a tool wearing a portfolio's clothes. Roughly two thirds of this repository was DAW code, which made the site slow to load and slower to change, and the audio it produced was stored in IndexedDB — visible only in the browser that made it. A portfolio should show finished work to other people; a DAW should be its own product. Splitting them was the right call.

一套完整的录音和混音环境,完全跑在浏览器里——没有服务器,没有上传步骤,没有原生插件宿主。 它曾经作为这个站的一部分上线过一段时间,在这个站从工具变成作品集的时候退役了。 代码还留在仓库的历史里。

它能做什么

带预卷的录音。 2/4/6 秒的预备拍会从 startAt - preroll 开始以 40% 增益播放伴奏, 然后在正拍上淡入。秒表在预卷期间倒着走负数,这样你在开口之前能先感觉到这一小节。

零操作对齐。 在浏览器里对着伴奏录人声,难的地方是延迟: 监听延迟、缓冲区大小和设备偏移会把这一条挪走几十毫秒,而且每次录的量都不一样。 所以与其让用户去手动拖波形,不如在停止录音时自动对齐:

// Cross-correlate RMS envelopes of the dry vocal and the backing track over the
// first 20 s, scanning a physically plausible latency window.
const WINDOW_MS = { min: -40, max: 240 };

function bestOffset(vocal: Float32Array, backing: Float32Array) {
  let best = { offsetMs: 0, score: -Infinity };
  for (let ms = WINDOW_MS.min; ms <= WINDOW_MS.max; ms += 1) {
    const score = dot(vocal, shift(backing, ms));
    if (score > best.score) best = { offsetMs: ms, score };
  }
  return best.offsetMs;
}

对能量包络做相关而不是对原始采样做相关,是这件事能成的关键: 唱出来的辅音和一记军鼓不共享波形,但它们共享一个瞬态。 相关的峰值就是歌手的重音和节拍对上的那个点。

一个插件机架。 带可拖拽曲线节点的参量均衡、压缩器、门限、立体声声像、 饱和、限幅、合唱、延迟,以及一个卷积混响。 外加一个自动修音,基于自相关的音高检测和时域 overlap-add 变调,根音和音阶可选。

无损导出。 合并不是实时重录一遍。导出时会起一个 OfflineAudioContext, 在内存里几毫秒渲染完对齐后的多轨,然后打包成 16 位 48 kHz 立体声 PCM WAV。

一个 3D 频谱图。 实时 FFT 频段投影到等轴网格上,作为滚动的地形, 画在一个普通的 2D canvas 上,透视变换是手写的。

试试活下来的那部分

录音机没了,但合成和频谱图还在。编一个 pattern,它会跑在和当年 DAW 同一类的图上—— 振荡器和滤波噪声按音频时钟调度,喂给一个实时 FFT。

Beat
1...2...3...4...
Kick
Snare
Clap
Hat

16 steps = one bar of sixteenth notes. Synthesised, not sampled — the kick is a pitch-swept sine, snare and hat are filtered noise. The terrain is a real FFT of what you programmed.

这个站上别处的曲子驱动不了这个显示:它们的 CDN 不发 Access-Control-Allow-Origin 头, 所以从那些元素接出来的 AnalyserNode 收到的是静音而不是采样。 本地合成的音频没这个问题,这也是这里的地形是真实数据而不是装饰动画的唯一原因。

我会留下什么

自动对齐。机架里其他东西都有更好的原生替代品, 但「按下停止,这一条已经在拍上了」确实比我用过的任何桌面 DAW 都舒服, 而它能成立只是因为浏览器同时拿得到两路信号。

为什么归档了

它是一个穿着作品集衣服的工具。这个仓库大概三分之二是 DAW 代码, 这让站点加载慢、改起来更慢,而且它产出的音频存在 IndexedDB 里—— 只有做出它的那个浏览器看得见。作品集应该把做完的东西展示给别人; DAW 应该是它自己的产品。把它们拆开是对的。

Web Audio APIDSPReactAudio ProgrammingWeb Audio APIDSPReact音频编程