OSORIO.SYS

Audio Fingerprinting Basics in Rust

Audio fingerprinting lets us answer questions like:

“Does this noisy recording match anything in my catalog?”

without storing or comparing raw waveforms.

In this article we’ll:

  1. Load WAV files with symphonia.
  2. Compute magnitude spectra using rustfft.
  3. Build a very naive hash just to get the pipeline in place.

Loading audio with Symphonia

We start with 16-bit mono WAV files to keep things simple:

rust
use symphonia::core::{
    audio::{AudioBufferRef, Signal},
    codecs::CODEC_TYPE_NULL,
    formats::FormatOptions,
    io::MediaSourceStream,
    meta::MetadataOptions,
    probe::Hint,
};
use std::{fs::File, path::Path};

pub fn load_samples(path: &Path) -> anyhow::Result<Vec<f32>> {
    let file = File::open(path)?;
    let mss  = MediaSourceStream::new(Box::new(file), Default::default());

    let mut hint = Hint::new();
    hint.with_extension("wav");

    let probed = symphonia::default::get_probe().format(
        &hint,
        mss,
        &FormatOptions::default(),
        &MetadataOptions::default(),
    )?;

    let mut format = probed.format;
    let track = format
        .tracks()
        .iter()
        .find(|t| t.codec_params.codec != CODEC_TYPE_NULL)
        .ok_or_else(|| anyhow::anyhow!("no audio track"))?;

    let mut decoder = symphonia::default::get_codecs().make(&track.codec_params, &Default::default())?;

    let mut samples = Vec::new();

    loop {
        let packet = match format.next_packet() {
            Ok(packet) => packet,
            Err(_) => break,
        };

        let decoded = decoder.decode(&packet)?;

        if let AudioBufferRef::S16(buf) = decoded {
            for frame in buf.chan(0) {
                samples.push(*frame as f32 / i16::MAX as f32);
            }
        }
    }

    Ok(samples)
}

From samples to spectra

We’ll use a fixed-size FFT (e.g. 4096 points). Each window produces one spectrum:

rust
use rustfft::{FftPlanner, num_complex::Complex32};

pub fn spectrogram(samples: &[f32], fft_size: usize, hop: usize) -> Vec<Vec<f32>> {
    let mut planner = FftPlanner::new();
    let fft = planner.plan_fft_forward(fft_size);

    let mut out = Vec::new();
    let mut buf = vec![Complex32::ZERO; fft_size];

    let mut i = 0;
    while i + fft_size <= samples.len() {
        for (dst, &s) in buf.iter_mut().zip(&samples[i..i + fft_size]) {
            *dst = Complex32::new(s, 0.0);
        }

        fft.process(&mut buf);

        let magnitudes = buf
            .iter()
            .map(|c| c.norm())
            .collect::<Vec<_>>();

        out.push(magnitudes);
        i += hop;
    }

    out
}

A toy fingerprint

Real systems use clever peak picking and hashing. For now we’ll do something intentionally simple:

  • Divide each spectrum into N bands.
  • For each band, record the index of the maximum bin.
  • Concatenate those indices into a small vector.
rust
pub fn toy_fingerprint(spec: &[Vec<f32>], bands: usize) -> Vec<u16> {
    let bins = spec[0].len();
    let band_width = bins / bands;

    spec.iter()
        .map(|frame| {
            let mut fp = 0u16;

            for b in 0..bands {
                let start = b * band_width;
                let end   = ((b + 1) * band_width).min(bins);
                let (idx, _) = frame[start..end]
                    .iter()
                    .enumerate()
                    .max_by(|a, b| a.1.partial_cmp(b.1).unwrap())
                    .unwrap();

                // 4 bits per band (toy!), shift and or
                fp <<= 4;
                fp |= (idx as u16 & 0x0F);
            }

            fp
        })
        .collect()
}
This is deliberately bad

Don’t ship this to production 🙂
The goal is to have something concrete we can later replace with a better, well-researched scheme.

Sanity-checking in the terminal

Once everything compiles, we run a quick smoke test:

cargo run --release --bin toy-fp ./data/snare.wav
frames: 214
fingerprints: 214
first 10: [0x3AF2, 0x39F1, 0x39F1, 0x39E1, 0x39E1, 0x39D1, 0x39C1, 0x39C1, 0x39C1, 0x39C1]

If two different takes of the same audio produce a similar sequence of toy fingerprints, we’re on the right track.

In later posts we’ll:

  • Replace the toy hash with a constellation hash.
  • Add a database backend for large catalogs.
  • Benchmark the system on real-world noisy recordings.