Watermark Spectrogram Comparison
Visual comparison of where watermark energy appears across the 10 methods.
A Comprehensive Measurement Study — benchmarking 10 open-source audio watermarking methods against digital, physical, and AI-induced removal attacks.
We study whether modern audio watermarks survive practical removal attacks while preserving the usefulness of the audio, connecting a component-wise taxonomy of 26 schemes with a benchmark of 10 reproducible methods across speech and music.
We reproduce 10 open-source watermarking methods spanning both AI-based and traditional signal-processing approaches, each with different embedding strategies and robustness characteristics.
| # | Method | Type | Watermark Pattern | Frequency Range | Message Repetition | Training Distortions |
|---|---|---|---|---|---|---|
| 1 | AudioSeal | AI | Encoder network → waveform additive | 0 – 8000 Hz | Frame level | 14 |
| 2 | WavMark | AI | Invertible Neural Network (INN) | 0 – 8000 Hz | Block level | 10 |
| 3 | SilentCipher | AI | Encoder network → waveform additive | 0 – 4000 Hz | Frame level | 6 |
| 4 | Timbre | AI | Encoder network → waveform additive | 0 – 8000 Hz | Frame level | 1 |
| 5 | RobustDNN | AI | Encoder network → waveform additive | 0 – 8000 Hz | Frame level | 3 |
| 6 | AWARE | AI | Adversarial optimization | 0 – 8000 Hz | None | 0 |
| 7 | audiowmark | Traditional | FFT bin modification | 861 – 4307 Hz | Block level | N/A |
| 8 | FSVC | Traditional | DCT coefficient modification | 1088 – 1448 Hz | None | N/A |
| 9 | Patchwork | Traditional | DCT coefficient modification (multi-layer) | 3000 – 7000 Hz | Layer level | N/A |
| 10 | Norm-space | Traditional | DCT coefficient modification (norm-space) | 0 – 8000 Hz | None | N/A |
We evaluate watermark robustness across three categories of removal attacks: digital-level, physical-level, and AI-induced distortions.
Visual comparison of where watermark energy appears across the 10 methods.
Heatmap of watermark bit recovery accuracy and perceptual quality under pitch-shift and time-stretch attacks. Quality thresholds: ViSQOL > 3.0, MUSHRA-Music > 79, MUSHRA-Speech > 60.
Re-recording introduces room, distance, microphone, and speaker effects. Distance causes larger accuracy drops than device changes.
Distance accuracy
Device accuracy
Distance setting
Device setting
Voice conversion and TTS regeneration preserve perceived speech utility while discarding watermark traces from the waveform.
Zero-shot VC
Few-shot / Adaptive VC
Zero-shot TTS
TTS + G-Lim
TTS + H-GAN
TTS + H-GAN*
| Metric | No Attack |
Voice Conversion | Text-to-Speech | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AdaIn-VC | FragmentVC | YourTTS VC | MediumVC | RVC | YourTTS | Tacotron2 | FastSpeech2 | ||||||
| G-Lim | H-GAN | H-GAN | G-Lim | H-GAN | H-GAN | ||||||||
| Zero-shot | Zero-shot | Zero-shot | Few-shot | Adaptive | Zero-shot | Adaptive | Adaptive | Adaptive* | Adaptive | Adaptive | Adaptive* | ||
| ViSQOL | 4.87 | 3.82 | 3.02 | 3.96 | 2.54 | 4.04 | 2.35 | 1.81 | 2.69 | 2.72 | 1.82 | 2.71 | 2.75 |
| SECS | 0.99 | 0.84 | 0.75 | 0.90 | 0.67 | 0.93 | 0.83 | 0.69 | 0.87 | 0.87 | 0.71 | 0.89 | 0.90 |
| MUSHRA-Speech | 100 | 70 | 64 | 65 | 76 | 97 | 78 | 81 | 85 | 78 | 79 | 96 | 72 |
Listen to representative audio samples for each distortion setting. Samples are grouped by distortion type. Expand a group to browse individual settings, then expand a setting to listen to more samples.