Go to personal site →

Research Scenario: What Can Voice Do Beyond Talk?

Voice is not just how we talk. It is how we communicate with machines and each other, how systems verify who we are, and how devices sense the physical world. Every one of these roles is now powered by AI, and every one of them can be attacked.

Communication icon

Communication

Voice carries commands and conversations, from voice assistants in the car and at home, to underwater acoustic links where radio cannot reach.

In-car voice assistant
Voice control in-car
Underwater acoustic communication
Underwater acoustic links
Biometric icon

Biometrics

“My voice is my password.” Banks, call centers, and smartphones authenticate people by how they sound, a biometric you cannot change.

Voice recognition on a smartphone
Smartphone voice ID
Chase voice biometrics banking
Voice banking (Chase)
Sensing icon

Sensing

Sound waves double as radar: inaudible ultrasound tracks gestures, monitors heartbeats, and verifies that a real human is present.

Acoustic gesture sensing with a smartphone
Gesture sensing
Wearable cardiac monitoring
Health monitoring

Threat Models: Who Attacks Whom?

We organize threats by the interaction they poison. Each row below is a real attack surface that our lab has demonstrated or defended in top security venues.

User at a laptop × Mobile device

Human ↔ Mobile Thrust 1 ↓

Attackers fool the speaker authentication and speech recognition models inside phones, earbuds, and voice assistants.

  • Mis-command
  • Mis-access control
  • Inaudible injection
  • Adversarial speech
  • Backdoored voice models
User × Another user

Human ↔ Human Thrust 2 ↓

Generative AI turns anyone’s voice into a weapon: deepfake scam calls, cloned identities, secretly recorded and misused speech.

  • Deepfake speech online
  • Scam calls
  • Unauthorized recording
  • Speech dataset misuse
  • Sensing privacy leakage
User × Malicious AI agent

Human ↔ Agent Thrust 3 ↓

Autonomous LLM agents browse, listen, retrieve, and act. A poisoned page, document, or audio clip can hijack them, and they can impersonate us.

  • Malicious action
  • Malicious output
  • Voice jailbreak
  • Copyright violation
  • Money loss & fraud

Try It Yourself

Our research is demo-driven. These interactive sites let you experience the attacks, defenses, and datasets first-hand.

Agent Security Playground preview
Interactive Playground

Agent Security Playground

Launch real prompt-injection, tool-exploitation, and chained attacks against a sandboxed LLM agent, then watch it get hijacked.

EchoGuard preview
Live System

EchoGuard

Continuous ultrasound sensing that tells real humans apart from AI agents operating your computer, using only the speaker and microphone.

Audio Watermarking SoK preview
Benchmark + Audio Demos

SoK: Audio Watermarking

26 watermark schemes, 127 attack settings, listenable samples. Find out which watermarks actually survive deepfake pipelines.

WiSenseHub preview
Open Dataset

WiSenseHub

A curated wireless-sensing dataset catalog. Explore CSI-Bench and other datasets for sensing and security research.

Thrust 1: Human–Mobile Security

How can we secure interactions between people and mobile, embedded, and edge-intelligent systems?

Voice assistants, speech recognition, speaker verification, and wireless sensing operate in noisy, adversarial acoustic and radio environments. We uncover practical attack surfaces on edge devices and design defenses that remain usable in the real world.

Watch it in action

SurfingAttack (NDSS'21)

Inaudible ultrasonic guided waves travel through the table and command a phone while the owner sits right next to it.

GhostTalk (NDSS'22)

A modified charging cable injects voice commands into a smartphone through the power line, with no sound needed.

Projects

WavePurifier

WavePurifier

Purifying audio adversarial examples via hierarchical diffusion models (MobiCom'24).

MASTERKEY

MASTERKEY

Practical backdoor attack against speaker verification systems (MobiCom'23).

PiezoBud

PiezoBud

A piezo-aided secure earbud with practical speaker authentication (SenSys'24).

SpecPatch

SpecPatch

Human-in-the-loop adversarial spectrogram patch attack on ASR, Best Paper Honorable Mention (CCS'22).

PhantomSound

PhantomSound

Black-box, query-efficient audio adversarial attack via split-second phoneme injection (RAID'23).

SuperVoice

SuperVoice

Speaker verification using ultrasound energy in human speech (AsiaCCS'22).

SurfingAttack

SurfingAttack

Ultrasonic guided-wave attacks on voice assistants (NDSS'21).

GhostTalk

GhostTalk

Interactive attack on smartphone voice systems through the power line (NDSS'22).

Full paper list with authors →

Thrust 2: Human–Human Security

How can we protect human communication, identity, privacy, and digital content from recording, impersonation, sensing, and misuse?

Unauthorized recording, voice cloning, deepfake fraud, and privacy leakage from acoustic, RF, and mmWave sensing threaten interpersonal trust. We build defenses that prevent capture, block impersonation, preserve sensing privacy, and trace unauthorized content use.

Watch it in action

Ultrasound Watermark

Embedding watermarks into live speech with ultrasound, so unauthorized recordings can be traced and verified.

Secure-IRS (ICNC'25)

An intelligent reflecting surface defends against adversarial physical-layer sensing in ISAC systems.

Projects

Learning to Evade (AWM)

Learning to Evade

Adaptive attacks on audio watermarking that bypass distribution-based defenses (Interspeech'26).

AUDIO WATERMARK

AUDIO WATERMARK

Dynamic and harmless watermark for black-box voice dataset copyright protection (USENIX Security'25).

NEC

NEC

Speaker-selective cancellation via neural enhanced ultrasound shadowing (DSN'22). Covered by New Scientist.

VSMask

VSMask

Real-time predictive perturbation against voice synthesis attacks (WiSec'23).

Secure-IRS

Secure-IRS

Defending against adversarial physical-layer sensing in ISAC systems (ICNC'25).

Full paper list with authors →

Thrust 3: Human–Agent Security

How can we secure interactions among people, autonomous AI agents, websites, and knowledge systems?

Agents that browse the web, interpret audio, retrieve documents, generate content, and take actions create a new threat surface. We study jailbreaks, scraper abuse, RAG integrity, generative-system manipulation, and how to prove a real human is behind the keyboard.

Watch it in action

How Mobile Agents Work, and How We Tell Them from Humans

See how a mobile GUI agent perceives the screen, plans, and taps through an app, and how we can detect that the actions come from an agent rather than a human.

Demo credit: incoming Ph.D. student Hengrui Yan, Tsinghua University

Projects

EchoGuard agent detection

EchoGuard

Continuous ultrasound-based human presence verification that tells real users apart from AI agents.

Agent Security Playground

Agent Security Playground

Interactive teaching platform: launch data-ingestion, tool-use, and chained attacks on a sandboxed agent.

WebCloak

WebCloak

Characterizing and mitigating LLM-driven web agents as intelligent scrapers (S&P'26).

Audio Jailbreak Attacks

Audio Jailbreak Attacks

Exposing vulnerabilities in SpeechGPT in a white-box framework (DSN-DSML'25).

Full paper list with authors →

Foundations & Earlier Systems Work

Our security agenda builds on award-winning systems work in wireless communication, LoRa, underwater navigation, and mobile sensing.

NELoRa

NELoRa

Ultra-low SNR LoRa communication with neural-enhanced demodulation, Best Paper Award (SenSys'21).

U-Star

U-Star

An underwater navigation system based on passive 3D optical identification tags (MobiCom'22).

DSIC

DSIC

Deep-learning self-interference cancellation for in-band full duplex wireless, Best Paper Award (GLOBECOM'19).

FedIoT

FedIoT

Federated IoT interaction vulnerability analysis (ICDE'23).

Full paper list with authors →