Speaker Identity Verification via Articulator Motion Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Replay attacks pose a significant threat to speaker identity verification systems, as they can mimic the liveness of a speaker's video and voice, making it difficult to distinguish between a real and recorded speech signal.
Innovation Solution
A system that combines motion sensors to capture physical articulator movements and acoustic sensors to receive speech signals, correlating the data to verify the source of the speech and ensure it is from a living human, using remote sensing techniques to measure internal and external articulator vibrations and oscillations, thereby preventing spoofing and replay attacks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If video and acoustic sensors are used to verify speaker liveness, then speaker identity verification accuracy is improved, but the system becomes susceptible to replay attacks
Solution Approach 1:
The patent introduces an intermediary physical quantity (jaw motion, neck motion, or throat motion) that mediates between the speaker's vocal cord vibration and the acoustic signal. This intermediary motion serves as a trusted indicator that is difficult to replicate in replay attacks, while still being correlated with the acoustic speech signal for verification purposes.
Solution Approach 2:
The patent replaces reliance on acoustic signal analysis alone with mechanical motion sensing of articulators. By using motion sensors to detect physical jaw, neck, or throat movements, the system substitutes acoustic verification with mechanical verification, which is inherently more resistant to replay attacks since recorded mechanical motions cannot be easily replicated.
2Object-affected harmful factors
If motion sensors are added to capture articulator movements, then resistance to replay attacks is improved, but device complexity increases
Solution Approach 1:
The patent makes the motion sensing system multi-functional by enabling it to detect different types of articulator motions (jaw, neck, or throat) depending on implementation. The same basic motion sensing architecture can be adapted to sense various physical movements, reducing overall system complexity through component reuse and standardized processing pipelines.
Solution Approach 2:
The patent changes the measurement parameter from acoustic signal characteristics to mechanical motion characteristics. By shifting from analyzing sound wave properties to analyzing physical displacement, velocity, or acceleration of articulators, the system achieves replay attack resistance while using simpler motion sensing technology rather than complex acoustic analysis.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively confirms the identity of the speaker by ensuring that the speech signal is produced by a functioning human vocal apparatus, providing robust resistance to both live and mechanical impersonation attempts, thus preventing successful spoofing and replay attacks.
Implementation Method 1
capture physical motion of at least one articulator that contributes to the production of speech
Implementation Method 2
measure internal and external articulator vibrations and oscillations
Implementation Method 3
acoustic signal sensor to receive acoustic signals
Data Source
AI summary
A system for confirming that a subject is the source of spoken audio and the identity of the subject providing the spoken audio is described. The system includes at least one motion sensor operable to capture physical motion of at least one articulator that contributes to the production of speech, at least one acoustic signal sensor to receive acoustic signals, and a processing device comprising a memory and communicatively coupled to the at least one motion sensor and the at least one acoustic signal sensor. The processing device is programmed to correlate physical motion data with acoustical signal data to uniquely characterize the subject for purposes of verifying the subject is the source of the acoustical signal data and the identity of the subject.


