Voice Mimic Detection Using Composite Training Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-activated systems fail to account for vocal impersonation, compromising parental control settings when a child mimics a parent's voice to access restricted content.
Innovation Solution
A neural network is trained using composite voice signals from individuals within the same and different households to distinguish between original and mimicked voice inputs, employing diarization, tempo adjustment, and superimposition techniques to create Class A and Class B data for classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice-activated systems use user profile settings to restrict content access, then parental control is improved, but the system becomes vulnerable to vocal impersonation by children
Solution Approach 1:
The system performs preliminary actions by capturing and storing voiceprints of authorized users (parents) in advance, creating a reference database before any content access occurs. This preliminary voiceprint collection enables the system to detect and prevent vocal impersonation attempts when children try to access restricted content, thereby maintaining parental control reliability while addressing the vulnerability to vocal impersonation.
2Measurement precision
If the system trains a neural network to detect mimicked voice inputs, then detection accuracy is improved, but system complexity increases
Solution Approach 1:
The system creates synthetic mimicked voice samples by combining and manipulating recorded voiceprints of different users. These synthetic samples serve as training data that replicate the characteristics of actual vocal impersonation attempts. By using these copied and manipulated voice patterns as training data, the neural network learns to detect mimicked voices without requiring complex real-time analysis during actual use, thus improving detection accuracy while managing system complexity.
Solution Approach 2:
The system transforms voice signals by adjusting parameters such as pitch, tone, and acoustic characteristics to create varied training samples. The neural network is trained on these parameter-transformed voice data, enabling it to recognize subtle variations and inconsistencies in mimicked voices. This parameter transformation approach allows the system to achieve high detection accuracy by learning from diverse voice patterns while keeping the underlying system architecture relatively simple.
3Productivity
If the system combines voice signals from multiple individuals to create composite data, then training effectiveness is improved, but data processing complexity increases
Solution Approach 1:
The system merges voice signals from multiple individuals by superimposing and combining their acoustic waveforms to create composite voice data. This merging process generates training samples that represent various scenarios of vocal impersonation and mixed-voice conditions. By combining multiple voice signals in this controlled manner, the system improves training effectiveness by exposing the neural network to diverse voice patterns while managing data processing complexity through automated signal processing techniques.
Data Source
AI summary
Methods and systems are disclosed herein for training a network to detect mimicked voice input, so that it can be determined whether a voice input signal is a mimicked voice signal. First voice data is received. The first voice data comprises at least a voice signal of a first individual and another voice signal. The voice signal of the first individual and at least one other voice signal is combined to create a composite voice signal. Second voice data is received. The second voice data comprises at least a voice signal of the first individual. The network is trained using at least the composite voice signal and the second voice data to determine whether a voice input signal is a mimicked voice input signal.


