Voice Morphing for Speaker De-identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems struggle to mask speaker identity while maintaining intelligibility, as existing voice transformation methods that sufficiently de-identify speakers often reduce transcription accuracy and intelligibility to unacceptable levels, especially in noisy and reverberant environments.
Innovation Solution
The proposed solution involves a voice morphing process that applies pitch and frequency shifts to audio clips, allowing for effective de-identification of speakers while minimizing loss of intelligibility, using a method that includes randomizing pitch and frequency shifts to create variance and secure the morphing process either on the client browser or server-side, ensuring that only morphed clips are used for labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional voice transforms are applied at sufficient strength to mask speaker identity, then speaker de-identification is improved, but speech intelligibility deteriorates to unacceptable levels
Solution Approach 1:
The patent applies multiple voice transformation parameters including pitch shift (e.g., +12 semitones), frequency shift (e.g., +500 Hz), and formant modification to achieve effective speaker de-identification. By adjusting these parameters within specific ranges, the system masks speaker identity while preserving speech intelligibility better than conventional single-parameter transforms.
Solution Approach 2:
The system combines multiple voice transformation techniques (pitch shifting, frequency shifting, formant modification) into a composite transformation approach. This multi-layered transformation strategy achieves robust speaker masking while maintaining acceptable intelligibility levels, overcoming the limitations of single-method conventional transforms.
2Reliability
If voice transformation strength is increased to mask speaker identity, then speaker de-identification is improved, but transcription accuracy deteriorates
Solution Approach 1:
The system modifies multiple acoustic parameters (pitch, frequency, formants) simultaneously with controlled magnitudes. This multi-parameter approach achieves strong speaker masking while preserving enough speech structure for accurate transcription, unlike single-parameter transforms that sacrifice too much intelligibility.
Solution Approach 2:
The patent applies voice transformation at a strength that is partially sufficient for masking (not maximum possible transformation), thereby achieving adequate de-identification while preserving transcription accuracy. The transformation is calibrated to provide just enough masking without over-transforming and losing intelligibility.
3Object-affected harmful factors
If existing voice transformation methods are used, then speaker masking is achieved, but speech becomes more difficult to understand and transcribe
Solution Approach 1:
The system carefully selects and adjusts transformation parameters (pitch shift range, frequency shift magnitude, formant modification depth) to achieve speaker masking while minimizing impact on speech understanding. This controlled parameter modification preserves transcription ease better than conventional transforms.
Solution Approach 2:
The voice transformation acts as an intermediary that modifies speaker-specific characteristics while preserving speech content. The transformation serves as a mediator between complete privacy protection and full speech intelligibility, achieving a balanced state that maintains both masking effectiveness and transcription ease.
Data Source
AI summary
A system and method for masking an identity of a speaker of natural language speech, such as speech clips to be labeled by humans in a system generating voice transcriptions for training an automatic speech recognition model. The natural language speech is morphed prior to being presented to the human for labeling. In one embodiment, morphing comprises pitch shifting the speech randomly either up or down, then frequency shifting the speech, then pitch shifting the speech in a direction opposite the first pitch shift. Labeling the morphed speech comprises at least one or more of transcribing the morphed speech, identifying a gender of the speaker, identifying an accent of the speaker, and identifying a noise type of the morphed speech.


