Voice Morphing for Speaker De-identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems struggle to mask speaker identity while maintaining intelligibility, as existing voice transformation methods that sufficiently de-identify speakers often reduce transcription accuracy and intelligibility to unacceptable levels, especially in noisy and reverberant environments.

Innovation Solution

The proposed solution involves a voice morphing process that applies pitch and frequency shifts to audio clips, allowing for effective de-identification of speakers while minimizing loss of intelligibility, using a method that includes randomizing pitch and frequency shifts to create variance and secure the morphing process either on the client browser or server-side, ensuring that only morphed clips are used for labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional voice transforms are applied at sufficient strength to mask speaker identity, then speaker de-identification is improved, but speech intelligibility deteriorates to unacceptable levels

Engineering Contradiction:
Improvespeaker de-identification effectivenessVSAvoidspeech intelligibility
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent applies multiple voice transformation parameters including pitch shift (e.g., +12 semitones), frequency shift (e.g., +500 Hz), and formant modification to achieve effective speaker de-identification. By adjusting these parameters within specific ranges, the system masks speaker identity while preserving speech intelligibility better than conventional single-parameter transforms.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system combines multiple voice transformation techniques (pitch shifting, frequency shifting, formant modification) into a composite transformation approach. This multi-layered transformation strategy achieves robust speaker masking while maintaining acceptable intelligibility levels, overcoming the limitations of single-method conventional transforms.

Inventive Principle:
Principle #40Composite materials

2Reliability

If voice transformation strength is increased to mask speaker identity, then speaker de-identification is improved, but transcription accuracy deteriorates

Engineering Contradiction:
Improvespeaker de-identification effectivenessVSAvoidtranscription accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system modifies multiple acoustic parameters (pitch, frequency, formants) simultaneously with controlled magnitudes. This multi-parameter approach achieves strong speaker masking while preserving enough speech structure for accurate transcription, unlike single-parameter transforms that sacrifice too much intelligibility.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies voice transformation at a strength that is partially sufficient for masking (not maximum possible transformation), thereby achieving adequate de-identification while preserving transcription accuracy. The transformation is calibrated to provide just enough masking without over-transforming and losing intelligibility.

Inventive Principle:
Principle #16Partial or excessive action

3Object-affected harmful factors

If existing voice transformation methods are used, then speaker masking is achieved, but speech becomes more difficult to understand and transcribe

Engineering Contradiction:
Improvespeaker identity maskingVSAvoidspeech understanding and transcription ease
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The system carefully selects and adjusts transformation parameters (pitch shift range, frequency shift magnitude, formant modification depth) to achieve speaker masking while minimizing impact on speech understanding. This controlled parameter modification preserves transcription ease better than conventional transforms.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The voice transformation acts as an intermediary that modifies speaker-specific characteristics while preserving speech content. The transformation serves as a mediator between complete privacy protection and full speech intelligibility, achieving a balanced state that maintains both masking effectiveness and transcription ease.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240370667A1System and method for voice morphing in a data annotator tool
Publication Date: 2024.11.07 SOUNDHOUND AI IP LLC
  • US20240370667A1 patent drawing
  • US20240370667A1 patent drawing
  • US20240370667A1 patent drawing

AI summary

A system and method for masking an identity of a speaker of natural language speech, such as speech clips to be labeled by humans in a system generating voice transcriptions for training an automatic speech recognition model. The natural language speech is morphed prior to being presented to the human for labeling. In one embodiment, morphing comprises pitch shifting the speech randomly either up or down, then frequency shifting the speech, then pitch shifting the speech in a direction opposite the first pitch shift. Labeling the morphed speech comprises at least one or more of transcribing the morphed speech, identifying a gender of the speaker, identifying an accent of the speaker, and identifying a noise type of the morphed speech.