Speech Anonymization via Neural Network Vocal Characteristic Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech-processing systems struggle to effectively anonymize speech by removing vocal characteristics of a source voice and adding those of a target voice while retaining phoneme characteristics.
Innovation Solution
A speech-processing system utilizing neural-network models as encoders and decoders processes audio data to extract and transfer phoneme and vocal characteristics, allowing for the anonymization of speech by replacing source voice characteristics with those of a target voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If existing speech-processing systems attempt to anonymize speech by removing vocal characteristics, then speaker identity is masked, but phoneme characteristics are lost or distorted
Solution Approach 1:
The speech signal is segmented into distinct components: phoneme characteristics and vocal characteristics. The neural network model separately processes and transforms these components, allowing phoneme information to be preserved while vocal characteristics are modified for anonymization.
Solution Approach 2:
A neural network model acts as an intermediary between the source speech and the anonymized output. The model learns to map source phoneme characteristics to target phoneme characteristics while simultaneously transforming source vocal characteristics to target vocal characteristics, thereby preserving essential speech content while achieving anonymization.
2Object-affected harmful factors
If neural network models are used to transfer vocal characteristics between voices, then speech anonymization is achieved, but computational complexity increases
Solution Approach 1:
The neural network model is designed to perform multiple functions simultaneously: it extracts phoneme characteristics, transforms vocal characteristics, and synthesizes the anonymized speech output. This multi-functionality reduces the need for separate processing stages and simplifies the overall system architecture despite the complexity of the neural network itself.
Data Source
AI summary
A speech-processing system receives first audio data correspond to a first voice and second audio data corresponding to a second voice. The speech-processing system determines vocal characteristics of the second voice and determines output corresponding to the first audio data and the vocal characteristics.


