Voice Morphing Apparatus Spectral Parameter Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice processing systems face challenges in providing speaker anonymity while maintaining audio fidelity, as existing methods often produce unintelligible or distorted outputs, and struggle to de-identify speakers while preserving non-identifying characteristics like noise and accent.
Innovation Solution
A voice morphing apparatus is trained using an objective function that balances speaker identification and audio fidelity, employing an artificial neural network architecture and existing speech processing components to adjust parameters through gradient descent, ensuring the output audio is recognizable by speech processing systems while masking the speaker's identity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speaker de-identification mapping is applied to mask speaker identity, then speaker identification certainty is reduced, but audio fidelity and speech intelligibility deteriorate
Solution Approach 1:
The patent transforms the voice signal from time-domain to frequency-domain representation (spectrogram), applies morphing operations in the frequency domain by modifying spectral parameters, and then reconstructs the time-domain signal. This parameter transformation approach allows effective speaker anonymization while preserving speech intelligibility by operating on spectral characteristics rather than raw audio waves.
Solution Approach 2:
The patent introduces an intermediate spectrogram representation as a mediator between the input audio signal and the output morphed audio. The voice morphing apparatus processes the audio through this intermediate frequency-domain representation, allowing speaker identity masking while maintaining speech content integrity through the spectral transformation and reconstruction process.
2Reliability
If voice morphing is applied to anonymize speaker identity, then speaker identification becomes difficult, but distinctive speech characteristics like accent and noise are lost
Solution Approach 1:
The patent applies different processing strategies to different aspects of the voice signal: aggressive morphing to speaker-specific spectral features to mask identity, while preserving linguistic and prosodic characteristics that convey speech meaning, accent, and emotional content. This localized quality modification approach maintains speech characteristics while achieving speaker anonymization.
Solution Approach 2:
The voice morphing apparatus dynamically adjusts morphing parameters during processing to balance speaker anonymization and preservation of speech characteristics. The system adaptively controls the degree of spectral modification to ensure speaker unidentifiability while maintaining natural speech quality and distinctive characteristics like accent and noise patterns.
3Reliability
If neural network mapping is used to transform speaker voice, then speaker identity is obscured, but output becomes unintelligible or heavily distorted
Solution Approach 1:
The patent replaces direct time-domain audio processing with frequency-domain spectral processing. By substituting the mechanical audio signal transformation with spectral parameter manipulation in the frequency domain, the system achieves speaker anonymization while preserving speech intelligibility through proper spectral reconstruction and inverse transformation.
Solution Approach 2:
The system incorporates feedback mechanisms during the voice morphing process to monitor and maintain speech intelligibility. The neural network is trained with objectives that include both speaker anonymization and intelligibility preservation, using feedback from speech recognition systems and quality metrics to adjust morphing parameters and ensure the output remains intelligible while obscuring speaker identity.
Data Source
AI summary
A voice morphing apparatus having adjustable parameters is described. The disclosed system and method include a voice morphing apparatus that morphs input audio to mask a speaker's identity. Parameter adjustment uses evaluation of an objective function that is based on the input audio and output of the voice morphing apparatus. The voice morphing apparatus includes objectives that are based adversarially on speaker identification and positively on audio fidelity. Thus, the voice morphing apparatus is adjusted to reduce identifiability of speakers while maintaining fidelity of the morphed audio. The voice morphing apparatus may be used as part of an automatic speech recognition system.


