Speech Synthesizer Phase Modulation for Audio Watermarking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesizers face challenges in embedding audio watermarking without deteriorating sound quality, as the insertion of watermarking information often leads to a decline in the quality of the synthesized speech.
Innovation Solution
A speech synthesizer configuration that includes a sound source generator, phase modulator, and vocal tract filter unit, where the phase of the pulse signal is modulated based on audio watermarking information, allowing for the insertion of watermarking without significantly affecting sound quality, using a phase modulation rule that varies over time or frequency bins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If audio watermarking information is embedded into synthesized speech, then speech identification and copyright protection are improved, but sound quality is deteriorated
Solution Approach 1:
The patent applies local quality by modulating the phase of pulse signals at specific locations (pitch marks) rather than uniformly across the entire signal. The phase modulation is applied selectively at pitch mark positions identified through autocorrelation analysis, allowing watermark embedding only in specific local regions where it least affects overall sound quality.
Solution Approach 2:
The patent employs dynamic phase modulation where the modulation amount varies over time based on the speech signal characteristics. The phase modulation amount is adjusted dynamically according to the instantaneous properties of the speech signal, allowing the system to adapt to changing signal conditions and minimize quality degradation while maintaining watermark detectability.
2Measurement precision
If phase modulation is applied to embed watermarking, then watermark detectability is improved, but speech naturalness is deteriorated
Solution Approach 1:
The patent applies partial action by modulating only a portion of the signal - specifically, only the phase at pitch mark positions rather than the entire signal spectrum. This partial modulation approach provides sufficient watermark information for detection while minimizing the overall impact on speech naturalness by leaving the majority of the signal unchanged.
Solution Approach 2:
The patent changes the phase parameter of the pulse signal at pitch marks to embed watermark information. By modifying only the phase parameter rather than amplitude or frequency, the system achieves watermark embedding with minimal perceptual impact, as phase changes are less noticeable to human listeners compared to amplitude or frequency changes.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables the insertion of audio watermarking into synthesized speech without causing noticeable sound quality deterioration, ensuring that the speech remains of high quality even with the embedded watermarking information.
Implementation Method 1
a phase modulator that modulates, with respect to the sound source signal generated by the sound source generator, a phase of the pulse signal at each pitch mark based on audio watermarking information
Data Source
AI summary
According to an embodiment, a speech synthesizer includes a source generator, a phase modulator, and a vocal tract filter unit. The source generator generates a source signal by using a fundamental frequency sequence and a pulse signal. The phase modulator modulates, with respect to the source signal generated by the source generator, a phase of the pulse signal at each pitch mark based on audio watermarking information. The vocal tract filter unit generates a speech signal by using a spectrum parameter sequence with respect to the source signal in which the phase of the pulse signal is modulated by the phase modulator.


