Speech Synthesizer Phase Modulation for Audio Watermarking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesizers face challenges in embedding audio watermarking without deteriorating sound quality, as the insertion of watermarking information often leads to a decline in the quality of the synthesized speech.

Innovation Solution

A speech synthesizer configuration that includes a sound source generator, phase modulator, and vocal tract filter unit, where the phase of the pulse signal is modulated based on audio watermarking information, allowing for the insertion of watermarking without significantly affecting sound quality, using a phase modulation rule that varies over time or frequency bins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If audio watermarking information is embedded into synthesized speech, then speech identification and copyright protection are improved, but sound quality is deteriorated

Engineering Contradiction:
Improvewatermark information embeddingVSAvoidsound quality
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

The patent applies local quality by modulating the phase of pulse signals at specific locations (pitch marks) rather than uniformly across the entire signal. The phase modulation is applied selectively at pitch mark positions identified through autocorrelation analysis, allowing watermark embedding only in specific local regions where it least affects overall sound quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs dynamic phase modulation where the modulation amount varies over time based on the speech signal characteristics. The phase modulation amount is adjusted dynamically according to the instantaneous properties of the speech signal, allowing the system to adapt to changing signal conditions and minimize quality degradation while maintaining watermark detectability.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If phase modulation is applied to embed watermarking, then watermark detectability is improved, but speech naturalness is deteriorated

Engineering Contradiction:
Improvewatermark detection accuracyVSAvoidspeech naturalness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies partial action by modulating only a portion of the signal - specifically, only the phase at pitch mark positions rather than the entire signal spectrum. This partial modulation approach provides sufficient watermark information for detection while minimizing the overall impact on speech naturalness by leaving the majority of the signal unchanged.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the phase parameter of the pulse signal at pitch marks to embed watermark information. By modifying only the phase parameter rather than amplitude or frequency, the system achieves watermark embedding with minimal perceptual impact, as phase changes are less noticeable to human listeners compared to amplitude or frequency changes.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables the insertion of audio watermarking into synthesized speech without causing noticeable sound quality deterioration, ensuring that the speech remains of high quality even with the embedded watermarking information.

Implementation Method 1

a phase modulator that modulates, with respect to the sound source signal generated by the sound source generator, a phase of the pulse signal at each pitch mark based on audio watermarking information

Methodology Applied
Scientific EffectPhase modulation: Phase Modulation

Data Source

PatentUS10109286B2Speech synthesizer, audio watermarking information detection apparatus, speech synthesizing method, audio watermarking information detection method, and computer program product
Publication Date: 2018.10.23 TOSHIBA DIGITAL SOLUTIONS CORP
  • US10109286B2 patent drawing
  • US10109286B2 patent drawing
  • US10109286B2 patent drawing

AI summary

According to an embodiment, a speech synthesizer includes a source generator, a phase modulator, and a vocal tract filter unit. The source generator generates a source signal by using a fundamental frequency sequence and a pulse signal. The phase modulator modulates, with respect to the source signal generated by the source generator, a phase of the pulse signal at each pitch mark based on audio watermarking information. The vocal tract filter unit generates a speech signal by using a spectrum parameter sequence with respect to the source signal in which the phase of the pulse signal is modulated by the phase modulator.