Speech Signal Watermarking via Phase Sequence Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current watermarking techniques for speech signals are inadequate in ensuring authentication and quality, particularly with the rise of text-to-speech technologies that can mimic human speech, leading to unauthorized use and spoofing.
Innovation Solution
A method involving the application of a watermark signal to speech signals using a phase sequence and bit encoding, spread across frequency bins, with a frequency-dependent gain factor, and incorporating Pretty Good Privacy (PGP) or public key cryptography, to authenticate and secure speech signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current watermarking techniques are applied to speech signals, then authentication capability is improved, but audio quality deteriorates
Solution Approach 1:
The patent transforms the watermarking approach by changing the domain parameters from time-domain to frequency-domain representation using spectrograms. By operating in the frequency domain and manipulating phase sequences rather than directly modifying amplitude, the watermark becomes imperceptible to human hearing while maintaining robust authentication capability.
Solution Approach 2:
The patent replaces traditional mechanical/audio signal manipulation with cryptographic methods. By incorporating PGP or public key cryptography into the watermark signal, the system achieves strong authentication without requiring high watermark power, thus avoiding audio quality degradation.
2Reliability
If watermarking is applied to prevent unauthorized copying, then security is improved, but susceptibility to splicing attacks worsens
Solution Approach 1:
The patent divides the speech signal into multiple frames and applies watermarking to each frame's spectrogram independently. The phase sequence watermark is distributed across frequency bins of different frames, making it difficult for attackers to splice segments without detecting the watermark discontinuities.
Solution Approach 2:
The patent moves the watermarking operation from the time domain to the frequency-time domain using spectrograms. By embedding watermarks in the phase domain rather than amplitude domain, and distributing them across frequency bins, the system creates multiple dimensions of protection that resist splicing attacks.
3Measurement precision
If traditional watermarking methods are used, then detection capability is improved, but inaudibility worsens
Solution Approach 1:
The patent changes the parameter domain from amplitude to phase in the frequency representation. By modifying phase sequences rather than amplitude, the watermark remains imperceptible to human hearing (which is more sensitive to amplitude changes) while maintaining detectability through phase correlation analysis.
Solution Approach 2:
The patent uses the spectrogram as an intermediary representation between the original speech signal and the watermark. The phase sequence operates in this intermediate frequency-time domain, allowing detection through spectral analysis while remaining inaudible in the time-domain audio signal.
Data Source
AI summary
A method for applying a watermark signal to a speech signal to prevent unauthorized use of speech signals, the method may include receiving an original speech signal; determining a corresponding spectrogram of the original speech signal; selecting a phase sequence of fixed frame length and uniform distribution; and generating an encoded watermark signal based on the corresponding spectrogram and phase sequence.


