Speech Signal Watermarking via Phase Sequence Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current watermarking techniques for speech signals are inadequate in ensuring authentication and quality, particularly with the rise of text-to-speech technologies that can mimic human speech, leading to unauthorized use and spoofing.

Innovation Solution

A method involving the application of a watermark signal to speech signals using a phase sequence and bit encoding, spread across frequency bins, with a frequency-dependent gain factor, and incorporating Pretty Good Privacy (PGP) or public key cryptography, to authenticate and secure speech signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If current watermarking techniques are applied to speech signals, then authentication capability is improved, but audio quality deteriorates

Engineering Contradiction:
Improveauthentication capabilityVSAvoidaudio quality
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent transforms the watermarking approach by changing the domain parameters from time-domain to frequency-domain representation using spectrograms. By operating in the frequency domain and manipulating phase sequences rather than directly modifying amplitude, the watermark becomes imperceptible to human hearing while maintaining robust authentication capability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical/audio signal manipulation with cryptographic methods. By incorporating PGP or public key cryptography into the watermark signal, the system achieves strong authentication without requiring high watermark power, thus avoiding audio quality degradation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If watermarking is applied to prevent unauthorized copying, then security is improved, but susceptibility to splicing attacks worsens

Engineering Contradiction:
ImprovesecurityVSAvoidsplicing attack vulnerability
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent divides the speech signal into multiple frames and applies watermarking to each frame's spectrogram independently. The phase sequence watermark is distributed across frequency bins of different frames, making it difficult for attackers to splice segments without detecting the watermark discontinuities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent moves the watermarking operation from the time domain to the frequency-time domain using spectrograms. By embedding watermarks in the phase domain rather than amplitude domain, and distributing them across frequency bins, the system creates multiple dimensions of protection that resist splicing attacks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If traditional watermarking methods are used, then detection capability is improved, but inaudibility worsens

Engineering Contradiction:
Improvedetection capabilityVSAvoidaudibility
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent changes the parameter domain from amplitude to phase in the frequency representation. By modifying phase sequences rather than amplitude, the watermark remains imperceptible to human hearing (which is more sensitive to amplitude changes) while maintaining detectability through phase correlation analysis.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses the spectrogram as an intermediary representation between the original speech signal and the watermark. The phase sequence operates in this intermediate frequency-time domain, allowing detection through spectral analysis while remaining inaudible in the time-domain audio signal.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240386898A1Tamper-robust watermarking of speech signals
Publication Date: 2024.11.21 CERENCE OPERATING CO
  • US20240386898A1 patent drawing
  • US20240386898A1 patent drawing
  • US20240386898A1 patent drawing

AI summary

A method for applying a watermark signal to a speech signal to prevent unauthorized use of speech signals, the method may include receiving an original speech signal; determining a corresponding spectrogram of the original speech signal; selecting a phase sequence of fixed frame length and uniform distribution; and generating an encoded watermark signal based on the corresponding spectrogram and phase sequence.