Adversarial Watermark Embedding in Synthetic Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice conversion systems lack robustness against watermark removal, as simple watermarks can be easily broken by bad actors, and there is a need for a watermark that cannot be easily detected by humans but is detectable by machine learning systems.

Innovation Solution

A method and system that use adversarial neural networks and a watermark robustness module to generate and embed a detectable watermark in synthetic speech, ensuring the watermark remains undetectable to humans while being detectable by a trained machine learning system, even in noisy environments or when subjected to transformations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a simple watermark is embedded in synthetic speech, then the watermark is easily detectable by machine learning systems, but the watermark can be easily removed or broken by bad actors

Engineering Contradiction:
Improvewatermark robustnessVSAvoidwatermark removal capability
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system applies preliminary anti-action by training the generator to embed watermarks that are specifically designed to resist common attack transformations. The watermark is pre-configured with properties that make it difficult to remove, such as being embedded in multiple frequency domains and being robust to common audio processing operations like filtering, compression, and noise addition.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The patent applies parameter changes by transforming the watermark through multiple domains and applying various transformations to the speech data. The watermark is embedded in the frequency domain and time domain, and the system tests robustness by applying transformations such as adding noise, filtering, and compression to verify the watermark remains detectable.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If a watermark is made undetectable to humans, then the speech sounds natural and realistic, but the watermark becomes difficult to detect by machine learning systems

Engineering Contradiction:
Improvehuman perception of naturalnessVSAvoidwatermark detectability by machine learning system
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system applies local quality by embedding the watermark in specific local regions of the audio signal in a way that is imperceptible to human ears. The watermark is distributed across the audio spectrum and time domain in a localized manner that does not create audible artifacts, while still maintaining detectability through the trained machine learning system that knows where to look for the watermark signature.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses an intermediary approach by introducing a trained machine learning detector as the mediator between the embedded watermark and human perception. The detector is specifically trained to identify the watermark pattern, while humans perceive only the natural speech. The intermediary detector bridges the gap between the invisible watermark and human auditory perception.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the watermark is embedded in the speech data, then the watermark can be detected, but the watermark may be affected by transformations and noise

Engineering Contradiction:
Improvewatermark detection accuracyVSAvoidwatermark vulnerability to transformations
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system applies preliminary action by pre-training the generator to embed watermarks that are inherently robust to common transformations. The watermark is embedded in a way that anticipates potential attacks, and the training process includes exposing the system to various transformations to ensure the watermark remains detectable after such processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies beforehand cushioning by designing the watermark embedding to withstand common audio transformations and noise. The watermark is embedded with sufficient strength and in multiple domains to cushion against the effects of filtering, compression, and noise addition that may be applied to the speech data.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS11538485B2Generation and detection of watermark for real-time voice conversion
Publication Date: 2022.12.27 MODULATE INC
  • US11538485B2 patent drawing
  • US11538485B2 patent drawing
  • US11538485B2 patent drawing

AI summary

A method watermarks speech data by using a generator to generate speech data including a watermark. The generator is trained to generate the speech data including the watermark. The training process generates first speech from the generator. The first speech data is configured to represent speech. The first speech data includes a candidate watermark. The training also produces an inconsistency message as a function of at least one difference between the first speech data and at least authentic speech data. The training further includes transforming the first speech data, including the candidate watermark, using a watermark robustness module to produce transformed speech data including a transformed candidate watermark. The transformed speech data includes a transformed candidate watermark. The training further produces a watermark-detectability message, using a watermark detection machine learning system, relating to one or more desirable watermark features of the transformed candidate watermark.