Cross-Domain Speech Deepfake Detection Under Compression Shifts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech deepfake detection methods struggle with high false-positive rates and performance degradation under real-world conditions due to domain shifts and environmental distortions caused by speech compression and transmission variability, particularly in internet-based communication systems.

Innovation Solution

A cross-domain speech deepfake detection method integrating auxiliary cross-domain data generation, self-supervised feature extraction, domain-invariant representation learning, and one-class learning to enhance detection accuracy and robustness across diverse scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech forensic methods are used to analyze acoustic features, then detection can be performed on speech patterns, but false-positive rates increase and performance degrades under real-world conditions

Engineering Contradiction:
Improvedetection accuracyVSAvoidfalse-positive rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces domain adaptation as an intermediary technique that bridges the gap between controlled training environments and real-world detection scenarios. By learning domain-invariant representations and adapting features across different domains (e.g., from clean studio recordings to compressed VoIP transmissions), the system maintains high detection accuracy while reducing false positives caused by environmental variations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent dynamically adjusts detection parameters and feature weighting based on the detected domain characteristics. When speech compression or transmission artifacts are detected, the system modifies its analysis parameters to account for these distortions, thereby maintaining measurement precision across varying real-world conditions without increasing false-positive rates.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If detection methods focus on identifying anomalies in synthesized speech, then deepfake artifacts can be detected, but performance degrades under speech compression and channel distortions

Engineering Contradiction:
Improvedeepfake detection accuracyVSAvoidrobustness to compression
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent performs domain adaptation and compression simulation in advance during the training phase. By pre-exposing the model to various compression artifacts and channel distortions through auxiliary cross-domain data generation, the system learns to recognize deepfake anomalies independently of compression effects, thereby maintaining detection accuracy across different transmission channels.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent converts the harmful effect of speech compression into a beneficial training mechanism. By deliberately compressing training data through various codecs and transmission simulations, the system learns to distinguish between compression-induced artifacts and deepfake anomalies, thereby improving robustness while maintaining detection precision.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Ease of manufacture

If neural network-based deepfake models are used to generate speech, then realistic speech content is produced, but detection becomes increasingly difficult under varying transmission conditions

Engineering Contradiction:
Improvespeech generation qualityVSAvoiddetection complexity
Core Design Contradiction:
Ease of manufactureVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the detection process into multiple domain-specific stages. Instead of attempting to detect deepfakes in a single unified model, the system divides detection into domain-adapted modules that handle different transmission conditions separately. This segmentation allows the system to maintain simplicity in each module while achieving comprehensive detection capability across diverse scenarios.

Inventive Principle:
Principle #1Segmentation

4Reliability

If existing detection methods are applied to internet-based communication systems, then deepfake detection can be performed, but false alarms increase due to domain shifts

Engineering Contradiction:
Improvedetection capabilityVSAvoidfalse alarm rate
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent creates a universal detection framework that functions across multiple domains and transmission channels. By training on auxiliary cross-domain data that simulates various internet-based communication scenarios, the system learns domain-invariant features that generalize across different platforms, thereby maintaining reliable detection while minimizing false alarms caused by domain-specific variations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12475896B1Method for cross-domain speech deepfake detection
Publication Date: 2025.11.18 ZHEJIANG GONGSHANG UNIVERSITY
  • US12475896B1 patent drawing
  • US12475896B1 patent drawing
  • US12475896B1 patent drawing

AI summary

A method for detecting speech deepfakes under real-world conditions is disclosed, where factors such as audio compression, transmission channels, and codec transformations may distort the speech signal. The invention combines self-supervised learning pre-trained models, domain-invariant representation learning, and one-class learning to improve the robustness and generalization of deepfake detection systems. By aligning feature distributions across domains, the system ensures consistent representation of genuine speech, even under varying conditions, and effectively identifies synthetic speech. The proposed method offering a robust solution for real-world audio deepfake detection applications.