Privacy enhancement type voice forgery detection method based on WeSpeaker architecture

Through the voice forgery detection method based on the WeSpeaker architecture, through acoustic-semantic decoupling and lightweight improved feature extraction, the problems of privacy leakage, high complexity and insufficient generalization ability in the existing technology are solved, and efficient voice forgery detection and privacy protection are achieved.

CN120748412AActive Publication Date: 2025-10-03NANJING LONGYUAN INFORMATION TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511013648.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-03
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing speech forgery detection technologies have problems such as privacy leakage risks, high model complexity and mismatch between target and task, and limited generalization ability.

Method used

A privacy-enhanced voice forgery detection method based on the WeSpeaker architecture is adopted. Acoustic-semantic decoupling technology is used for privacy protection preprocessing, a lightweight improved WeSpeaker architecture is used for audio feature extraction, and a lightweight fully connected binary classification layer is used for forgery discrimination, which weakens the dependence on semantic content and enhances generalization ability.

Benefits of technology

It improves the privacy protection capability of the model, reduces the complexity of the model, and enhances the generalization ability of unknown forgery techniques without leaking the semantic content of the speech. It is suitable for edge deployment on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748412A_ABST
    Figure CN120748412A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing, in particular to a privacy enhancement type voice forgery detection method based on a WeSpeaker architecture, and in specific use, the method comprises three stages: the first stage is an audio input and privacy protection preprocessing stage, and in the stage, privacy protection of voice content is realized through an acoustic-semantic decoupling technology; and the second stage is a feature extraction stage based on improved WeSpeaker, and depth extraction of audio features is carried out by using a lightweight improved WeSpeaker architecture. And the third stage is a forgery discrimination and decision-making stage, wherein the extracted features are subjected to'true / forgery 'discrimination through a lightweight full-connection dichotomy layer. And finally, giving a detection result about whether the audio is forged or not. According to the method, the technical problems of privacy leakage risk, high model complexity, mismatching of target tasks and limited generalization ability in actual use of a voice forgery detection technology in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture. Background Art

[0002] With the rapid development of artificial intelligence-generated content (AIGC) technology, text-to-speech (TTS) systems have become capable of generating highly natural, almost convincing human speech. In particular, recent advances in deep neural networks (DNNs), diffusion models, multi-speaker modeling, and voice transfer technologies have enabled attackers to exploit publicly available TTS tools to quickly synthesize the target speaker's impersonation, enabling them to carry out security attacks such as identity fraud, voice manipulation, and social deception. Consequently, anti-spoofing detection, a key research area in voice security, aims to automatically identify whether speech is generated by a synthesis system or a replay attack. It is a core security defense in the deployment of modern speech recognition, voiceprint recognition, voice assistants, and other systems. Furthermore, with increasing emphasis on data privacy, a growing number of researchers are focusing on how to perform speech recognition and voice security testing without exposing semantic information.

[0003] Existing solutions include voiceprint-based detection, deep learning-based end-to-end detection, and forgery detection mechanisms based on content stripping and privacy decoupling. However, these existing approaches have the following problems:

[0004] 1. The strong reliance on semantic content leads to unavoidable privacy risks: Detection methods based on voiceprint features or deep learning typically directly process raw speech or its spectral representation, which contains a wealth of information related to semantic content and speaker identity. In actual deployments, this approach inevitably collects, processes, and even stores the user's complete voice data, posing a serious risk of privacy leakage. This approach is particularly difficult to implement in regulatory or privacy-critical environments (such as smart assistants and financial voice interaction systems).

[0005] 2. High model complexity and a mismatch between the target task and the model make effective migration difficult. For example, while some existing high-performance voiceprint recognition architectures (such as WeSpeaker) possess powerful speaker modeling capabilities, their final classification layer is designed for speaker identity verification and cannot be directly applied to the "real / fake" binary classification scenario. Directly using the original architecture results in redundant speaker feature learning, making it difficult to adapt to the forgery detection task. Furthermore, the complex structure of these models presents challenges for deployment on resource-constrained devices.

[0006] 3. Overfitting to specific speech features limits generalization: Current deep learning-based forgery detection models tend to become dependent on the synthesis method, semantic content, or speaker style of the training data during training, resulting in significant performance degradation when faced with unseen forgery methods. Especially in recent years, with the continuous updates and iterations of TTS synthesis models, the detection model's ability to adapt to unknown attacks has become a key bottleneck to its practicality.

[0007] In summary, existing speech forgery detection technologies have the risks of privacy leakage, high model complexity, mismatch between target and task, and limited generalization ability when used in practice. Summary of the Invention

[0008] The purpose of the present invention is to provide a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture, aiming to solve the technical problems of the existing voice forgery detection technology in actual use, such as the risk of privacy leakage, high model complexity, mismatch between target tasks, and limited generalization ability.

[0009] To achieve the above objectives, the present invention adopts a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture, which includes the following stages:

[0010] Phase 1: Receive the audio input to be detected and perform privacy-preserving preprocessing based on acoustic-semantic decoupling technology;

[0011] Phase 2: Deep extraction of audio features using a lightweight and improved WeSpeaker architecture;

[0012] The third stage: The extracted features are judged as "real / forged" through a lightweight fully connected binary classification layer, and the detection result of whether the audio is forged is given.

[0013] Among them, in the first stage, the specific methods are as follows:

[0014] First, receive the audio input to be detected and perform basic feature extraction;

[0015] Then, an improved neural audio codec architecture is adopted to realize acoustic-semantic decoupling processing, decomposing the speech signal into acoustic representation and semantic representation.

[0016] Through a designed filter network, the system strips semantically relevant information from extracted features, retaining only the acoustic portion without semantic information for subsequent processing. This process effectively prevents the acquisition, reconstruction, or leakage of semantic content in speech, making it unusable for speech-to-text tasks. The semantic recovery success rate is less than 5%, and even in the face of specially designed semantic reconstruction attacks, the system maintains a high level of privacy protection, meeting high-level privacy and security requirements.

[0017] In the second stage, the specific method is as follows: using the streamlined and optimized WeSpeaker as the feature extraction backbone network;

[0018] Structural streamlining and optimization of WeSpeaker include:

[0019] Remove the speaker mapping layer to weaken speaker dependence;

[0020] Preserve and optimize the ECAPA-TDNN residual structure;

[0021] Improved statistical pooling mechanism to capture timing anomalies;

[0022] A channel attention mechanism is introduced to enhance spectrum perception capabilities.

[0023] Among them, when using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the speaker mapping layer is removed and the speaker-dependent speech is weakened to enhance the system's generalization ability for different forged audio generators, avoiding misjudging "speaker changes" as "forgery".

[0024] Among them, when using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the ECAPA-TDNN residual structure is retained and optimized, and the residual network structure of the multi-scale temporal context modeling and channel attention mechanism in ECAPA-TDNN is inherited, which enhances the time-frequency domain coupling expression capability of the features without adding additional parameter overhead; by deeply extracting the local and global acoustic structure information in the speech signal, the system can more easily identify the unnatural distortion of the synthesized speech at the micro level.

[0025] Among them, when using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the statistical pooling mechanism is improved to capture timing anomalies. The configuration of the statistical pooling layer is tuned to make it more sensitive to capturing nonlinear timing changes in speech, thereby enhancing the ability to detect rhythm / rhythm anomalies in synthesized speech.

[0026] Among them, when using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the channel attention mechanism is introduced to enhance spectrum perception capabilities. On the basis of the original TDNN framework, a lightweight SE channel attention module is added to adaptively weight the importance of each frequency channel, so that the model can automatically highlight the abnormal energy distribution of forged speech in specific frequency bands, thereby further improving the accuracy of forged recognition.

[0027] In the third stage, a newly added lightweight fully connected binary classification layer is used to directly judge the "real / fake" of the extracted features;

[0028] The discriminant model is trained with multi-source forged data, including forged data generated by a variety of representative TTS technologies. The diversity of training data is enhanced through reverberation, noise, and speed change techniques, and a domain adversarial learning strategy is introduced to reduce the overfitting of specific features of the synthesis technology.

[0029] Among them, various representative TTS technologies used to generate fake data include autoregressive synthesis technology, non-autoregressive streaming synthesis technology, diffusion model synthesis technology, vocoder fusion technology and neural sound conversion technology.

[0030] The present invention discloses a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture. When used specifically, the method includes three stages. The first stage is the audio input and privacy protection preprocessing stage, which implements privacy protection of voice content through acoustic-semantic decoupling technology. The second stage is the feature extraction stage based on the improved WeSpeaker, which uses the lightweight improved WeSpeaker architecture to perform in-depth extraction of audio features. The third stage is the forgery discrimination and decision-making stage, which uses a lightweight fully connected binary classification layer to perform "real / forged" discrimination on the extracted features. Finally, a detection result of whether the audio is forged will be given. In this way, the technical problems of the existing voice forgery detection technology in actual use, such as the risk of privacy leakage, high model complexity, mismatch between target tasks, and limited generalization ability, are solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 This is a flow chart of the privacy-enhanced voice forgery detection method based on the WeSpeaker architecture of the present invention. DETAILED DESCRIPTION

[0033] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be understood as limiting the present invention.

[0034] See also Figure 1 , Figure 1 This is a flow chart of the privacy-enhanced voice forgery detection method based on the WeSpeaker architecture of the present invention.

[0035] The present invention provides a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture, which includes the following stages:

[0036] Phase 1: Receive the audio input to be detected and perform privacy-preserving preprocessing based on acoustic-semantic decoupling technology;

[0037] For this specific implementation, firstly, the audio input to be detected is received and basic features are extracted;

[0038] Then, an improved neural audio codec architecture is adopted to realize acoustic-semantic decoupling processing, decomposing the speech signal into acoustic representation and semantic representation.

[0039] Through a designed filter network, the system strips semantically relevant information from extracted features, retaining only the acoustic portion without semantic information for subsequent processing. This process effectively prevents the acquisition, reconstruction, or leakage of semantic content in speech, making it unusable for speech-to-text tasks. The semantic recovery success rate is less than 5%, and even in the face of specially designed semantic reconstruction attacks, the system maintains a high level of privacy protection, meeting high-level privacy and security requirements.

[0040] Phase 2: Deep extraction of audio features using a lightweight and improved WeSpeaker architecture;

[0041] For this specific implementation, the streamlined and optimized WeSpeaker is used as the feature extraction backbone network;

[0042] Structural streamlining and optimization of WeSpeaker include:

[0043] The speaker mapping layer is removed to reduce speaker dependency. The original WeSpeaker design included a mapping layer for speaker classification, but this layer has been completely removed in this paper. This significantly reduces the risk of the model overfitting to speaker characteristics, thereby enhancing the system's generalization across different forged audio generators and avoiding misclassification of "speaker changes" as "forgeries."

[0044] The ECAPA-TDNN residual structure is retained and optimized. To this end, the residual network structure of the multi-scale temporal context modeling and channel attention mechanism in ECAPA-TDNN is inherited, which enhances the time-frequency domain coupling expression capability of features without adding additional parameter overhead. By deeply extracting the local and global acoustic structure information in the speech signal, the system can more easily identify unnatural distortions of synthesized speech at the micro level.

[0045] Improve the statistical pooling mechanism to capture timing anomalies; for this, in response to the common characteristics of forged audio such as "lack of frequency periodicity" and "overly smooth time dynamics", the present invention adjusts the configuration of the statistical pooling layer to make it more sensitive to capturing nonlinear timing changes in speech and enhance the ability to detect rhythm / rhythm anomalies in synthesized speech.

[0046] A channel attention mechanism is introduced to enhance spectrum perception. To address common characteristics of forged audio, such as a lack of frequency periodicity and overly smooth temporal dynamics, the present invention optimizes the configuration of the statistical pooling layer, making it more sensitive to nonlinear temporal variations in speech and improving the ability to detect rhythmic / prosodic anomalies in synthesized speech.

[0047] The third stage: The extracted features are judged as "real / forged" through a lightweight fully connected binary classification layer, and the detection result of whether the audio is forged is given.

[0048] In this specific implementation, the system uses a newly added lightweight, fully connected binary classification layer to directly distinguish "real / fake" from extracted features. This discriminator is trained using multi-source forged data, including data generated by five representative TTS technologies: autoregressive synthesis, non-autoregressive streaming synthesis, diffusion model synthesis, vocoder fusion, and neural voice conversion. Techniques such as reverberation, noise, and speed variation are used to enhance training data diversity, and a domain adversarial learning strategy is introduced to reduce overfitting of specific features of the synthesis techniques.

[0049] By removing the original WeSpeaker's speaker mapping layer (the fully connected layer used for speaker classification), its residual structure and feature extraction capabilities are retained. A lightweight fully connected binary classification layer is added at the end to directly serve the task of distinguishing "real / forged" audio. This improvement inherits WeSpeaker's strong modeling capabilities while reducing the risk of speaker overfitting, making it more suitable for forgery detection tasks.

[0050] By integrating the core mechanism of the neural audio codec structure, the original speech is subjected to acoustic-semantic decoupling processing, and only the acoustic part without semantic information is retained for subsequent feature extraction and discrimination, preventing the collection, reconstruction or leakage of speech semantic content, and achieving a higher level of privacy protection requirements.

[0051] This paper uses a variety of representative TTS techniques (five mainstream TTS methods) to construct forged data during the training phase. During the validation phase, it uses a novel, unseen TTS technique for evaluation, ensuring that the proposed model has good generalization capabilities across synthesis systems. Through privacy-enhancing mechanisms, the model's reliance on semantic content is further reduced, allowing it to focus more on unnatural acoustic structural differences in forged speech.

[0052] The beneficial effects of the present invention are as follows:

[0053] 1. Acoustic-semantic decoupling to ensure voice privacy: This invention introduces a neural audio codec architecture at the audio input stage to deeply decouple the acoustics and semantics of the voice signal, retaining only non-semantic acoustic features for forgery detection. This effectively prevents the voice content from being reversely recovered or misused, significantly improving the system's privacy protection capabilities and making it suitable for practical scenarios with higher data compliance requirements.

[0054] 2. Optimized WeSpeaker feature extractor: While retaining the advantages of the ECAPA-TDNN residual network architecture, the system removes the speaker mapping layer, optimizes the statistical pooling strategy, and incorporates a channel attention mechanism. This enhances the model's ability to detect features such as spectral anomalies and timing inconsistencies in speech forgeries, improving the generalization detection of unknown forgery types.

[0055] 3. Training strategy integrating multi-source forged samples: The system introduces multiple types of forged TTS speech data (including diffusion models, neural vocoders, etc.) for training during the discrimination phase. At the same time, it enhances sample diversity through methods such as reverberation and noise. Combined with domain adversarial training, it further alleviates the model's overfitting to specific technologies, enabling the system to maintain high accuracy and robustness in complex environments and under unknown forgery technologies.

[0056] 4. Lightweight design adapts to edge deployment requirements: The system that implements this method adopts a modular and lightweight design, which takes into account both detection performance and privacy protection without significantly increasing the computational burden. It has strong practicality and is suitable for low-power, real-time application scenarios such as smart terminals, voice customer service, and in-vehicle voice.

[0057] The present invention uses a privacy-enhanced voice forgery detection method based on the WeSpeaker architecture. When used specifically, the method includes three stages. The first stage is the audio input and privacy protection preprocessing stage, which implements the privacy protection of the voice content through the acoustic-semantic decoupling technology. The second stage is the feature extraction stage based on the improved WeSpeaker, which uses the lightweight improved WeSpeaker architecture to perform in-depth extraction of audio features. The third stage is the forgery discrimination and decision-making stage, which uses a lightweight fully connected binary classification layer to perform "real / forged" discrimination on the extracted features. Finally, a detection result of whether the audio is forged will be given. In this way, the technical problems of the existing voice forgery detection technology in actual use, such as the risk of privacy leakage, high model complexity, mismatch between target tasks, and limited generalization ability, are solved.

[0058] The above disclosure is only a preferred embodiment of the present invention, and certainly cannot be used to limit the scope of the rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.

Claims

1. A privacy-enhanced voice forgery detection method based on the WeSpeaker architecture, characterized in that: The following stages are included: Phase 1: Receive the audio input to be detected and perform privacy-preserving preprocessing based on acoustic-semantic decoupling technology; Phase 2: Deep extraction of audio features using a lightweight and improved WeSpeaker architecture; The third stage: The extracted features are judged as "real / forged" through a lightweight fully connected binary classification layer, and the detection result of whether the audio is forged is given.

2. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 1, characterized in that: In the first phase, the specific methods are as follows: First, receive the audio input to be detected and perform basic feature extraction; Then, an improved neural audio codec architecture is adopted to realize acoustic-semantic decoupling processing, decomposing the speech signal into acoustic representation and semantic representation.

3. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 1, characterized in that: In the second stage, the specific method is as follows: using the streamlined and optimized WeSpeaker as the feature extraction backbone network; Structural streamlining and optimization of WeSpeaker include: Remove the speaker mapping layer to weaken speaker dependence; Preserve and optimize the ECAPA-TDNN residual structure; Improved statistical pooling mechanism to capture timing anomalies; A channel attention mechanism is introduced to enhance spectrum perception capabilities.

4. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 3, characterized in that: When using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the speaker mapping layer is removed and speaker-dependent speech is weakened to enhance the system's generalization ability for different forged audio generators and avoid misjudging "speaker changes" as "forgery".

5. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 3, characterized in that: When using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the ECAPA-TDNN residual structure is retained and optimized, inheriting the residual network structure of the multi-scale temporal context modeling and channel attention mechanism in ECAPA-TDNN, enhancing the time-domain-frequency domain coupling expression capability of features without adding additional parameter overhead; by deeply extracting local and global acoustic structure information from speech signals, the system can more easily identify unnatural distortions in synthesized speech at the micro level.

6. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 3, characterized in that: When using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the statistical pooling mechanism is improved to capture timing anomalies. The configuration of the statistical pooling layer is tuned to make it more sensitive to capturing nonlinear timing changes in speech, thereby enhancing the ability to detect rhythm / prosody anomalies in synthesized speech.

7. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 3, characterized in that: When using the streamlined and optimized WeSpeaker as the feature extraction backbone network, the channel attention mechanism is introduced to enhance spectrum perception capabilities. On the basis of the original TDNN framework, a lightweight SE channel attention module is added to adaptively weight the importance of each frequency channel, enabling the model to automatically highlight the abnormal energy distribution of forged speech in specific frequency bands, thereby further improving the accuracy of forged speech recognition.

8. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 1, characterized in that: In the third stage, a newly added lightweight fully connected binary classification layer is used to directly judge the "real / fake" of the extracted features; The discriminant model is trained with multi-source forged data, including forged data generated by a variety of representative TTS technologies. The diversity of training data is enhanced through reverberation, noise, and speed change techniques, and a domain adversarial learning strategy is introduced to reduce the overfitting of specific features of the synthesis technology.

9. The privacy-enhanced voice forgery detection method based on the WeSpeaker architecture as claimed in claim 8, characterized in that: Various representative TTS technologies for generating fake data include autoregressive synthesis technology, non-autoregressive streaming synthesis technology, diffusion model synthesis technology, vocoder fusion technology, and neural voice conversion technology.

Citation Information

Patent Citations

  • Intelligent voice forgery attack detection method based on attention mechanism

    CN116416997A

  • Voice detection method, voice detection device, electronic equipment and storage medium

    CN116543794A

  • Synthetic speech detection method based on rhythm characteristics

    CN116665649A

  • Multi-countermeasure discriminative forged audio detection system

    CN118280389A

  • Decoupling type voice self-supervision pre-training method

    CN118841029A