Counterfeit voice detection method based on feature reconstruction generation

By using the front-end feature generation module in forged speech detection, and combining the weighted fusion and confidence scores of cosine similarity conversion of the back-end classification module, the problem of insufficient detection accuracy and robustness of forged speech in the prior art is solved, and more efficient and stable detection performance is achieved.

CN120108424APending Publication Date: 2025-06-06NINGBO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231323.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Among the existing forged speech detection methods, traditional manual feature extraction methods are difficult to capture the subtle differences in forged speech. Deep learning methods have high demands on data and computing resources, and insufficient generalization capabilities, resulting in insufficient detection accuracy and robustness, and are unable to effectively deal with the problems of complex scenarios and resource-constrained environments.

Method used

A forged speech detection method is proposed for reconstructing and generating features. The original spectrum features are reconstructed through the front-end feature generation module, amplify the specific features of the forged speech, and further improve the detection accuracy through the confidence scores of weighted fusion and cosine similarity conversion in the back-end classification module.

Benefits of technology

It significantly improves the detection capability and accuracy of the forged speech detection model, enhances the difference in authentic and false speech characteristics, improves the performance stability of the model under different test conditions, and reduces the dependence on a large amount of labeled data and powerful computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108424A_ABST
    Figure CN120108424A_ABST
Patent Text Reader

Abstract

The invention discloses a forged voice detection method for feature reconstruction generation. A front-end feature generation module is utilized to reconstruct original spectrum features extracted by a front-end feature extraction module, amplify the specific features of forged voice, enhance the difference between true and false voice features and improve the detection capability of a forged voice detection model; carrying out weighted fusion on the original encoder features and the reconstructed encoder features, enabling the fused features to pass through a classifier in a rear-end classification module to obtain a first confidence score, calculating cosine similarity between the original encoder features and the reconstructed encoder features, converting the cosine similarity into a second confidence score, and obtaining a second confidence score; the speech authenticity detection confidence is obtained through weighted calculation of the two confidence scores, and the detection precision of a forged speech detection model is improved; in addition, the method reduces the dependence on a large amount of annotation data and powerful computing resources, improves the training efficiency and real-time performance of the model, and enables the model to be more suitable for being applied to a resource-limited scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of information security technology, and relates to a forged voice detection technology, and in particular to a forged voice detection method that reconstructs and generates features. Background Art

[0002] With the rapid development of deep learning technology and multimedia processing technology, artificial intelligence generated content (AIGC) has become a key technology that has attracted much research attention. AIGC technology uses advanced technologies in the fields of deep learning and multimedia processing to successfully promote the widespread application of synthetic conversion speech in actual scenarios, bringing excellent interactive experience to human-computer interaction scenarios, especially in the fields of virtual assistants and voice navigation. By imitating the voice of a specific individual, synthetic conversion speech technology gives human-computer dialogue a more natural feeling and provides a higher degree of personalized services, demonstrating the significant advantages of AIGC technology in creativity and innovation.

[0003] However, this widespread application is also accompanied by potential information security issues, especially in the field of voice forgery. Voice forgery technology allows attackers to convert the voice of one speaker into the voice of another speaker while keeping the voice content unchanged, providing attackers with the opportunity to imitate the voice characteristics of a specific individual, thereby achieving the purpose of illegally obtaining benefits. In fields with extremely high requirements for information security, such as finance and security, this may lead to false identity authentication and fraud, which will cause serious social problems.

[0004] In fact, fake voice technology has indeed begun to be used by criminals in illegal or malicious scenarios. To ensure the benign application of AIGC technology, researchers urgently need to pay attention to and solve related information security issues, especially in the field of synthetic voice conversion technology. It is crucial to develop a detection mechanism for fake voice.

[0005] With the continuous advancement of speech generation technology, especially the emergence of speech synthesis and conversion methods based on deep learning, forged speech has become more and more realistic, which has prompted the continuous development of forged speech detection technology. Existing forged speech detection models can be divided into front-end and back-end separation pipeline models and end-to-end models.

[0006] The model structure of the front-end and back-end separation pipeline model is as follows Figure 1As shown in the figure, the forged speech detection task is divided into two parts: front-end feature extraction and back-end classification. The front-end feature extraction is mainly responsible for extracting effective speech features from the original speech signal, while the back-end classification is responsible for classifying these speech features and determining whether the original speech signal is forged speech. Since the front-end and back-end are separated, the feature extraction process is easier to understand, and the feature extraction method can be adjusted as needed. However, if the front-end speech feature extraction is not good enough, the performance of the back-end classification will also be affected. The quality of speech feature extraction directly affects the performance of the back-end classification. At present, speech feature extraction methods are mainly divided into two categories: traditional manual feature extraction and feature learning based on deep learning.

[0007] Traditional manual feature extraction methods, such as MFCC (Mel-Frequency Cepstral Coefficients) extraction method and LFCC (Linear Frequency Cepstral Coefficients) extraction method, can stably extract features from the original speech through fixed formulas and steps. These methods have the advantages of high computational efficiency and simple implementation, and are not easily affected by external conditions such as background noise. However, its fixed calculation method is difficult to amplify and optimize the specific features of forged speech, and it is impossible to optimize the generation of key features that can distinguish between real speech and forged speech. For example, when faced with high-imitation speech forged by generative models, the features extracted by manual features often fail to highlight the subtle differences and abnormal features between real speech and forged speech. In addition, the features extracted by manual features are universal and less dependent on data, but lack targeted optimization of forged speech features. This method can only rely on the back-end classifier to complete the speech classification task, and the classification performance is heavily dependent on the quality of feature extraction. In the forged speech detection scenario, this method of extracting features from the original speech through fixed formulas and steps will reduce the ability of the back-end classifier to distinguish different types of forged speech, especially in the face of cross-dataset or cross-language scenarios. Therefore, the traditional manual feature extraction method has obvious limitations in forged speech detection and is difficult to meet the requirements of high precision, high robustness and high efficiency.

[0008] In recent years, feature learning methods based on deep learning have made significant progress in the field of forged speech detection, mainly including convolutional neural networks (CNN), recurrent neural networks (RNN) and their variants (such as long short-term memory networks LSTM and gated recurrent units GRU), and pre-trained models. Convolutional neural networks automatically extract local features of speech signals through convolutional layers and pooling layers, and can capture the time series features in speech signals. They perform well in processing speech signals and can focus on and extract the most important features for the current task. Recurrent neural networks and their variants can process sequence data and capture the time dependencies in speech signals. They are particularly suitable for processing the time series characteristics of speech signals and can effectively capture dynamic changes in speech. In addition, pre-trained models are a type of deep learning model. By pre-training on large-scale unsupervised or weakly supervised data, they learn general feature representations, such as BERT, Wav2Vec2, etc., which have achieved great success in the fields of natural language processing and speech processing. The pre-trained model transfers the general features learned in the pre-training stage to specific tasks through transfer learning, and adapts to specific application scenarios through fine-tuning, which can significantly improve the accuracy and robustness of forged speech detection. However, its parameters are usually large (usually reaching several GB), and its storage and operation require high hardware resources, making it difficult to deploy efficiently on low-performance embedded devices.

[0009] Although deep learning-based feature learning methods have performed well in forged speech detection, these methods also have some limitations. Deep learning models usually require a large amount of annotated data and powerful computing resources to train. For example, convolutional neural networks and recurrent neural networks require a large amount of data to learn complex feature representations, and the training process of pre-trained models requires massive data and computing resources. The generalization ability of deep learning models in cross-domain or low-resource scenarios may be insufficient. For example, the number of parameters of pre-trained models is usually large, and the hardware resources required for storage and operation are high. It is difficult to deploy efficiently on low-performance embedded devices, which limits its application in resource-constrained scenarios (such as portable devices or real-time processing). If there is a lack of sufficient data support, the discriminative performance of deep learning models can only maintain good performance in the field where the training data is located. Once applied to multi-speaker scenarios or complex background sound environments, its performance may drop significantly, and the final discrimination accuracy may even be lower than that of manual feature extraction.

[0010] The end-to-end model is different from the above-mentioned front-end and back-end separation pipeline model. Its model structure is as follows Figure 2As shown in the figure, it aims to directly output the discrimination result of forged speech starting from the original speech signal input without explicit intermediate steps or feature extraction processes. The end-to-end model is trained end-to-end through a unified, differentiable network such as a deep neural network (DNN), and can automatically learn the mapping relationship between the original speech signal and the forged speech discrimination. For example, in speech recognition, the end-to-end model can directly map from the speech signal to the text without extracting features first and then classifying. The overall process of the end-to-end model is simple and direct, and there is no need to manually design features. Through training, the most relevant features for forged speech discrimination are automatically learned, and more details can be obtained from the original speech signal. However, since the end-to-end model needs to process the entire speech signal, it requires a large amount of training data and computing resources. In addition, when the data set is insufficient or contains multiple noises, the generalization ability of the end-to-end model is usually poor, and it is easy to overfit due to insufficient model training, which ultimately leads to performance degradation in practical applications.

[0011] In addition, voiceprint recognition and multimodal feature fusion methods have also been used to improve detection performance, but they still face the problem of balancing real-time performance and computational efficiency, especially in resource-constrained scenarios. Forged voice detection technology needs to develop in the direction of high precision, high robustness, and high efficiency.

[0012] The shortcomings of the above-mentioned prior art make forged voice detection face many difficulties in practical applications, and a new technical solution is urgently needed to solve these problems. Summary of the invention

[0013] In view of the fact that traditional manual feature extraction methods in existing forged speech detection methods are difficult to capture the subtle differences of forged speech and lack targeted optimization, and deep learning methods have high requirements for data and computing resources and insufficient generalization ability, resulting in insufficient detection accuracy and robustness, and cannot effectively deal with complex scenarios and resource-constrained environments, a forged speech detection method for feature reconstruction is provided. This method reconstructs the original spectral features extracted by the front-end feature extraction module through the front-end feature generation module, amplifies the specific features of forged speech, enhances the difference between true and false speech features, and improves the detection ability of the forged speech detection model; at the same time, the weighted calculation of the confidence score output by the classifier in the back-end classification module and the confidence score converted from the cosine similarity between the original encoder features and the reconstructed encoder features further improves the detection accuracy of the forged speech detection model.

[0014] The technical solution adopted by the present invention to solve the above technical problem is: a method for detecting forged speech by reconstructing and generating features, characterized by comprising the following steps:

[0015] Step 1: Select a speech dataset and divide the speech dataset into a training set and a test set, wherein each of the training set and the test set contains several real speech samples and several fake speech samples;

[0016] Step 2: Select a front-end and back-end separation pipeline model, which includes a front-end feature extraction module and a back-end classification module. The back-end classification module consists of an encoder and a classifier. The front-end feature extraction module is used to receive a speech sample and extract the original spectrum features corresponding to the speech sample. The encoder is used to receive the original spectrum features corresponding to the speech sample and generate the original encoder features corresponding to the speech sample. The classifier is used to receive the original encoder features corresponding to the speech sample and obtain the authenticity detection result corresponding to the speech sample, wherein the real speech sample and the forged speech sample are both speech samples; then the back-end classification module is trained using the original spectrum features extracted by all the speech samples in the training set after passing through the front-end feature extraction module to obtain a trained back-end classification module;

[0017] Step 3: Use the front-end feature extraction module and the trained back-end classification module in the front-end and back-end separation pipeline model to build a fake speech detection model, which adds a front-end feature generation module on the basis of the front-end feature extraction module and the trained back-end classification module; first, input the speech sample into the front-end feature extraction module to obtain the original spectrum features corresponding to the speech sample, wherein both the real speech sample and the fake speech sample are speech samples; then input the original spectrum features into the front-end feature generation module, reconstruct the original spectrum features, and generate the reconstructed spectrum features corresponding to the speech sample; then input the original spectrum features and the reconstructed spectrum features into the trained back-end classification module together. In the end classification module, the original spectrum features are passed through the encoder in the trained back-end classification module to obtain the original encoder features, and the reconstructed spectrum features are passed through the encoder in the trained back-end classification module to obtain the reconstructed encoder features; then, the original encoder features and the reconstructed encoder features are weightedly fused to obtain the fused features, and the fused features are passed through the classifier in the trained back-end classification module to obtain the first confidence score; then, the cosine similarity of the original encoder features and the reconstructed encoder features is calculated, and the cosine similarity is converted into a second confidence score; finally, the first confidence score and the second confidence score are weightedly calculated to obtain the confidence of speech authenticity detection;

[0018] Step 4: After the front-end feature extraction module in the forged speech detection model extracts the original spectrum features corresponding to all speech samples in the training set, the front-end feature generation module in the forged speech detection model is trained using these original spectrum features. After the front-end feature generation module generates the reconstructed spectrum features corresponding to all speech samples in the training set, the total loss is calculated. After the training is completed, the trained front-end feature generation module is obtained.

[0019] Step 5: Input each speech sample in the test set into the forged speech detection model including the trained front-end feature generation module and the trained back-end classification module to obtain the speech authenticity detection confidence of each speech sample in the test set, which is the authenticity detection result.

[0020] In step 2 and step 3, the front-end feature extraction module uses a manual feature extraction method to obtain the original spectrum features corresponding to the speech sample.

[0021] In step 3, the front-end feature generation module is composed of a VQVAE generator, and the VQVAE generator includes an encoder, a code table and a decoder.

[0022] In step 3, when the original encoder features and the reconstructed encoder features are weightedly fused, the sum of the weight parameters of the original encoder features and the weight parameters of the reconstructed encoder features is 1, and the weight parameters of the original encoder features are greater than the weight parameters of the reconstructed encoder features; when the first confidence score and the second confidence score are weightedly calculated, the sum of the weight parameters of the first confidence score and the weight parameters of the second confidence score is 1.

[0023] The specific process of step 4 is as follows:

[0024] Step 4.1: Import the training parameters of the trained backend classification module;

[0025] Step 4.2: Randomly divide all speech samples in the training set into multiple batches, so that each batch contains batchsize speech samples;

[0026] Step 4.3: Take one of the batches of the training set, and use all the speech samples in this batch as the input of the forged speech detection model. Input the original spectrum features corresponding to all the speech samples in this batch extracted by the front-end feature extraction module into the front-end feature generation module. The front-end feature generation module generates the reconstructed spectrum features corresponding to all the speech samples in this batch. The original spectrum features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the original encoder features. At the same time, the reconstructed spectrum features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the reconstructed encoder features.

[0027] Step 4.4: For all speech samples in this batch, the quantization loss L is calculated to constrain the discretized vector selected from the code table of the front-end feature generation module, i.e., the VQVAE generator, to be close to the features generated by the encoder in the VQVAE generator.quantize , a commitment loss L used to constrain the features generated by the encoder in the VQVAE generator to be close to the discretized vector selected in the code table of the VQVAE generator commit , the real speech consistency loss L used to ensure that the original encoder features corresponding to the real speech sample are as close as possible to the reconstructed encoder features to improve the fidelity of the VQVAE generator to the real speech sample TSC , the real and fake speech discrimination loss L used to maximize the difference between the reconstructed encoder features corresponding to the real speech sample and the reconstructed encoder features corresponding to the fake speech sample TFD , the forged speech distance loss L is used to control the original spectrum features corresponding to the forged speech sample and the reconstructed spectrum features to maintain an appropriate distance FSD , a triplet loss L used to ensure that the original spectral features corresponding to the real speech sample are closer to the reconstructed spectral features and farther away from the original spectral features corresponding to the forged speech sample triplet , the real speech constraint loss L used to strengthen the similarity between the original spectral features corresponding to the real speech sample and the reconstructed spectral features positive , and the total loss L total To constrain the VQVAE generator, where L total =L quantize +L commit +L tSC +L TFd +L FSD +L triplet +L positive ;

[0028] Step 4.5: For all speech samples in this batch, the total loss L total After the calculation, the Adam optimizer with the learning rate setting is used to train the parameters of the front-end feature generation module to complete the training of the front-end feature generation module for this batch;

[0029] Step 4.6: Repeat the process from step 4.3 to step 4.5 until all batches of the training set have undergone a round of training for the front-end feature generation module;

[0030] Step 4.7: Repeat the process from step 4.2 to step 4.6 for a total of number of training rounds, and finally obtain the trained front-end feature generation module.

[0031] In step 4.4,

[0032] Where M represents the number of speech samples in this batch, 1≤m≤M, ‖·‖ 2 is the l2 norm operator, xm represents the original spectrum feature corresponding to the mth speech sample in this batch, z e (x m ) means to convert x m The features generated after passing through the encoder in the VQVAE generator, sg[·] indicates the stop gradient operation, e m represents the code selected from the VQVAE generator and x m The corresponding discretized vector, β represents the weight hyperparameter of the commitment loss, N represents the number of real speech samples in this batch, 1≤i≤N, cos(·) represents the cosine similarity, It represents the original encoder feature obtained by passing the original spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module. represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module, K represents the number of fake speech samples in this batch, 1≤j≤K, It represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module. It represents the original encoder features obtained by passing the original spectrum features corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module.

[0033] In step 4.4, Among them, max(·,·) means taking the maximum value, d(·,·) means finding the Euclidean distance, margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the original encoder features corresponding to the fake speech sample, and positive_margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the reconstructed encoder features.

[0034] In the step 1, a Rawboost data enhancement operation is performed on each speech sample in the training set and the test set, and the Rawboost data enhancement operation includes noise superposition, signal mixing, frequency perturbation, time shift and dynamic range compression.

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] 1) The method of the present invention proposes a new method for optimizing manually extracted features. When the original spectrum features are difficult to effectively amplify the difference between true and false speech, the method of the present invention reconstructs the original spectrum features through the front-end feature generation module (VQVAE generator), which can amplify the specific features of the forged speech and enhance the difference between the true and false speech features. This makes it easier for the classifier in the trained back-end classification module to grasp the key features for distinguishing true from false, thereby significantly improving the detection capability of the forged speech detection model.

[0037] 2) The original encoder features and the reconstructed encoder features are weightedly fused, and the fused features are passed through the classifier in the trained back-end classification module to obtain a first confidence score. At the same time, the cosine similarity between the original encoder features and the reconstructed encoder features is calculated, and the cosine similarity is converted into a second confidence score. The first confidence score and the second confidence score are weightedly calculated to obtain the confidence of speech authenticity detection, which further improves the detection accuracy of the forged speech detection model.

[0038] 3) Use multiple loss functions (such as quantization loss, commitment loss, and real speech consistency loss) to constrain the VQVAE generator to ensure that the VQVAE generator can stably learn robust feature representations during training. This enables the VQVAE generator to better amplify the abnormal features of forged speech, thereby enhancing the forged speech detection model's ability to detect forged speech. At the same time, this stable feature representation helps to improve the performance stability of the forged speech detection model under different test conditions and reduce performance fluctuations caused by data changes or noise interference.

[0039] 4) Improve the robustness and generalization ability of the forged speech detection model to diverse forged speech through data enhancement operations (such as Rawboost). This helps the forged speech detection model maintain good performance in multi-speaker scenarios, complex background sound environments, and cross-domain scenarios, and reduces overfitting.

[0040] 5) The method of the present invention reduces the reliance on a large amount of annotated data and powerful computing resources by optimizing the feature extraction process, thereby improving the training efficiency of the forged speech detection model. This makes the forged speech detection model more suitable for application in resource-constrained scenarios, such as portable devices or real-time processing systems.

[0041] 6) The method of the present invention combines the efficiency of traditional methods and the strong feature learning ability of deep learning methods by fusing manual feature extraction methods (such as LFCC) and VQVAE generator reconstruction features, thereby improving the overall detection performance and the adaptability of the forged voice detection model to different forgery technologies. This fusion method enables the forged voice detection model to better cope with various complex forged voice scenarios.

[0042] 7) The trained backend classification module is used to supervise the VQVAE generator, which significantly improves the learning effect of the VQVAE generator. Through supervised training, the VQVAE generator can better adapt to the forged speech detection task, thereby improving the overall performance of the forged speech detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a schematic diagram of the model structure of the front-end and back-end separated pipeline model;

[0044] Figure 2 Schematic diagram of the model structure of the end-to-end model;

[0045] Figure 3 Schematic diagram of the implementation process of the forged speech detection model;

[0046] Figure 4 A schematic diagram of the cosine similarity distribution percentage between the reconstructed spectral features generated by the VQVAE generator and the original spectral features extracted by the LFCC extraction method on the standard fake speech test dataset ASVspoof2021;

[0047] Figure 5 This is a schematic diagram of the cosine similarity distribution percentage between the reconstructed encoder features and the original encoder features in the method of the present invention on the standard fake speech test dataset ASVspoof2021. DETAILED DESCRIPTION

[0048] The present invention is further described in detail below with reference to the accompanying drawings.

[0049] The present invention proposes a method for detecting forged speech by reconstructing and generating features, which comprises the following steps:

[0050] Step 1: Select a speech dataset and divide the speech dataset into a training set and a test set, wherein the training set and the test set each contain several real speech samples and several forged speech samples.

[0051] The preferred solution is that each speech sample in the training set and the test set performs a Rawboost data enhancement operation, and the Rawboost data enhancement operation includes noise superposition, signal mixing, frequency perturbation, time shift and dynamic range compression.

[0052] Step 2: Select a front-end and back-end separation pipeline model, which includes a front-end feature extraction module and a back-end classification module. The back-end classification module consists of an encoder and a classifier. The front-end feature extraction module is used to receive speech samples and extract original spectrum features corresponding to the speech samples. The encoder is used to receive the original spectrum features corresponding to the speech samples and generate original encoder features corresponding to the speech samples. The classifier is used to receive the original encoder features corresponding to the speech samples and obtain the authenticity detection results corresponding to the speech samples, wherein the real speech samples and the forged speech samples are both speech samples; then the back-end classification module is trained using the original spectrum features extracted by all the speech samples in the training set after passing through the front-end feature extraction module to obtain a trained back-end classification module.

[0053] Here, the front-end and back-end separation pipeline model directly adopts existing technologies, such as Figure 1 As shown; the existing technology is used to train the back-end classification module in the front-end and back-end separation pipeline model.

[0054] Specifically, in step 2, the front-end feature extraction module uses a manual feature extraction method to obtain the original spectrum features corresponding to the speech sample.

[0055] Step 3: Use the front-end feature extraction module and the trained back-end classification module in the front-end and back-end separation pipeline model to build a fake speech detection model, which adds a front-end feature generation module on the basis of the front-end feature extraction module and the trained back-end classification module; Figure 3 As shown, first, the speech sample is input into the front-end feature extraction module to obtain the original spectrum feature corresponding to the speech sample, wherein the real speech sample and the forged speech sample are both speech samples; then the original spectrum feature is input into the front-end feature generation module, the original spectrum feature is reconstructed, and the reconstructed spectrum feature corresponding to the speech sample is generated, and the reconstructed spectrum feature can highlight the difference between the true and false speech; then the original spectrum feature and the generated reconstructed spectrum feature are input into the trained back-end classification module together, the original spectrum feature is encoded by the encoder in the trained back-end classification module to obtain the original encoder feature, and the reconstructed spectrum feature is encoded by the encoder in the trained back-end classification module to obtain the reconstructed encoder feature; then the original encoder feature and the reconstructed encoder feature are weightedly fused to obtain the fused feature, and the fused feature is passed through the classifier in the trained back-end classification module to obtain the first confidence score; then the cosine similarity of the original encoder feature and the reconstructed encoder feature is calculated, and the cosine similarity is converted into a second confidence score; finally, the first confidence score and the second confidence score are weightedly calculated to obtain the confidence of speech authenticity detection.

[0056] Specifically, in step 3, the front-end feature generation module is composed of a VQVAE generator, which includes an encoder, a code table and a decoder; the encoder maps the original spectral features corresponding to the speech sample to a continuous latent space to generate a continuous latent vector; the code table maps the continuous latent vector to a discrete code table vector through a quantization operation; the decoder restores the discrete code table vector to the reconstructed spectral features, that is, obtains the reconstructed spectral features.

[0057] In order to construct a generator that can amplify the specific features of forged speech in the original speech, the present invention uses a VQVAE generator to reconstruct the original spectral features corresponding to the speech sample. The VQVAE generator maps the original spectral features to a discrete latent vector space through an encoder, extracts representative features through a quantization operation, and finally restores the spectrum by a decoder, that is, obtains the reconstructed spectral features. In this process, the VQVAE generator can capture high-order features in the spectrum and amplify the abnormal characteristics of the forged speech, thereby significantly improving the recognizability of the forged speech. Compared with the traditional manual feature extraction method, the VQVAE generator uses a data-driven approach to more comprehensively mine the specific information in the spectrum during the feature extraction and reconstruction stages, significantly improving the accuracy and robustness of forged speech detection. In addition, the quantization characteristics of the VQVAE generator give the model strong stability, avoiding the common instability problems of the generator during training, and at the same time has unique advantages in capturing complex speech patterns.

[0058] Compared with other generators (such as GAN or VAE), the characteristic of the VQVAE generator lies in the design of its discrete latent vector space. The VAE generator uses a continuous latent space when generating, which easily leads to fuzzy features during reconstruction, which confuses the real and fake features of the speech and makes it difficult to effectively amplify the specificity of the fake speech; although the GAN generator can generate clearer features, its training process is prone to instability problems, such as mode collapse or convergence difficulties. The VQVAE generator, by introducing a discretized quantization process, can not only retain the detailed features of the speech spectrum, but also effectively stabilize the training process of the generator. Therefore, choosing the VQVAE generator can not only take into account the stability and performance of the model, but also better highlight the abnormal features of the fake speech through its discretization characteristics, providing a more reliable solution for fake speech detection.

[0059] Specifically, in step 3, when the original encoder feature and the reconstructed encoder feature are weightedly fused, the sum of the weight parameter of the original encoder feature and the weight parameter of the reconstructed encoder feature is 1, and the weight parameter of the original encoder feature is greater than the weight parameter of the reconstructed encoder feature; when the first confidence score and the second confidence score are weighted, the sum of the weight parameter of the first confidence score and the weight parameter of the second confidence score is 1. In this embodiment, the weight parameter of the original encoder feature is 0.7, and the weight parameter of the reconstructed encoder feature is 0.3; the weight parameter of the first confidence score and the weight parameter of the second confidence score are both 0.5.

[0060] The purpose of weighted fusion of original encoder features and reconstructed encoder features is to balance the advantages of original encoder features and reconstructed encoder features and improve the robustness of the forged speech detection model. Reconstructed encoder features have significant advantages in amplifying the difference between forged speech, but due to the slight errors that may be introduced in the generation process, relying solely on reconstructed encoder features may affect the classification performance. By fusing the original encoder features with the reconstructed encoder features, the forged speech detection model can amplify the difference between true and false speech features while using the stability of the original encoder features to offset potential generation errors, thereby ensuring that the fused features input to the classifier are more accurate and complete. This fusion method can effectively enhance the adaptability to different speech features and improve the reliability of classification results.

[0061] In addition to directly inputting the reconstructed spectral features and original spectral features generated by the VQVAE generator into the back-end classification module to obtain the first confidence score, the present invention further explores the potential of the VQVAE generator. By analyzing the differences in the generation effects of the VQVAE generator on real speech and fake speech, a confidence score fusion method based on cosine similarity is designed. Specifically, since the reconstructed spectral features generated by the VQVAE generator for real speech samples have a high similarity with the original spectral features, while the reconstructed spectral features generated for fake speech samples show a low similarity, the cosine similarity between the reconstructed spectral features and the original spectral features is calculated, converted into a set of independent confidence scores, namely the second confidence scores, and fused with the first confidence scores output by the classifier. This design can maximize the use of the discriminative ability of the VQVAE generator and improve the accuracy of classification from different angles.

[0062] By calculating the cosine similarity between the original encoder features and the reconstructed encoder features and converting it into a second confidence score, and fusing it with the first confidence score output by the classifier, the discrimination ability of the classifier is further enhanced. This design makes full use of the sensitivity of the VQVAE generator to the difference between real and fake speech generation, making cosine similarity an effective auxiliary discrimination signal. The fused speech authenticity detection confidence not only integrates the classifier's direct discrimination ability for features, but also introduces additional information that the VQVAE generator amplifies the difference between real and fake features, thereby significantly improving the accuracy of speech authenticity detection. Through this multi-signal fusion approach, the fake speech detection model achieves stronger classification robustness and accurate detection of fake speech.

[0063] Step 4: After using the front-end feature extraction module in the forged speech detection model to extract the original spectrum features corresponding to all the speech samples in the training set, use these original spectrum features to train the front-end feature generation module in the forged speech detection model. After the front-end feature generation module generates the reconstructed spectrum features corresponding to all the speech samples in the training set, calculate the total loss; after the training is completed, a trained front-end feature generation module is obtained.

[0064] In this embodiment, the specific process of step 4 is:

[0065] Step 4.1: Import the training parameters of the trained backend classification module.

[0066] Step 4.2: Randomly divide all speech samples in the training set into multiple batches, so that each batch contains batchsize speech samples.

[0067] Step 4.3: Take one of the batches of the training set, and use all the speech samples in this batch as the input of the forged speech detection model. Input the original spectral features corresponding to all the speech samples in this batch extracted by the front-end feature extraction module into the front-end feature generation module. The front-end feature generation module generates the reconstructed spectral features corresponding to all the speech samples in this batch. The original spectral features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the original encoder features. At the same time, the reconstructed spectral features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the reconstructed encoder features.

[0068] Step 4.4: For all speech samples in this batch, the quantization loss L is calculated to constrain the discretized vector selected from the code table of the front-end feature generation module, i.e., the VQVAE generator, to be close to the features generated by the encoder in the VQVAE generator. quantize , a commitment loss L used to constrain the features generated by the encoder in the VQVAE generator to be close to the discretized vector selected in the code table of the VQVAE generator commit , the real speech consistency loss L used to ensure that the original encoder features corresponding to the real speech sample are as close as possible to the reconstructed encoder features to improve the fidelity of the VQVAE generator to the real speech sample TSC (True Speech Consistency Loss), the real and fake speech discrimination loss L used to maximize the difference between the reconstructed encoder features corresponding to the real speech sample and the reconstructed encoder features corresponding to the fake speech sample TFD True-Fake Discrimination Loss, a fake speech distance loss L used to control the original spectral features corresponding to the fake speech sample to maintain an appropriate distance from the reconstructed spectral features FSD (Fake Speech Divergence Loss), a triplet loss L used to ensure that the original spectral features corresponding to the real speech sample are closer to the reconstructed spectral features and farther away from the original spectral features corresponding to the fake speech sample triplet , the real speech constraint loss L used to strengthen the similarity between the original spectral features corresponding to the real speech sample and the reconstructed spectral features positive , and the total loss L total To constrain the VQVAE generator, where L total =L quantize +L commit +L TSC +L TFD +L FSD +L triplet +L positive .

[0069] In order to enable the VQVAE generator to successfully learn the features related to distinguishing true from false speech, the present invention uses multiple loss functions for back propagation to train the forged speech detection model so that it can learn robust features for distinguishing true from false speech.

[0070] Here, Among them, quantization loss and commitment loss are the basic loss functions in the VQVAE generator training process. The quantization loss minimizes e m With z e (x m), thereby ensuring that the code table can effectively represent the features generated by the encoder, and the commitment loss can ensure that the features generated by the encoder are adapted to the code table. This two-way constraint mechanism works together to make the VQVAE generator perform better in capturing high-order features and abnormal characteristics of speech, providing strong feature support for the detection task of forged speech. M represents the number of speech samples in this batch, 1≤m≤M, ‖·‖ 2 is the l2 norm operator, x m represents the original spectrum feature corresponding to the mth speech sample in this batch, z e (x m ) means to convert x m The features generated after passing through the encoder in the VQVAE generator, sg[·] means stopping the gradient operation and not updating it when calculating the gradient, e m represents the code selected from the VQVAE generator and x m The corresponding discretized vector, β represents the weight hyperparameter of the commitment loss, which is usually taken as 0.25, N represents the number of real speech samples in this batch, 1≤i≤N, cos(·) represents the cosine similarity, It represents the original encoder feature obtained by passing the original spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module. represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module, K represents the number of fake speech samples in this batch, 1≤j≤K, It represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module. It represents the original encoder features obtained by passing the original spectrum features corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module.

[0071] In order to further enable the code table to learn robust features related to the speech authenticity detection task, the present invention draws on the ideas of knowledge distillation and adversarial training, uses the back-end classification module trained in step 2 as a fixed supervision model, designs additional loss terms through the supervision signal provided by it, constrains the behavior of the encoder, code table and decoder in the VQVAE generator, and introduces real speech consistency loss.

[0072] In order to enhance the VQVAE generator's ability to distinguish between real and fake speech, and to maximize the difference between the reconstructed encoder features corresponding to real speech samples and the reconstructed encoder features corresponding to fake speech samples, a loss for distinguishing between real and fake speech is introduced. This loss enhances the fake speech detection model's ability to detect fake speech. By focusing on the core feature differences between real and fake speech, the VQVAE generator can deeply learn and amplify the intrinsic differences between real and fake features while reconstructing the features that the supervised model focuses on. Without this loss, the VQVAE generator may try to learn all the features of fake speech, which will make the VQVAE generator generate fake speech too realistically during reconstruction, covering up the abnormal features of fake speech. This constraint mechanism effectively avoids the VQVAE generator from only copying the common features of real speech, but instead accurately captures the subtle anomalies and potential traces of fake speech in fake speech, thereby improving the distinguishing power of feature distribution and providing more robust and task-related feature support for the supervised model.

[0073] In order to prevent the reconstructed spectral features corresponding to the forged speech samples from completely covering the corresponding original spectral features, the forged speech distance loss is introduced to control the original spectral features corresponding to the forged speech samples to maintain an appropriate distance from the reconstructed spectral features. This design does not simply pursue high similarity, but encourages the VQVAE generator to retain a certain characteristic difference when generating the reconstructed spectral features. If the generated reconstructed spectral features are too close to the original spectral features, it may cause the VQVAE generator to learn details that are irrelevant to the forged speech, thereby weakening its ability to extract forged traces. Through the above-mentioned loss function, the VQVAE generator can learn features with a certain deviation for the forged speech in the feature space. This deviation is intended to amplify the potential anomalies in the forged speech (such as the unnatural features of the VQVAE generator in the forged speech), providing a stronger discriminant signal for the back-end classification module. The introduction of this loss essentially adds a kind of "adversarial constraint" to the generated reconstructed spectral features. In order to reduce this loss, the VQVAE generator must generate feature representations that are consistent with the forged speech distribution and different from the original spectral features in the feature space. This constraint makes the VQVAE generator mainly learn the unique features of forged speech, rather than the common features that exist in both real and forged speech, making it difficult for the reconstructed spectral features of forged speech to hide the anomalies in forged speech, thereby further improving the distinguishability of forged speech. Ultimately, this design enables the VQVAE generator to amplify the key differences from real speech while retaining the overall distribution of forged speech, providing the supervision model with more significant forgery trace features and improving the overall detection capability.

[0074] Here, Among them, max(·,·) means taking the maximum value, d(·,·) means finding the Euclidean distance, margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the original encoder features corresponding to the fake speech sample. In this embodiment, margin is set to 10, and positive_margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the reconstructed encoder features. In this embodiment, positive_margin is set to 5.

[0075] The above-mentioned losses related to the VQVAE generator are mainly used to supervise the reconstruction quality of the generated speech and learn the difference between true and false speech. By constraining the distribution of the feature space of the VQVAE generator, it ensures that the generated reconstructed spectral features are consistent with the original spectral features, while amplifying the distinguishing characteristics of true and false speech. The triplet loss further supervises the training of the forged speech detection model from the perspective of discrimination, emphasizing the relative position relationship between the real speech, generated speech and forged speech in the feature space, so that the forged speech detection model can more accurately learn the essential difference between true and false speech. The combination of the two not only improves the speech reconstruction ability of the VQVAE generator, but also provides a more robust feature representation for the forged speech detection model, effectively assisting the generation task, while strengthening the detection ability of forged speech.

[0076] The goal of the triplet loss is to ensure that the original spectral features corresponding to the real speech sample are closer to the reconstructed spectral features, while being away from the original spectral features corresponding to the fake speech sample. In addition to the triplet loss, in order to further strengthen the similarity between the original spectral features corresponding to the real speech sample and the reconstructed spectral features, the real speech constraint loss is added.

[0077] The losses associated with the VQVAE generator focus mainly on the quality of the generated speech and its performance in the feature space. The main goal is to retain the high-quality features of the real speech while enhancing the discriminability of the fake speech features. The triplet loss focuses more on the relative positional relationship between the real speech, generated speech, and fake speech in the model feature space, and improves the distinguishability between real and fake speech through clear geometric constraints. One focuses on the quality of speech generation and the amplification of abnormal characteristics, and the other focuses on the distribution optimization of the feature space. The two complement each other and jointly improve the performance of the VQVAE generator in generating high-quality speech and enhancing the ability to distinguish between real and fake speech.

[0078] Step 4.5: For all speech samples in this batch, the total loss L total After the calculation, the Adam optimizer with the learning rate setting is used to train the parameters of the front-end feature generation module to complete the training of the front-end feature generation module for this batch.

[0079] Step 4.6: Repeat the process from step 4.3 to step 4.5 until all batches of the training set have undergone a round of training for the front-end feature generation module.

[0080] Step 4.7: Repeat the process from step 4.2 to step 4.6, and train for a total of Number rounds to finally obtain a trained front-end feature generation module. In this embodiment, Number is set to 100.

[0081] Step 5: Input each speech sample in the test set into the forged speech detection model including the trained front-end feature generation module and the trained back-end classification module, and obtain the speech authenticity detection confidence of each speech sample in the test set, which is the authenticity detection result. Specifically, input each speech sample in the test set into the forged speech detection model, obtain the original spectrum feature corresponding to each speech sample through the front-end feature extraction module, reconstruct the original spectrum feature corresponding to each speech sample through the trained front-end feature generation module to generate the corresponding reconstructed spectrum feature, the original spectrum feature and the reconstructed spectrum feature corresponding to each speech sample are passed through the encoder in the trained back-end classification module together to obtain the original encoder feature and the reconstructed encoder feature, the fusion feature obtained by weighted fusion of the original encoder feature and the reconstructed encoder feature corresponding to each speech sample is passed through the classifier in the trained back-end classification module to obtain the first confidence score, and the second confidence score obtained by cosine similarity conversion of the original encoder feature and the reconstructed encoder feature corresponding to each speech sample is weighted with the first confidence score to obtain the speech authenticity detection confidence, which is the authenticity detection result. In this embodiment, if the voice authenticity detection confidence is greater than or equal to 0 and less than or equal to 0.5, the voice sample is detected as forged voice; otherwise, the voice sample is detected as real voice.

[0082] In order to further verify the feasibility and effectiveness of the method of the present invention, experiments were conducted on the method of the present invention.

[0083] In the experiment, the ASVspoof2019 dataset, which is widely used in the field of language forgery forensics, is used as a training set, and performance tests are performed on the ASVspoof2019 and ASVspoof2021 datasets. The ASVspoof2019 and ASVspoof2021 datasets are used as test sets to verify the effectiveness and generalization ability of the forged speech detection model.

[0084] The ASVspoof2019 dataset is a standard dataset provided by the Automatic Voice Verification Anti-Fraud Challenge, focusing on the problem of voice forgery detection, including two forms of fraud: voice conversion (VC) and speech synthesis (TTS).

[0085] The LA (Logical Access) task in the ASVspoof2019 dataset mainly targets synthetic and converted voice attacks. It contains 2,580 training audios, 24,800 development audios, and 71,700 test audios, covering a variety of forgery techniques, emphasizing the robustness of the model to diverse forgery forms.

[0086] The ASVspoof2021 dataset is an extended version of the ASVspoof2019 dataset, which further improves the diversity and challenge of the dataset, especially in terms of generalization performance testing in real scenarios. The ASVspoof2021 dataset has added more complex forgery techniques, and places special emphasis on adaptability to cross-language and complex background sound scenarios. The LA task continues to focus on synthetic and converted speech forgery, and the amount of data has increased significantly, including 20,215 training audios, 22,180 development audios, and 62,903 test audios, further enhancing the coverage of complex forgery techniques. In addition, the ASVspoof2021 dataset is closer to real scenarios in task design, aiming to evaluate the performance of the model in multi-language and multi-background sound environments.

[0087] By testing on two versions of the dataset, we can fully verify the robustness and generalization ability of the forged voice detection model in dealing with diverse forged voice scenarios.

[0088] Implementation details:

[0089] In the experiment, all training audio files were resampled to 16kHz. LFCC with shape (60, 750) was used as the original spectrogram feature, and Resnet18 was used as the classifier in the back-end classification module. In terms of experimental parameters, the margin in the triple loss was set to 10 and the positive_margin was set to 5 according to experience. In the weighted fusion stage, the original encoder features and the reconstructed encoder features generated by the encoder in the back-end classification module were weighted and fused, and a weighting strategy of 0.7 for the original encoder features and 0.3 for the reconstructed encoder features was adopted. Such a weight parameter allocation can not only retain the core features in the original encoder features and ensure the integrity of the basic information, but also highlight the amplified forged features in the reconstructed encoder features, thereby enhancing the sensitivity of the model to forged features. This weighted fusion method not only improves the robustness of the features, but also significantly enhances the specificity of the model in distinguishing real speech from forged speech. In addition, weighted fusion can balance the feature distribution to a certain extent, alleviate the problem of insufficient generalization ability caused by the excessive dominance of a single feature, and thus improve the performance of the model in complex scenarios and diversified data.

[0090] In order to make the reconstructed spectral features generated by the VQVAE generator more robust, the present invention performs Rawboost data enhancement operations on the data to improve the robustness and generalization ability of the model to diversified forged speech. These include noise superposition, signal mixing, frequency perturbation, time offset and dynamic range compression. Noise superposition simulates various noise interference scenarios in real environments by adding background noise to the audio. Signal mixing increases data complexity by mixing multiple voice signals together, helping the model to better learn forged features in complex backgrounds. Frequency perturbation perturbs the spectral characteristics of the voice signal to simulate the frequency distortion problem that may be introduced under different recording conditions. Time offset enhances the model's adaptability to changes in time synchronization by randomly adjusting the time position of the voice signal. Dynamic range compression further approaches the actual recording scene by simulating the dynamic range compression effect of the recording device. Through these strategies, the Rawboost data enhancement operation significantly expands the diversity of the data set, enabling the model to have stronger detection capabilities in environments such as unknown voice communication protocols and channels.

[0091] In the experiment, the equal error rate (EER) and the minimum detection cost function (min-tDCF) were used as evaluation indicators of model performance. The equal error rate (EER) measures the model's ability to classify real and fake speech by balancing the false rejection rate (FRR) and the false acceptance rate (FAR). The error rate when FRR and FAR are equal is the EER. The lower the value, the better the classification performance of the model. The minimum detection cost function (min-tDCF) is used to evaluate the comprehensive performance of the forgery detection system working in conjunction with the automatic speech verification (ASV) system. The min-tDCF takes into account the rejection errors, acceptance errors of the ASV system, and the misclassification of the forgery detection system. By setting the cost weights of different errors, it can reflect the economy and practicality of the system in real application scenarios. The combination of EER and min-tDCF can comprehensively evaluate the detection performance and practical application capabilities of the model, thereby providing a more scientific and accurate performance measurement standard for the forgery speech detection task.

[0092] Experimental results:

[0093] The following is the experimental results of the proposed method and its comparative analysis with other methods. Table 1 shows the comparison of the baseline performance of the proposed method and other methods on the ASVspoof 2019 and ASVspoof2021 datasets.

[0094] Table 1 Comparison of baseline performance of the proposed method and other methods on the ASVspoof 2019 and ASVspoof2021 datasets

[0095]

[0096] Table 1 lists the EER and min-tDCF indicators of the current mainstream and effective end-to-end models and front-end and back-end separated pipeline models. RawNet 2 in Table 1 was proposed in the ASVSpoof2021 competition. It directly processes the original speech waveform to achieve an end-to-end forgery detection baseline model; AASIST (Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention) is a model that integrates time-frequency analysis and graph neural network (GNN). This method is used to model the spatiotemporal correlation of speech signals. This model is an evolved version of the champion solution in the ASVSpoof2021 competition; Rawformer is a model that combines the waveform processing capabilities of RawNet and the global modeling capabilities of Transformer. This model is an improved solution for RawNet; LFCC+GMM and CQCC+GMM are the pipeline model baseline systems proposed in the ASVSpoof2021 competition; Wav2Vec2+Transformer is a model that uses the pre-trained model Wav2Vec2 and the Transformer model to improve the forgery detection performance; VQ, LFCC, and Resnet18 in VQ-LFCC+Resnet18 refer to the reconstructed spectral features, original spectral features, and classifier Resnet18 generated by the VQVAE generator, respectively.

[0097] From the comparison data in Table 1, it can be found that although the backend of the method of the present invention only uses the relatively simple ResNet18, the detection performance on the ASVspoof2019 dataset is still very good, with an equal error rate (EER) of 1.02. On the more challenging ASVspoof2021 dataset, compared with other methods that rely on deep learning models as the front end and require higher computing resources and longer training time, the method of the present invention still shows significant advantages, with an EER of 5.02, and the min-tDCF index is further better than other methods, which is only 0.3.

[0098] On the standard fake speech test dataset ASVspoof2021, the cosine similarity distribution percentage between the reconstructed spectral features generated by the VQVAE generator and the original spectral features extracted by the LFCC extraction method is as follows: Figure 4As shown in Figure 2, the difference between real speech and fake speech can be effectively amplified by cosine similarity distribution. Figure 4 Real is the cosine similarity distribution percentage between the reconstructed spectral features and the original spectral features corresponding to the real speech sample, and Spoof is the cosine similarity distribution percentage between the reconstructed spectral features and the original spectral features corresponding to the forged speech sample. Figure 4 It can be seen that the distribution of the two is significantly different. The cosine similarity corresponding to the real speech sample is more concentrated in the high segment, while the cosine similarity corresponding to the forged speech sample is obviously distributed in the lower segment. The unique advantage of the method of the present invention is that it integrates the reconstructed spectrum features generated by the VQVAE generator and the original spectrum features extracted by the LFCC extraction method, that is, the reconstructed spectrum features and the original spectrum features are input into the encoder in the back-end classification module, and the encoder outputs the reconstructed encoder features and the original encoder features accordingly, and then the reconstructed encoder features and the original encoder features are weighted fused. This fusion strategy not only retains the stability and interpretability of traditional manual features, but also enhances the amplification effect of forged speech features through VQVAE generator reconstruction, thereby improving the robustness and specificity of forgery detection. Compared with other more complex deep learning models, the training time of the method of the present invention is significantly shortened, the demand for computing resources is lower, and it can efficiently adapt to scenarios with limited resources. At the same time, it still surpasses many mainstream methods in detection performance and shows better comprehensive capabilities. Figure 5 The cosine similarity distribution percentages of the reconstructed encoder features and the original encoder features in the method of the present invention on the standard forged speech test dataset ASVspoof2021 are given. Real is the cosine similarity distribution percentage of the reconstructed encoder features and the original encoder features corresponding to the real speech sample, and Spoof is the cosine similarity distribution percentage of the reconstructed encoder features and the original encoder features corresponding to the forged speech sample. Figure 5 It can be seen that the cosine similarity corresponding to most real speech samples reaches 1, and the cosine similarity corresponding to the rest of the real speech samples is mostly maintained above 0.8, while the cosine similarity distribution corresponding to the fake speech samples is more discrete and low. This distribution characteristic further illustrates the VQVAE generator's ability to distinguish between real and fake speech.

[0099] from Figure 4 and Figure 5It can also be seen that after using a series of loss functions to optimize the fake speech detection model in a targeted manner, the VQVAE generator can reconstruct the original spectral features corresponding to the real speech sample excellently, and can accurately restore the detailed features of the real speech. The reconstruction effect of the original spectral features corresponding to the fake speech sample is poor, which can well amplify the feature differences related to fake in the fake speech. In addition, the cosine similarity of the original encoder features and the reconstructed encoder features is converted into a second confidence score, which can well distinguish the gap between the real speech and the fake speech, and provide a reliable basis for distinguishing the two. Through the introduction and fusion of cosine similarity (second confidence score), the present invention has important advantages in the following aspects. First, the fusion of cosine similarity can provide supplementary information for the classifier in the back-end classification module, effectively improving the accuracy and robustness of fake speech classification. Secondly, by using the characteristic differences of the VQVAE generator for the fake speech generation effect, the potential of the VQVAE generator can be further optimized, so that it not only has the reconstruction ability, but also can use cosine similarity to detect abnormal features in fake speech. Finally, this method provides an independent source of discrimination signals for the fake speech detection model without significantly increasing the complexity of the model, which significantly improves the performance of the fake speech detection model. The confidence score fusion method based on cosine similarity fully exploits the capabilities of the VQVAE generator, enabling it to not only reconstruct spectral features in the fake speech detection task, but also significantly improve the overall classification performance and accuracy by amplifying the fake speech features.

[0100] In order to verify the effectiveness of data enhancement and the improvement effect of the method of the present invention on the original model, Table 2 lists the performance comparison of the method of the present invention and other methods before and after using Rawboost data enhancement on the ASVspoof2021 dataset.

[0101] Table 2 Performance comparison of the method of the present invention and other methods before and after using Rawboost data enhancement on the ASVspoof2021 dataset

[0102]

[0103] In Table 2, LFCC+ResNet18 refers to the method of using only the original spectrum features combined with the classifier ResNet18. As can be seen from Table 2, when Rawboost data enhancement is not used, the EER of LFCC+ResNet18 is 22.64% and the min-tDCF is 0.694. After combining the VQ-LFCC features, its EER and min-tDCF are significantly reduced to 13.37% and 0.476, respectively. This shows that the model performance can be significantly improved without additional data enhancement by fusing the reconstructed spectrum features generated by the VQVAE generator alone.

[0104] Further analysis of the results after using Rawboost data enhancement shows that the performance of LFCC+ResNet18 is significantly improved, with EER reduced to 8.25% and min-tDCF reduced to 0.409. After fusing VQ-LFCC features, the performance is greatly improved again, with EER and min-tDCF reduced to 5.02% and 0.300 respectively. This result verifies the effectiveness of Rawboost data enhancement in improving detection performance, and also shows that the method of the present invention significantly improves the overall performance of the model in complex speech forgery detection tasks while enhancing the ability to distinguish forged features through VQ-LFCC feature fusion.

[0105] In summary, by combining the strategies of data enhancement and VQ-LFCC feature fusion, the performance of the proposed method on the ASVspoof2021 dataset is better than that of the traditional LFCC+ResNet18 model, which fully demonstrates its effectiveness and superiority.

Claims

1. A method for detecting forged speech by reconstructing and generating features, characterized in that The following steps are involved: Step 1: Select a speech dataset and divide the speech dataset into a training set and a test set, wherein each of the training set and the test set contains several real speech samples and several fake speech samples; Step 2: Select a front-end and back-end separation pipeline model, which includes a front-end feature extraction module and a back-end classification module. The back-end classification module consists of an encoder and a classifier. The front-end feature extraction module is used to receive a speech sample and extract the original spectrum features corresponding to the speech sample. The encoder is used to receive the original spectrum features corresponding to the speech sample and generate the original encoder features corresponding to the speech sample. The classifier is used to receive the original encoder features corresponding to the speech sample and obtain the authenticity detection result corresponding to the speech sample, wherein the real speech sample and the forged speech sample are both speech samples; then the back-end classification module is trained using the original spectrum features extracted by all the speech samples in the training set after passing through the front-end feature extraction module to obtain a trained back-end classification module; Step 3: Use the front-end feature extraction module and the trained back-end classification module in the front-end and back-end separation pipeline model to build a fake speech detection model, which adds a front-end feature generation module on the basis of the front-end feature extraction module and the trained back-end classification module; first, input the speech sample into the front-end feature extraction module to obtain the original spectrum features corresponding to the speech sample, wherein both the real speech sample and the fake speech sample are speech samples; then input the original spectrum features into the front-end feature generation module, reconstruct the original spectrum features, and generate the reconstructed spectrum features corresponding to the speech sample; then input the original spectrum features and the reconstructed spectrum features into the trained back-end classification module together. In the end classification module, the original spectrum features are passed through the encoder in the trained back-end classification module to obtain the original encoder features, and the reconstructed spectrum features are passed through the encoder in the trained back-end classification module to obtain the reconstructed encoder features; then, the original encoder features and the reconstructed encoder features are weightedly fused to obtain the fused features, and the fused features are passed through the classifier in the trained back-end classification module to obtain the first confidence score; then, the cosine similarity of the original encoder features and the reconstructed encoder features is calculated, and the cosine similarity is converted into a second confidence score; finally, the first confidence score and the second confidence score are weightedly calculated to obtain the confidence of speech authenticity detection; Step 4: After the front-end feature extraction module in the forged speech detection model extracts the original spectrum features corresponding to all speech samples in the training set, the front-end feature generation module in the forged speech detection model is trained using these original spectrum features. After the front-end feature generation module generates the reconstructed spectrum features corresponding to all speech samples in the training set, the total loss is calculated. After the training is completed, the trained front-end feature generation module is obtained. Step 5: Input each speech sample in the test set into the forged speech detection model including the trained front-end feature generation module and the trained back-end classification module to obtain the speech authenticity detection confidence of each speech sample in the test set, which is the authenticity detection result.

2. A method for detecting forged speech by reconstructing and generating features according to claim 1, characterized in that In step 2 and step 3, the front-end feature extraction module uses a manual feature extraction method to obtain the original spectrum features corresponding to the speech sample.

3. A method for detecting forged speech by reconstructing and generating features according to claim 1, characterized in that In step 3, the front-end feature generation module is composed of a VQVAE generator, and the VQVAE generator includes an encoder, a code table and a decoder.

4. A method for detecting forged speech by reconstructing and generating features according to claim 1, characterized in that In step 3, when the original encoder features and the reconstructed encoder features are weightedly fused, the sum of the weight parameters of the original encoder features and the weight parameters of the reconstructed encoder features is 1, and the weight parameters of the original encoder features are greater than the weight parameters of the reconstructed encoder features; when the first confidence score and the second confidence score are weightedly calculated, the sum of the weight parameters of the first confidence score and the weight parameters of the second confidence score is 1.

5. A method for detecting forged speech by reconstructing and generating features according to claim 3, characterized in that The specific process of step 4 is as follows: Step 4.1: Import the training parameters of the trained backend classification module; Step 4.2: Randomly divide all speech samples in the training set into multiple batches, so that each batch contains batchsize speech samples; Step 4.3: Take one of the batches of the training set, and use all the speech samples in this batch as the input of the forged speech detection model. Input the original spectrum features corresponding to all the speech samples in this batch extracted by the front-end feature extraction module into the front-end feature generation module. The front-end feature generation module generates the reconstructed spectrum features corresponding to all the speech samples in this batch. The original spectrum features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the original encoder features. At the same time, the reconstructed spectrum features corresponding to all the speech samples in this batch are obtained by passing through the encoder in the trained back-end classification module to obtain the reconstructed encoder features. Step 4.4: For all speech samples in this batch, the quantization loss L is calculated to constrain the discretized vector selected from the code table of the front-end feature generation module, i.e., the VQVAE generator, to be close to the features generated by the encoder in the VQVAE generator. quantize , a commitment loss L used to constrain the features generated by the encoder in the VQVAE generator to be close to the discretized vector selected in the code table of the VQVAE generator commit , the real speech consistency loss L used to ensure that the original encoder features corresponding to the real speech sample are as close as possible to the reconstructed encoder features to improve the fidelity of the VQVAE generator to the real speech sample TSC , the real and fake speech discrimination loss L used to maximize the difference between the reconstructed encoder features corresponding to the real speech sample and the reconstructed encoder features corresponding to the fake speech sample TFD , the forged speech distance loss L is used to control the original spectrum features corresponding to the forged speech sample and the reconstructed spectrum features to maintain an appropriate distance FSD , a triplet loss L used to ensure that the original spectral features corresponding to the real speech sample are closer to the reconstructed spectral features and farther away from the original spectral features corresponding to the forged speech sample triplet , the real speech constraint loss L used to strengthen the similarity between the original spectral features corresponding to the real speech sample and the reconstructed spectral features positive , and the total loss L total To constrain the VQVAE generator, where L total =L quantize +L commit +L TSC +L TFD +L FSD +L triplet +L positive ; Step 4.5: For all speech samples in this batch, the total loss L total After the calculation, the Adam optimizer with the learning rate setting is used to train the parameters of the front-end feature generation module to complete the training of the front-end feature generation module for this batch; Step 4.6: Repeat the process from step 4.3 to step 4.5 until all batches of the training set have undergone a round of training for the front-end feature generation module; Step 4.7: Repeat the process from step 4.2 to step 4.6 for a total of number of training rounds, and finally obtain the trained front-end feature generation module.

6. A method for detecting forged speech by reconstructing and generating features according to claim 5, characterized in that In step 4.4, Where M represents the number of speech samples in this batch, 1≤m≤M, ‖·‖2 is the l2 norm operator, and x m represents the original spectrum feature corresponding to the mth speech sample in this batch, z e (x m ) means to convert x m The features generated after passing through the encoder in the VQVAE generator, sg[·] indicates the stop gradient operation, e m represents the code selected from the VQVAE generator and x m The corresponding discretized vector, β represents the weight hyperparameter of the commitment loss, N represents the number of real speech samples in this batch, 1≤i≤N, cos(·) represents the cosine similarity, It represents the original encoder feature obtained by passing the original spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module. represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the i-th real speech sample in this batch through the encoder in the trained backend classification module, K represents the number of fake speech samples in this batch, 1≤j≤K, It represents the reconstructed encoder feature obtained by passing the reconstructed spectrum feature corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module. It represents the original encoder features obtained by passing the original spectrum features corresponding to the j-th forged speech sample in this batch through the encoder in the trained backend classification module.

7. A method for detecting forged speech by reconstructing and generating features according to claim 6, characterized in that In step 4.4, Among them, max(·,·) means taking the maximum value, d(·,·) means finding the Euclidean distance, margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the original encoder features corresponding to the fake speech sample, and positive_margin is a hyperparameter used to set the minimum interval between the original encoder features corresponding to the real speech sample and the reconstructed encoder features.

8. The method for detecting forged speech by reconstructing and generating features according to claim 1, characterized in that In the step 1, a Rawboost data enhancement operation is performed on each speech sample in the training set and the test set, and the Rawboost data enhancement operation includes noise superposition, signal mixing, frequency perturbation, time shift and dynamic range compression.