False Audio Detection Method and System Based on Self-Knowledge Distillation
The self-knowledge distillation method enhances shallow neural networks for fake audio detection by using a deep network as a teacher to guide shallow networks, improving detection accuracy by balancing feature disparities and refining extraction techniques.
Patent Information
- Application Number
- CN202310135374.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-09
AI Technical Summary
In the prior art, shallow networks have insufficient ability to capture features in false audio detection, and the characteristics of deep and shallow networks are significantly different, making it difficult to effectively balance, resulting in low detection accuracy.
The self-knowledge distillation method is used to divide the deep neural network into multiple parts, and the deepest network is used as a teacher model to guide the shallow network, and the feature capture ability of the shallow network is enhanced by calculating the loss function's balanced feature differences.
It significantly improves the accuracy of false audio detection, improves the feature capture capability of shallow networks, and enhances the performance of false audio detection.
Smart Images

Figure CN116312628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fake audio detection, and in particular to a fake audio detection method and system based on self-knowledge distillation. Background Art
[0002] With the rise of biometrics, speaker verification has been widely used. However, synthetic speech detection seriously threatens speaker verification systems. The attacks faced by speaker verification systems mainly include: voice replay, speech synthesis, and voice conversion. At present, these attack technologies are very mature, and the technologies for detecting these fake audio attacks have only been developed in recent years. Therefore, it is urgent for researchers to develop effective fake audio detection systems to detect the deception attacks of fake audio.
[0003] Audio forgery detection technology can effectively improve the performance of anti-deception systems. The current work mainly focuses on two aspects: 1) improving the acoustic features of audio; 2) designing new classification models. Existing research has shown that the shallow features of speech are very important for the task of fake audio detection, such as some spectrogram defects, silent segments, etc. Shallow networks are very sensitive to this information, but the ability of shallow networks to capture features is not as good as that of deep networks. Therefore, how to strengthen shallow networks is a challenging problem. In addition, due to the large gap between shallow networks and deep networks, their feature differences are also very obvious. Therefore, how to balance the feature differences between shallow and deep networks is also a challenging problem. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a fake audio detection method and system based on self-knowledge distillation, which can significantly improve the accuracy of fake audio detection.
[0005] To solve the above technical problem, a technical solution adopted by the present invention is: to provide a fake audio detection method based on self-knowledge distillation, including the following steps:
[0006] S1: Extract the log power spectrum features of the original speech waveform, and use the F0 sub-band of the log power spectrum features as the input features for fake audio detection;
[0007] S2: Use a deep neural network to model the input features, calculate the loss between the label and the network output, and use it to extract the hidden knowledge of the training set;
[0008] S3: Divide the network into n parts, calculate the loss of the prediction outputs of all shallow networks and the deepest network, use the deepest network as the teacher model, add a classifier to the backend of the shallow network to form a student model, and the teacher model guides the student model according to the prediction results to enhance the shallow network;
[0009] S4: In the feature dimension, the teacher model distills knowledge into the student model to balance the feature differences between the shallow and deep networks.
[0010] In a preferred embodiment of the present invention, the specific steps of step S1 include:
[0011] S101: Use the short-time Fourier transform (STFT) to convert the time-domain speech signal into a time-frequency domain speech signal:
[0012] X r [t, f] + i*X i [t, f] = STFT(X[k])(1)
[0013] where x[k] represents the speech signal in the time domain, k is the time index of the speech signal, and are the corresponding real and imaginary parts of the STFT, t is the index of the time frame number, and f is the index of the frequency unit;
[0014] S102: Perform a logarithmic operation on the real and imaginary parts of the STFT to obtain the log power spectrum features:
[0015]
[0016] where log represents the logarithmic operation, and LPS full is the full-band feature of the required log power spectrum;
[0017] S103: Apply the 0 - 400 Hz frequency band of the log power spectrum as the required F0 sub-band:
[0018] LPS F0 = LPS 0-400HZ (3)
[0019] In a preferred embodiment of the present invention, the specific steps of step S2 include:
[0020] Use the A_softmax function to calculate the loss between the label and the network prediction output:
[0021]
[0022] where A_softmax represents the A-softmax function, p n is the prediction output representing the deepest layer of the network, L is the label, is the loss between the label and the prediction output of the deepest layer of the network.
[0023] In a preferred embodiment of the present invention, the specific steps of step S3 include:
[0024] Calculate the loss between the prediction of the deepest layer of the network and the predictions of all shallow networks using the Kullback-Leible divergence function:
[0025]
[0026] where KL represents the Kullback-Leible divergence function, and p i is the prediction output representing the shallow network, is the sum of the prediction losses between all shallow networks and the deepest network.
[0027] In a preferred embodiment of the present invention, in step S4, calculate the loss between all shallow network features and the deepest network features to balance the feature differences between the shallow and deep network features.
[0028] Further, the specific steps of step S4 include:
[0029] S401: Calculate the loss between the deepest network features of the network and all shallow network features using the mean squared error function:
[0030]
[0031] where MSE represents the mean squared error function, is the shallow network feature, is the deepest network feature; is the sum of the feature losses between all shallow networks and the deepest network;
[0032] S402: Balance the three losses by setting hyperparameters:
[0033]
[0034] where α and β are the balance coefficients of the losses respectively, is the final loss.
[0035] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a fake audio detection system based on self-knowledge distillation, including:
[0036] A voice feature input module for extracting the logarithmic power spectrum features of the original voice waveform and using the F0 subband of the logarithmic power spectrum as the input features for fake audio detection;
[0037] A modeling module for modeling the input features obtained by the voice feature input module using a deep neural network;
[0038] A hidden knowledge extraction module for calculating the loss between the label and the network prediction output and obtaining the hidden knowledge of the training set;
[0039] Enhanced shallow network module, which is used to calculate the loss of the prediction outputs of all shallow networks and the deepest network. The deepest network serves as the teacher model to guide all shallow networks so as to enhance the shallow networks.
[0040] Balanced feature difference module, which is used to calculate the loss of the features of all shallow networks and the deepest network, so as to balance the feature differences between the shallow and deep network features.
[0041] In a preferred embodiment of the present invention, the steps for the enhanced shallow network module to enhance the shallow network include:
[0042] First, divide the network into n parts. Take the deepest network as the teacher model, set classifiers after all shallow networks to construct n - 1 student models, and according to the prediction results, the teacher model guides all student models.
[0043] The beneficial effects of the present invention are as follows: The present invention proposes a self - knowledge distillation method for fake audio detection. This method uses the F0 sub - band as the input feature. First, divide the network into n parts, take the deepest network as the teacher model, add classifiers at the back ends of all shallow networks to form student models, and the redundant classifiers can be removed during inference. In addition, calculate the loss between the prediction output of the deepest network and the label to obtain the hidden knowledge of the training set. Then, calculate the loss between the prediction output of the teacher model and the prediction output of the student model to enhance the shallow networks. Next, calculate the loss between the features of the deepest network and the shallow network features to balance the differences between the shallow and deep network features. Finally, balance the three losses through two hyperparameters. The present invention is very helpful for fake audio detection and can significantly improve the accuracy of fake audio detection technology without increasing the model load. Description of the Drawings
[0044] Figure 1 is the flowchart of the fake audio detection method based on self - knowledge distillation of the present invention;
[0045] Figure 2 is the schematic diagram of the network framework in the fake audio detection method and system based on self - knowledge distillation;
[0046] Figure 3 is the structural block diagram of the fake audio detection system based on self - knowledge distillation. Detailed Embodiment
[0047] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0048] Please refer toFigure 1 and Figure 2 , embodiments of the present invention include:
[0049] A method for detecting fake audio based on self-knowledge distillation, comprising the following steps:
[0050] S1: Extract the logarithmic power spectrum features of the original speech waveform, and use the F0 sub-band of the logarithmic power spectrum features as the input features for fake audio detection; the specific steps include:
[0051] S101: Use the short-time Fourier transform STFT to convert the time-domain speech signal into a time-frequency domain speech signal:
[0052] X r [t, f]+i*X i [t, f] = STFT(X[k])(1)
[0053] where x[k] represents the speech signal in the time domain, k is the time index of the speech signal, and are the corresponding real and imaginary parts of STFT, t is the index of the number of time frames, and f is the index of the frequency unit;
[0054] S102: Perform a logarithmic operation on the real and imaginary parts of STFT to obtain the logarithmic power spectrum features:
[0055]
[0056] where log represents the logarithmic operation, and LPS full is the full-band feature of the required logarithmic power spectrum;
[0057] S103: Apply the 0-400 Hz frequency band of the logarithmic power spectrum as the required F0 sub-band:
[0058] LPS F0 = LPS 0-400HZ (3)
[0059] S2: Use a deep neural network to model the input features, calculate the loss between the label and the network output, and use it to extract the hidden knowledge of the training set;
[0060] Specifically, use the A_softmax function to calculate the loss between the label and the prediction output of the deepest layer of the network:
[0061]
[0062] where A_softmax represents the A-softmax function, and p n is the prediction output representing the deepest layer of the network, and L is the label, is the loss of the label and the prediction output of the deepest layer network.
[0063] S3: Divide the network into n parts, calculate the losses of the prediction outputs of all shallow layer networks and the deepest layer network. Take the deepest layer network as the teacher model, add a classifier to the backend of the shallow layer network to form the student model, and the teacher model guides the student model according to the prediction results to enhance the shallow layer network;
[0064] Specifically, use the Kullback-Leible divergence function to calculate the losses between the prediction of the deepest layer of the network and the predictions of all shallow layer networks:
[0065]
[0066] where KL represents the Kullback-Leible divergence function, and p i represents the prediction output of the shallow layer network, is the sum of the prediction losses between all shallow layer networks and the deepest layer network.
[0067] S4: In the feature dimension, the teacher model distills knowledge into the student model to balance the feature differences between the shallow and deep layer networks. Specifically, calculate the losses between the features of all shallow layer networks and the features of the deepest layer network to balance the feature differences between the features of the shallow and deep layer networks. The specific steps include:
[0068] S401: Use the mean squared error function to calculate the losses between the features of the deepest layer network of the network and the features of all shallow layer networks:
[0069]
[0070] where MSE represents the mean squared error function, represents the features of the shallow layer network, are the features of the deepest layer network; is the sum of the feature losses between all shallow layer networks and the deepest layer network;
[0071] S402: Balance the three losses loss by setting hyperparameters:
[0072]
[0073] where α and β are the balance coefficients of the losses respectively, is the final loss.
[0074] It should be noted that the deep neural networks used in the present invention are ECANet and SENet. In addition, Adam is used as the optimizer for all networks, and the learning rate is set to 0.0001. Combined with Figure 2, the network is divided into four parts, namely Block1 - Block4. AL is a classifier. The deepest network is used as the teacher model, and the backend of the shallow network is connected to the classifier as the student model. First, the A-softmax function is used to calculate the loss between the output of the deepest network and the label, that is, the Hard loss. Then, the KL divergence function is used to calculate the loss between the prediction outputs of the deepest network and all shallow networks, that is, the Soft loss. Next, the L2 function is used to calculate the loss between the features of the deepest network and all shallow networks, that is, the Feature loss. Finally, these three losses are balanced by hyperparameters. The classification result of the deepest network's classifier (AL layer4, predict) is used as the teacher, and the rest of the shallow ones are used as students, guiding from two aspects of features and prediction results, that is, Feature loss and Soft loss. And a certain number of training rounds are set for the deep neural network classifier for training. Finally, the best model during training is selected for testing. When testing, the classifier at the backend of the shallow network is removed, and the output of the deepest network is used as the prediction result.
[0075] In the embodiment of the present invention, refer to Figure 3 , a false audio detection system based on self-knowledge distillation is further provided, including:
[0076] A speech feature input module, which is used to extract the log power spectrum features of the original speech waveform, and use the F0 sub-band of the log power spectrum as the input features for false audio detection;
[0077] A modeling module, which is used to model the input features obtained by the speech feature input module by using a deep neural network;
[0078] A hidden knowledge extraction module, which is used to calculate the loss between the label and the network prediction output, and obtain the hidden knowledge of the training set;
[0079] A shallow network enhancement module, which is used to calculate the loss between the prediction outputs of all shallow networks and the deepest network. The deepest network is used as the teacher model to guide all shallow networks to enhance the shallow networks;
[0080] A feature difference balancing module, which is used to calculate the loss between the features of all shallow networks and the deepest network to balance the feature differences between the shallow and deep network features.
[0081] Among them, the steps for the shallow network enhancement module to enhance the shallow network include:
[0082] First, the network is divided into n parts. The deepest network is used as the teacher model. Classifiers are set after all shallow networks to construct n - 1 student models. According to the prediction results, the teacher model guides all student models.
[0083] The present invention was experimentally tested on the public datasets ASVspoof 2019 LA and PA. To quantitatively evaluate the results of spoofed audio detection, equal error rate (EER) and minimum normalized tandem detection cost function (min-tDCF) were used as evaluation metrics.
[0084] Table 1
[0085]
[0086] Table 2
[0087]
[0088] Table 1 shows the experimental results based on different architectures of ECANet, and Table 2 shows the experimental results based on different architectures of SENet, where baseline represents the original network and SD represents the results after using self-knowledge distillation. As can be seen from Table 1 and Table 2, compared with the baseline system, the performance has been significantly improved after using self-knowledge distillation. This is because shallow features are very important for distinguishing real and spoofed speech, and the capture ability of shallow networks is insufficient. The self-distillation method of the present invention uses the deepest network as the teacher model to guide the shallow network, strengthening the ability of the shallow network to capture features, thus further improving the performance of spoofed audio detection.
[0089] These results prove the effectiveness of the method proposed by the present invention. In addition, these results also show that self-knowledge distillation can greatly exploit the discriminative information of spoofed audio detection features.
[0090] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for detecting fake audio based on self-knowledge distillation, characterized in that, It includes the following steps: S1: Extract the logarithmic power spectrum features of the original speech waveform, and use the F0 sub-band of the logarithmic power spectrum features as the input features for fake audio detection; S2: Use a deep neural network to model the input features, calculate the loss between the label and the network output, and extract the hidden knowledge of the training set; S3: Divide the network into n parts, calculate the loss of the prediction outputs of all shallow networks and the deepest network, use the deepest network as the teacher model, and add a classifier to the backend of the shallow networks to form the student model. The teacher model guides the student model according to the prediction results to enhance the shallow networks; S4: In the feature dimension, the teacher model distills knowledge into the student model to balance the feature differences between the shallow and deep networks, and calculate the loss between the features of all shallow networks and the deepest network to balance the feature differences between the shallow and deep network features; The specific steps include: S401: Use the mean square error function to calculate the loss between the features of the deepest network and all shallow networks of the network; , where MSE represents the mean square error function, represents the features of the shallow network, is the feature of the deepest network; is the sum of the feature losses between all shallow networks and the deepest network; S402: Balance the three losses by setting hyperparameters; , Among them, and are the balance coefficients of losses respectively, is the final loss.
2. The method for detecting fake audio based on self-knowledge distillation according to claim 1, wherein The specific steps of step S1 include: S101: Use the short-time Fourier transform (STFT) to convert the time-domain speech signal into a time-frequency domain speech signal; , where \(x[k]\) represents the speech signal in the time domain, \(k\) is the time index of the speech signal, and are the corresponding real and imaginary parts of the STFT, \(t\) is the index of the number of time frames, and \(f\) is the index of the frequency unit; S102: Take the logarithm of the real and imaginary parts of the STFT to obtain the logarithmic power spectrum features; , where log represents the logarithm operation, which is the full-band feature of the required logarithmic power spectrum; S103: Apply the 0-400 Hz frequency band of the logarithmic power spectrum as the required F0 sub-band; 。 3. The method for detecting fake audio based on self-knowledge distillation according to claim 1, wherein, The specific steps of step S2 include: Use the function to calculate the loss between the label and the network prediction output: , Among them represents the A-softmax function is the predicted output of the deepest layer of the network, and L is the label is the loss between the label and the predicted output of the deepest layer of the network 4. The method for detecting fake audio based on self-knowledge distillation according to claim 1, wherein The specific steps of step S3 include: Use the Kullback-Leible divergence function to calculate the loss between the prediction of the deepest network and the predictions of all shallow networks; , where KL represents the Kullback-Leible divergence function, is the predicted output representing the shallow network, is the sum of the prediction losses between all shallow networks and the deepest network.
5. A fake audio detection system based on self-knowledge distillation, which adopts the fake audio detection method based on self-knowledge distillation as described in any one of claims 1 to 4, characterized in that, It includes: A speech feature input module, which is used to extract the logarithmic power spectrum features of the original speech waveform, and use the F0 sub-band of the logarithmic power spectrum as the input features for fake audio detection; A modeling module, which is used to model the input features obtained by the speech feature input module using a deep neural network; A hidden knowledge extraction module, which is used to calculate the loss between the label and the network prediction output and obtain the hidden knowledge of the training set; A shallow network enhancement module, which is used to calculate the loss of the prediction outputs of all shallow networks and the deepest network, and use the deepest network as the teacher model to guide all shallow networks to enhance the shallow networks; A feature difference balancing module, which is used to calculate the loss between the features of all shallow networks and the deepest network to balance the feature differences between the shallow and deep network features.
6. The false audio detection system based on self-knowledge distillation according to claim 5, wherein The steps for the shallow network enhancement module to enhance the shallow networks include: First, divide the network into n parts, use the deepest network as the teacher model, set classifiers after all shallow networks to construct n-1 student models, and according to the prediction results, the teacher model guides all student models.
Citation Information
Patent Citations
Speaker model compression system and method based on double-layer knowledge distillation
CN112712099A
Speech enhancement method based on cross-layer similarity knowledge distillation
CN114067819A