Cross-domain voice anti-spoofing method and system

By forging true and false speech data and using end-to-end models for speech classification, the problem of low accuracy in speech recognition in the prior art is solved, and a higher accuracy of cross-domain speech recognition is achieved.

CN116386648BActive Publication Date: 2025-06-27SHANGHAI NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310594301.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2025-06-27
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

The existing pronunciation pseudo-recognition model has low accuracy due to the domain mismatch problem of training data, which reduces the accuracy of pronunciation pseudo-recognition.

Method used

By falsifying the identity and voice of the speaker in the true and false voice data, forged voice samples are generated to reduce the problem of domain mismatch. Use end-to-end models for speech classification and improve the convergence of the model through parameter updates.

Benefits of technology

It effectively reduces the domain mismatch problem between true and false speech data and forged speech samples, improves the cross-domain type diversity of samples, and enables the trained end-to-end model to effectively detect speech to be tested in different domains, improving the accuracy of cross-domain speech to be false.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386648B_ABST
    Figure CN116386648B_ABST
Patent Text Reader

Abstract

The present invention provides a cross-domain voice anti-spoofing method and system. The method includes: respectively performing forgery processing on the speaker identity and voice in the genuine and fake voice data to obtain forged voice samples; extracting features from the genuine and fake voice data and the forged voice samples to obtain sample voice features, and inputting the sample voice features into an end-to-end model for voice classification; updating the parameters of the end-to-end model according to the authenticity labels of the sample voice features and the voice classification results, and inputting the voice to be measured into the converged end-to-end model for voice anti-spoofing to obtain cross-domain voice anti-spoofing results. By respectively performing forgery processing on the speaker identity and voice in the genuine and fake voice data, the present invention can effectively reduce the problem of domain mismatch between the genuine and fake voice data and the forged voice samples, improve the diversity of the cross-domain types of the samples, enable the trained end-to-end model to effectively perform voice anti-spoofing on the voice to be measured in different domains, and improve the accuracy of cross-domain voice anti-spoofing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice detection, and in particular, to a cross-domain voice anti-spoofing method and system. Background Art

[0002] Voice is one of the most natural and convenient means in human-computer interaction and biometric verification. In certain specific situations, such as in a telephone conference, voice may be the only available biometric technology. An automatic speaker verification system can provide a reliable voiceprint verification means, but the voiceprint verification means is not without security and privacy issues, and how to obtain a secure and reliable voiceprint verification system is a major challenge in the current field of voice deepfake detection.

[0003] The security issues in the voiceprint verification system are mainly related to the possibility of a successful spoofing attack on the voiceprint verification system by deepfake voices. Lawbreakers can use forged voices after operation that are very similar to a specific speaker to conduct spoofing attacks on the voiceprint verification system to obtain illegal access to protected services or resources. At the same time, the methods of deepfake voices are also developing rapidly, and their availability is becoming easier. Without sufficient protection, spoofing attacks may greatly reduce the security and reliability of the voiceprint verification system. Therefore, voice anti-spoofing methods are becoming more and more important.

[0004] In the existing voice anti-spoofing process, voice anti-spoofing is generally based on a voice anti-spoofing model. However, in the training process of the existing voice anti-spoofing model, there is a problem of domain mismatch in the training data, resulting in low accuracy of the trained voice anti-spoofing model and reducing the accuracy of voice anti-spoofing. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a cross-domain voice anti-spoofing method and system, aiming to solve the problem of low accuracy of existing voice anti-spoofing.

[0006] The embodiments of the present invention are implemented as follows. A cross-domain voice anti-spoofing method, the method includes:

[0007] Obtain true and false voice data, and respectively perform forgery processing on the speaker identity and voice in the true and false voice data to obtain forged voice samples;

[0008] Extract features from the true and false voice data and the forged voice samples to obtain sample voice features, and input the sample voice features into an end-to-end model for voice classification;

[0009] Update the parameters of the end-to-end model according to the authenticity labels of the sample voice features and the voice classification results until the end-to-end model converges;

[0010] Input the voice to be tested into the end-to-end model after convergence for voice anti-spoofing to obtain cross-domain voice anti-spoofing results.

[0011] Preferably, the forgery processing of the speaker identity and voice in the genuine and fake voice data respectively includes:

[0012] Respectively obtain the speech rates of each genuine and fake voice sample in the genuine and fake voice data, and perform speed perturbation on the speech rates of each genuine and fake voice sample to obtain voice perturbation samples, where the speed perturbation is used to forge the speaker identity corresponding to each genuine and fake voice sample;

[0013] Globally mix each genuine and fake voice sample with each voice perturbation sample to obtain the forged voice sample, where the global mixing is used to forge the voice categories of each genuine and fake voice sample and each voice perturbation sample.

[0014] Preferably, the feature extraction of the genuine and fake voice data and the forged voice sample includes:

[0015] Respectively sample the genuine and fake voice data and the forged voice sample to obtain sampled voices, and perform voice truncation processing or voice completion processing according to the voice durations of each sampled voice to obtain fixed-duration voices;

[0016] Input each fixed-duration voice into the band-pass filter bank in the end-to-end model for feature convolution processing to obtain the sample voice features, where the band-pass filter bank includes at least one band-pass filter initialized with Mel scale.

[0017] Preferably, the end-to-end model includes a band-pass filter bank, a residual module, a channel attention module, a time-domain attention module, a frequency-domain attention module, a pooling layer, and a classifier.

[0018] Preferably, the input of the sample voice features into the end-to-end model for voice classification includes:

[0019] Input the sample voice features into the residual module for residual processing to obtain residual features, and input the residual features into the channel attention module for channel dimension correction to obtain global channel features;

[0020] Respectively input the global channel features into the time-domain attention module and the frequency-domain attention module for time-domain dimension correction and frequency-domain dimension correction to obtain time-domain dimension features and frequency-domain dimension features;

[0021] Fuse the time-domain dimension features and the frequency-domain dimension features to obtain fused features, and input the fused features into the pooling layer for pooling processing;

[0022] Input the pooling processing result into the classifier for speech classification to obtain the speech classification result.

[0023] Preferably, the formula for globally mixing each real and fake speech sample with each speech perturbation sample includes:

[0024] X m = λX i +(1 - λ)X j

[0025] where X m is the forged speech sample, X i is a sample randomly selected from the real and fake speech samples, X j is a sample randomly selected from the speech perturbation samples, and λ is the proportionality coefficient of global mixing.

[0026] Preferably, the loss function for updating the parameters of the end-to-end model according to the authenticity label of the sample speech feature and the speech classification result includes:

[0027]

[0028] N is the sum of the number of samples between the real and fake speech data and the forged speech samples, W j (j = 0, 1) are the weights corresponding to the real and fake categories, y ij is the label corresponding to the i-th sample, is the label corresponding to the i-th sample after speech classification by the end-to-end model in the speech classification result.

[0029] Another object of the embodiments of the present invention is to provide a cross-domain speech forgery detection system, the system includes:

[0030] A sample forgery module, configured to obtain real and fake speech data, and respectively perform forgery processing on the speaker identity and speech in the real and fake speech data to obtain forged speech samples;

[0031] A speech classification module, configured to extract features from the real and fake speech data and the forged speech samples to obtain sample speech features, and input the sample speech features into an end-to-end model for speech classification;

[0032] A model training module, configured to update the parameters of the end-to-end model according to the authenticity label of the sample speech feature and the speech classification result until the end-to-end model converges;

[0033] A speech forgery detection module, configured to input the speech to be detected into the converged end-to-end model for speech forgery detection to obtain a cross-domain speech forgery detection result.

[0034] Preferably, the sample forgery module is further configured to: respectively obtain the speech rates of the true and false speech samples in the true and false speech data, and perform speed perturbation on the speech rates of the true and false speech samples to obtain speech perturbation samples, where the speed perturbation is used to forge the speaker identities corresponding to the true and false speech samples;

[0035] Globally mix each true and false speech sample with each speech perturbation sample to obtain the forged speech samples, where the global mixing is used to forge the speech categories of each true and false speech sample and each speech perturbation sample.

[0036] Preferably, the speech classification module is further configured to: respectively sample the true and false speech data and the forged speech samples to obtain sampled speech, and perform speech truncation processing or speech completion processing on the sampled speech according to the speech durations of the sampled speech to obtain fixed-duration speech;

[0037] Input each fixed-duration speech into the band-pass filter bank in the end-to-end model for feature convolution processing to obtain the sample speech features, where the band-pass filter bank includes at least one band-pass filter initialized with the Mel scale.

[0038] In the embodiment of the present invention, by respectively forging the speaker identities and speech in the true and false speech data, the problem of domain mismatch between the true and false speech data and the forged speech samples can be effectively reduced, the diversity of the cross-domain types of the samples can be improved, so that the trained end-to-end model can effectively perform speech forgery identification on the to-be-tested speech in different domains, and the accuracy of cross-domain speech forgery identification is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of a cross-domain speech forgery identification method provided by the first embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of a global mixing mechanism provided by the first embodiment of the present invention;

[0041] Figure 3 is a schematic diagram of an end-to-end model provided by the first embodiment of the present invention;

[0042] Figure 4 is a schematic diagram of a channel attention module provided by the first embodiment of the present invention;

[0043] Figure 5 is a schematic diagram of a time domain attention module provided by the first embodiment of the present invention;

[0044] Figure 6 is a schematic diagram of a frequency domain attention module provided by the first embodiment of the present invention;

[0045] Figure 7It is a flowchart of the cross - domain voice anti - spoofing method provided by the second embodiment of the present invention;

[0046] Figure 8 It is a schematic structural diagram of the cross - domain voice anti - spoofing system provided by the third embodiment of the present invention;

[0047] Figure 9 It is a schematic diagram of end - to - end model training provided by the third embodiment of the present invention;

[0048] Figure 10 It is a schematic structural diagram of the terminal device provided by the fourth embodiment of the present invention. Detailed implementation manners

[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0050] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.

[0051] Embodiment 1

[0052] Please refer to Figure 1 , which is a flowchart of the cross - domain voice anti - spoofing method provided by the first embodiment of the present invention. This cross - domain voice anti - spoofing method can be applied to any terminal device or system, and the cross - domain voice anti - spoofing method includes the steps:

[0053] Step S10, obtain real and fake voice data, and respectively perform forgery processing on the speaker identity and voice in the real and fake voice data to obtain forged voice samples;

[0054] Among them, the voice anti - spoofing model generally uses the dataset used in the ASVspoof 2021 DF task for training and testing. The training set and validation set for model training use the training set and validation set of the logical access (LA) task released by ASVspoof 2019, and the test set uses the test set of the deep fake (DF) task released by ASVspoof 2021.

[0055] The training set of ASVspoof 2019LA contains 20 speakers (8 men and 12 women) and 6 methods of forging voice. The validation set contains 10 speakers (4 men and 6 women) and the same 6 methods of forging voice. The test set of ASVspoof 2021DF contains far more than 20 different speakers and forged voices synthesized by more than 100 different algorithms. It can be seen that the differences in the number of speakers and the types of forged voice algorithms will bring great domain mismatch problems to the training set and the test set, and will also bring great challenges to the design of voice forgery authentication systems.

[0056] Therefore, when constructing a data set, this embodiment can effectively reduce the problem of domain mismatch between true and false voice data and forged voice samples by respectively forging the speaker identity and voice in the true and false voice data, and improve the diversity of cross-domain types of samples, so that the trained end-to-end model can effectively perform voice authentication on the test voices in different domains, thereby improving the accuracy of cross-domain voice authentication.

[0057] Optionally, in this step, the speaker identity and voice in the true and false voice data are respectively falsified, including:

[0058] Respectively obtaining the speech speed of each true and false speech sample in the true and false speech data, and performing speed perturbation on the speech speed of each true and false speech sample to obtain a speech perturbation sample;

[0059] Among them, speed perturbation is used to forge the identity of the speaker corresponding to each true and false voice sample, and to perturb the speed of each true and false voice sample, reducing the speed to 0.9 times the original speed and speeding it up to 1.1 times the original speed, affecting the pitch of the original speaker, thereby implicitly increasing the number of speakers in the data set;

[0060] Globally mixing each true and false speech sample with each speech disturbance sample to obtain the forged speech sample;

[0061] Among them, global mixing is used to forge the voice category of each true and false speech sample and each speech perturbation sample. By globally mixing each true and false speech sample with each speech perturbation sample, more types of forged speech samples can be obtained, thereby alleviating the domain mismatch problem related to the speaker identity and the forged speech method in the dataset, thereby improving the speech authentication performance.

[0062] See also Figure 2 , a global mixing mechanism is used to mix every two samples within and between classes of true and false speech samples and speech perturbation samples in a certain proportion, so as to generate more forged speech samples and increase the types of forged speech samples generated in the dataset.

[0063] Further, the formula for globally mixing each true / false voice sample and each voice perturbation sample includes:

[0064] X m = λX i + (1 - λ)X j

[0065] where X m is the forged voice sample, X i is a sample randomly selected from the true / false voice samples, X j is a sample randomly selected from the voice perturbation samples, and λ is the proportionality coefficient for global mixing, which follows Beta(α, α), and the value of α is 0.5.

[0066] Step S20: Extract features from the true / false voice data and the forged voice sample to obtain sample voice features, and input the sample voice features into an end-to-end model for voice classification;

[0067] Among them, referring to Figure 3 , the end-to-end model includes a band-pass filter bank, a residual module, a channel attention module, a time-domain attention module, a frequency-domain attention module, a pooling layer, and a classifier. Optionally, in this step, inputting the sample voice features into the end-to-end model for voice classification includes:

[0068] Input the sample voice features into the residual module for residual processing to obtain residual features, and input the residual features into the channel attention module for channel dimension correction to obtain global channel features;

[0069] Among them, the sample voice features are low-dimensional information representations with distinctiveness. Inputting the sample voice features into the residual neural network module in the residual module for residual processing obtains a high-dimensional residual feature, and then inputting the residual feature into the channel attention module obtains global channel features after correcting the inter-channel interaction information and feature importance;

[0070] As Figure 4 shown, for the high-dimensional information representation (residual feature) X ∈ R C×F×T extracted by a given residual module, first perform global information averaging in the channel dimension, and the specific implementation method is as follows:

[0071]

[0072] C, F, and T are the dimensions of the channel dimension, frequency domain dimension, and time dimension respectively, j and k are the j-th dimension of the frequency domain dimension and the k-th dimension of the time dimension respectively, and then obtain the weight value for correcting the channel dimension information. The implementation method of the corrected high-dimensional representation X c is as follows:

[0073]

[0074] Among them, W m (m = 1, 2) are the parameters of the m-th fully connected layer, r is the reduction multiple of the channel dimension, relu is the rectified linear unit, δ is the sigmoid activation function, × is matrix multiplication, is the element-wise multiplication.

[0075] Respectively input the global channel features into the time-domain attention module and the frequency-domain attention module for time-domain dimension correction and frequency-domain dimension correction, to obtain time-dimensional features and frequency-dimensional features;

[0076] Fuse the time-dimensional features and the frequency-dimensional features to obtain fused features, and input the fused features into a pooling layer for pooling processing;

[0077] Input the result of the pooling process into the classifier for speech classification to obtain the speech classification result;

[0078] Among them, the high-dimensional representation (global channel feature) X after channel dimension correction c is respectively input into the time-domain attention module and the frequency-domain attention module, and information representations with distinguishable feature importance among different time frames and different frequency-domain subbands can be obtained respectively.

[0079] Similar to the implementation method of the channel attention module, the time-domain attention module and the frequency-domain attention module also first obtain weight values for evaluating feature importance, and then perform element-wise multiplication with their respective corresponding feature dimensions to obtain distinguishable features of different time frames and different frequency-domain subbands after correction. As Figure 5 shown, the specific implementation method of the time-domain attention module is as follows:

[0080]

[0081] Among them, C, F, and T are the dimensions of the channel dimension, the frequency-domain dimension, and the time dimension respectively, i and j are the i-th dimension of the channel dimension and the j-th dimension of the frequency-domain dimension respectively, and then weight values for correcting time-dimensional information are obtained, and the corrected time-dimensional feature X t is implemented as follows:

[0082]

[0083] Among them, W m (m = 1, 2) are the parameters of the m-th fully connected layer, r is the reduction multiple of the channel dimension, relu is the rectified linear unit, δ is the sigmoid activation function, × is matrix multiplication, is the element-wise multiplication. And the frequency-domain attention module is asFigure 6 As shown below, the specific implementation method is as follows:

[0084]

[0085] Where C, F, and T are the dimensions of the channel dimension, frequency domain dimension, and time dimension respectively, and i and k are the i-th dimension of the channel dimension and the k-th dimension of the time dimension respectively. Then, the weight value for correcting the frequency domain dimension information is obtained, and the corrected frequency domain dimension feature X f The implementation method is as follows:

[0086]

[0087] Where, W m (m = 1, 2) are the parameters of the m-th fully connected layer, r is the reduction multiple of the channel dimension, relu is the rectified linear unit, δ is the sigmoid activation function, × is matrix multiplication, is the element-wise multiplication. Multiply the corrected time dimension feature X t and the corrected frequency domain dimension feature X f element-wise to achieve the feature fusion of different dimensions. Finally, send the fused feature into the pooling layer to convert the frame-level information representation into a sentence-level representation, which is convenient for the subsequent classifier to distinguish the authenticity of the input speech.

[0088] In this step, by sending the residual feature into the channel dimension-based attention mechanism to model the global information, the weight value corresponding to each channel is generated to re-measure the importance degree of the features on each channel, so as to reduce the impact of insufficient information interaction caused by modeling local information through convolution, thereby obtaining the globally high-dimensional information representation corrected based on the channel dimension;

[0089] By sending the globally high-dimensional information representation (global channel feature) corrected based on the channel dimension obtained into the attention mechanism modules based on the time dimension and the frequency domain dimension respectively, measure the importance degree of the features of different time frames and different frequency domain sub-bands respectively, and then fuse the features output by these two modules to obtain a more discriminative fusion feature for true and false speech;

[0090] By aggregating the frame-level information representation with fusion into a sentence-level information representation and then sending it into the backend classifier, the probability of the authenticity of the input speech can be obtained.

[0091] Step S30, update the parameters of the end-to-end model according to the true / false label of the sample speech feature and the speech classification result until the end-to-end model converges;

[0092] Among them, using the true / false labels of the sample speech features, the constructed end-to-end model based on the joint attention mechanism is trained to obtain a robust true / false speech discriminative representation for true / false speech discrimination. In this step, the designed representation extraction module based on the joint attention mechanism is mainly designed in three dimensions: time, frequency domain, and channel, and embedded in different positions of the model, aiming to extract the discriminative representation information of true / false speech hidden between different time frames, different frequency domain subbands, and different channels, and further improve the performance of the speech anti-spoofing system.

[0093] In this embodiment, RawNet2 with the best performance among the 4 baseline systems provided by the ASVspoof 2021 DF task is used as the most basic end-to-end model, and the specific construction of the model is as Figure 3 shown, including a feature extractor composed of a band-pass filter bank, 6 modules for extracting high-dimensional representations composed of residual modules and channel attention modules, time-domain and frequency-domain attention mechanism modules, a pooling layer composed of gated recurrent units, and a speech true / false classifier composed of two fully connected layers.

[0094] In this step, the design principles of the channel attention mechanism module, the time-domain attention mechanism module, and the frequency-domain attention mechanism module are similar. They all first perform global information pooling to obtain a weight value measuring the importance of features, and then correct the features in the corresponding dimension to strengthen the importance of discriminative features and weaken redundant feature information, so as to extract a more discriminative and robust high-dimensional representation of true / false speech. For specific details, please refer to Figures 4 to 6 . The specific experimental configuration during training: train for 100 rounds, the loss function is the weighted cross-entropy loss, train with the Adam optimizer, the learning rate is 0.0001, the batch size is 4, and the weight decay value is 0.0001.

[0095] Optionally, the training of the entire end-to-end model uses the weighted cross-entropy loss function. The loss function used to update the parameters of the end-to-end model according to the true / false labels of the sample speech features and the speech classification results includes:

[0096]

[0097] N is the sum of the number of samples between the true / false speech data and the forged speech samples, and W j (j = 0, 1) are the weights corresponding to the true / false categories, y ij is the label corresponding to the i-th sample, is the label corresponding to the i-th sample after speech classification by the end-to-end model in the speech classification result.

[0098] In this step, when the number of iterations of the end-to-end model is greater than the iteration threshold or the model loss value is less than the loss threshold, it is determined that the end-to-end model converges, and the trained end-to-end model is saved.

[0099] Step S40: Input the speech to be tested into the converged end-to-end model for voice anti-spoofing to obtain a cross-domain voice anti-spoofing result.

[0100] Among them, the converged end-to-end model is used to perform voice anti-spoofing on the speech to be tested to obtain the discrimination probability of the authenticity of the speech to be tested, so as to output the cross-domain voice anti-spoofing result of the speech to be tested.

[0101] Optionally, in this step, the EER index for measuring the anti-spoofing performance of the designed end-to-end voice anti-spoofing system can also be calculated using the scores of the test set speech. The EER result of the test set in the ASVspoof 2021 deepfake (DF) detection task in this embodiment is 20.03%. The equal error rate (EER) of the converged end-to-end model is improved by 10.50% compared with the result of 22.38% of the best baseline system RawNet2 provided officially in this task, proving the effectiveness of this embodiment.

[0102] The beneficial effects of this embodiment include:

[0103] 1. Implicit speaker identity generation method based on speaker speech rate variation

[0104] There will be a problem of domain mismatch in speaker identity between the training data and test data for voice anti-spoofing, and it is also urgent to reduce the degree of mismatch between them. Both genuine and fake voices have the corresponding speech rate and pitch of the speaker. Therefore, the speech rate and pitch of the speaker also become factors characterizing the speaker's identity. Among them, the pitch size mainly depends on the fundamental frequency of the sound wave. By perturbing the speech rate of the speaker, the fundamental frequency of the sound wave will be changed, and thus the pitch size will also be changed, which implicitly changes the identity attribute of the speaker to some extent. Based on the above, this embodiment designs an implicit speaker identity generation method based on speaker speech rate variation. First, change the speech rate of the speakers of genuine and fake voices to affect their pitch size, and implicitly add new speaker identities on the basis of the original speakers, so as to reduce the domain mismatch problem caused by speaker identity and improve the performance of the voice anti-spoofing system.

[0105] 2. Fake speech generation method based on global mixing mechanism

[0106] The training and testing of voice anti-spoofing are both for genuine and spoofed voices. From the perspective of application scenarios, spoofed voices pose a greater challenge to the security of the system. And how to detect more spoofed voices depends to a large extent on the types of spoofed voices used during training. Therefore, how to generate more spoofed voices on the original dataset has also become the focus of this invention. Based on the above, this embodiment designs a method for generating spoofed voices based on a global mixing mechanism, which adopts global mixing within and between genuine and spoofed voice classes with multiple speakers to generate more spoofed voices, and to a certain extent solves the domain mismatch problem between data caused by the method of generating spoofed voices, enhancing the robustness of voice anti-spoofing.

[0107] 3. Design of Feature Extraction Based on Joint Attention Mechanism

[0108] The quality of voice anti-spoofing performance depends to a large extent on whether the system can extract discriminative feature information for genuine and spoofed voices when modeling them. And how to automatically focus on this discriminative feature information is also an aspect that this invention focuses on. Based on the above, this embodiment designs a feature extraction based on a joint attention mechanism, which is designed respectively in three dimensions: time, frequency domain, and channel, and embedded in different positions of the system, aiming to extract the discriminative feature information of genuine and spoofed voices hidden between different time frames, different frequency domain sub-bands, and different channels, further improving the accuracy of voice anti-spoofing.

[0109] Embodiment 2

[0110] Please refer to Figure 7 , which is the flowchart of the cross-domain voice anti-spoofing method provided by the second embodiment of this invention. This embodiment is used to further refine step S20 in the first embodiment, including the steps:

[0111] Step S21, sample the genuine and spoofed voice data and the spoofed voice samples respectively to obtain sampled voices, and perform voice truncation processing or voice filling processing according to the voice duration of each sampled voice to obtain fixed-duration voices;

[0112] Among them, the original waveforms of the genuine and spoofed voice data and the spoofed voice samples are preprocessed into fixed-duration voice samples. In this step, the genuine and spoofed voice data and the spoofed voice samples can be sampled at 16 kHz to obtain sampled voices. Since there are differences in the durations of the sampled voices obtained by sampling, therefore, each sampled voice is truncated or replicated and filled to a fixed-duration voice to obtain this fixed-duration voice, which effectively facilitates the subsequent acquisition of features of each fixed-duration voice. This fixed duration can be set according to requirements, and the fixed duration of this step is set to 4 seconds;

[0113] Step S22: Input each fixed-duration voice into the band-pass filter bank in the end-to-end model for feature convolution processing to obtain the sample voice features.

[0114] Among them, the band-pass filter bank includes at least one band-pass filter after Mel-scale initialization. Input each fixed-duration voice into the feature extractor composed of 128 band-pass filters initialized by Mel-scale in the end-to-end model, so that the original waveform of each fixed-duration voice is convolved in the time domain with 128 band-pass filters, and a preliminary low-dimensional representation for distinguishing true and false voices is extracted to obtain the sample voice features.

[0115] In this embodiment, by performing voice truncation processing or voice completion processing on the voice duration of each sampled voice, the sampled voices can be effectively made to have the same duration, which effectively facilitates the extraction of the preliminary low-dimensional representation for distinguishing true and false voices. By performing feature convolution processing on each fixed-duration voice, the preliminary low-dimensional representation for distinguishing true and false voices can be effectively extracted.

[0116] Embodiment 3

[0117] Please refer to Figure 8 , which is a schematic structural diagram of the cross-domain voice anti-spoofing system 100 provided by the third embodiment of the present invention, including: a sample forgery module 10, a voice classification module 11, a model training module 12, and a voice anti-spoofing module 13, where:

[0118] The sample forgery module 10 is used to obtain true and false voice data, and perform forgery processing on the speaker identity and voice in the true and false voice data respectively to obtain forged voice samples.

[0119] Optionally, the sample forgery module 10 is further used to: respectively obtain the speech rates of the true and false voice samples in the true and false voice data, and perform speed perturbation on the speech rates of the true and false voice samples to obtain voice perturbation samples, and the speed perturbation is used to forge the speaker identity corresponding to each true and false voice sample;

[0120] Globally mix each true and false voice sample with each voice perturbation sample to obtain the forged voice sample, and the global mixing is used to forge the voice categories of each true and false voice sample and each voice perturbation sample.

[0121] The voice classification module 11 is used to extract features from the true and false voice data and the forged voice samples to obtain sample voice features, and input the sample voice features into the end-to-end model for voice classification.

[0122] Optionally, the voice classification module 11 is further configured to: sample the true and false voice data and the forged voice samples respectively to obtain sampled voices, and perform voice truncation processing or voice completion processing on the basis of the voice durations of the sampled voices to obtain fixed-duration voices;

[0123] Input each fixed-duration voice into the bandpass filter bank in the end-to-end model for feature convolution processing to obtain the sample voice features, where the bandpass filter bank includes at least one bandpass filter after Mel scale initialization.

[0124] The model training module 12 is configured to update the parameters of the end-to-end model according to the authenticity labels of the sample voice features and the voice classification results until the end-to-end model converges.

[0125] In this embodiment, the end-to-end model includes a bandpass filter bank, a residual module, a channel attention module, a time domain attention module, a frequency domain attention module, a pooling layer, and a classifier.

[0126] Optionally, the model training module 12 is further configured to: input the sample voice features into the residual module for residual processing to obtain residual features, and input the residual features into the channel attention module for channel dimension correction to obtain global channel features;

[0127] Input the global channel features into the time domain attention module and the frequency domain attention module respectively for time domain dimension correction and frequency domain dimension correction to obtain time dimension features and frequency domain dimension features;

[0128] Fuse the time dimension features and the frequency domain dimension features to obtain fused features, and input the fused features into the pooling layer for pooling processing;

[0129] Input the pooling processing result into the classifier for voice classification to obtain the voice classification result.

[0130] The voice forgery detection module 13 is configured to input the voice to be detected into the converged end-to-end model for voice forgery detection to obtain a cross-domain voice forgery detection result.

[0131] Please refer to Figure 9 , in the dataset construction stage, by varying the speaker's speaking speed of the original true and false voice samples, voice perturbation samples are obtained, each true and false voice sample is globally mixed with each voice perturbation sample to obtain the forged voice samples, and the forged voice samples and the true and false voice samples are combined to obtain true and false voice training samples;

[0132] In the training stage of the end-to-end model, true and false speech training samples are input into the end-to-end model for true and false classification. The loss function is calculated based on the true and false classification results, and the parameters of the end-to-end model are updated based on the loss function until the end-to-end model converges and then the model is saved;

[0133] In the testing stage, true and false speech samples to be tested are input into the converged end-to-end model for true and false speech prediction to obtain true and false speech prediction scores, and true and false speech discrimination results are generated based on the true and false speech prediction scores to demonstrate the accuracy of the converged end-to-end model.

[0134] In this embodiment, by respectively performing forgery processing on the speaker identity and speech in the true and false speech data, the problem of domain mismatch between the true and false speech data and the forged speech samples can be effectively reduced, the diversity of cross-domain types of samples is improved, so that the trained end-to-end model can effectively perform speech forgery detection on the speech to be tested in different domains, and the accuracy of cross-domain speech forgery detection is improved.

[0135] Embodiment 4

[0136] Figure 10 It is a structural block diagram of a terminal device 2 provided in the fourth embodiment of the present application. As Figure 10 shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for the cross-domain speech forgery detection method. When the processor 20 executes the computer program 22, the steps in each embodiment of the above-mentioned cross-domain speech forgery detection method are implemented.

[0137] Exemplarily, the computer program 22 can be divided into one or more modules, the one or more modules are stored in the memory 21, and are executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, the processor 20 and the memory 21.

[0138] The so-called processor 20 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0139] The memory 21 may be an internal storage unit of the terminal device 2, such as the hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both the internal storage unit of the terminal device 2 and the external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.

[0140] In addition, in each embodiment of the present application, each functional module may be integrated in a processing unit, may also exist physically alone for each unit, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0141] When an integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0142] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A cross-domain voice anti-spoofing method, characterized in that The method includes: Obtaining true and false speech data, and respectively performing forgery processing on the speaker identity and speech in the true and false speech data to obtain forged speech samples; Performing feature extraction on the true and false speech data and the forged speech samples to obtain sample speech features, and inputting the sample speech features into an end-to-end model for speech classification; Updating the parameters of the end-to-end model according to the authenticity labels of the sample speech features and the speech classification results until the end-to-end model converges; Inputting the speech to be tested into the converged end-to-end model for speech forgery verification to obtain a cross-domain speech forgery verification result; The end-to-end model includes a band-pass filter bank, a residual module, a channel attention module, a time-domain attention module, a frequency-domain attention module, a pooling layer, and a classifier.

2. The cross-domain voice anti-spoofing method according to claim 1, characterized in that The performing forgery processing on the speaker identity and speech in the true and false speech data respectively includes: Respectively obtaining the speech rates of the true and false speech samples in the true and false speech data, and performing speed perturbation on the speech rates of the true and false speech samples to obtain speech perturbation samples, where the speed perturbation is used to forge the speaker identity corresponding to each true and false speech sample; Globally mixing each true and false speech sample with each speech perturbation sample to obtain the forged speech samples, where the global mixing is used to forge the speech categories of each true and false speech sample and each speech perturbation sample.

3. The cross-domain voice anti-spoofing method according to claim 1, characterized in that, The performing feature extraction on the true and false speech data and the forged speech samples includes: Respectively sampling the true and false speech data and the forged speech samples to obtain sampled speech, and performing speech truncation processing or speech completion processing on the sampled speech according to the speech duration of each sampled speech to obtain fixed-duration speech; Inputting each fixed-duration speech into the band-pass filter bank in the end-to-end model for feature convolution processing to obtain the sample speech features, where the band-pass filter bank includes at least one band-pass filter initialized with Mel scale.

4. The cross-domain voice anti-counterfeiting method according to claim 1, characterized in that The inputting the sample speech features into the end-to-end model for speech classification includes: Inputting the sample speech features into the residual module for residual processing to obtain residual features, and inputting the residual features into the channel attention module for channel dimension correction to obtain global channel features; Respectively inputting the global channel features into the time-domain attention module and the frequency-domain attention module for time-domain dimension correction and frequency-domain dimension correction to obtain time-domain dimension features and frequency-domain dimension features; Fusing the time-domain dimension features and the frequency-domain dimension features to obtain fused features, and inputting the fused features into the pooling layer for pooling processing; Inputting the pooling processing result into the classifier for speech classification to obtain the speech classification result.

5. The cross-domain voice anti-spoofing method according to claim 2, wherein The formula used for globally mixing each true and false speech sample with each speech perturbation sample includes: ; wherein, X m is the forged voice sample, X i is a sample randomly selected from the true and false voice samples, X j is a sample randomly selected from the voice perturbation samples, and λ is the proportionality coefficient of global mixing.

6. The cross-domain voice anti-spoofing method according to any one of claims 1 to 5, characterized in that, The loss function used for updating the parameters of the end-to-end model according to the authenticity labels of the sample speech features and the speech classification results includes: ; N is the sum of the number of samples between the true / false voice data and the forged voice samples, W j is the weight corresponding to the true / false category, y ij is the i label corresponding to the ŷ ij In the voice classification result, the i label corresponding to the j th sample after voice classification by the end-to-end model, where = 0, 1.

7. A cross-domain voice anti-spoofing system, characterized in that, The system includes: A sample forgery module, configured to obtain true and false speech data, and respectively perform forgery processing on the speaker identity and speech in the true and false speech data to obtain forged speech samples; A voice classification module, which is used to extract features from the true and false voice data and the forged voice samples to obtain sample voice features, and input the sample voice features into an end-to-end model for voice classification; A model training module, which is used to update the parameters of the end-to-end model according to the authenticity labels of the sample voice features and the voice classification results until the end-to-end model converges; A voice forgery detection module, which is used to input the voice to be detected into the converged end-to-end model for voice forgery detection to obtain a cross-domain voice forgery detection result; The end-to-end model includes a band-pass filter bank, a residual module, a channel attention module, a time-domain attention module, a frequency-domain attention module, a pooling layer, and a classifier.

8. The cross-domain voice anti-spoofing system according to claim 7, characterized in that, The sample forgery module is further used for: Obtaining the speech rates of the true and false voice samples in the true and false voice data respectively, and performing speed perturbation on the speech rates of the true and false voice samples to obtain voice perturbation samples, where the speed perturbation is used to forge the speaker identities corresponding to the true and false voice samples; Globally mixing each true and false voice sample with each voice perturbation sample to obtain the forged voice sample, where the global mixing is used to forge the voice categories of each true and false voice sample and each voice perturbation sample.

9. The cross-domain voice anti-spoofing system according to claim 7, wherein The voice classification module is further used for: Sampling the true and false voice data and the forged voice samples respectively to obtain sampled voices, and performing voice truncation processing or voice completion processing on the sampled voices according to the voice durations of the sampled voices to obtain fixed-duration voices; Inputting each fixed-duration voice into the band-pass filter bank in the end-to-end model for feature convolution processing to obtain the sample voice features, where the band-pass filter bank includes at least one band-pass filter initialized with a Mel scale.