Audio counterfeit detection method and system based on audio potential feature comparative learning

By combining the comparative learning method with multi-dimensional data augmentation and specific model architecture, the shortcomings of existing audio forgery detection methods in detection accuracy and generalization capabilities are solved, and effective detection of complex scenes and highly realistic forgery audio is achieved.

CN120220731AActive Publication Date: 2025-06-27ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Patent Information

Application Number
CN202510699483.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing audio forgery detection methods have shortcomings in the problems of insufficient detection accuracy, weak generalization ability, low sensitivity to new forgery modes, and vulnerability to confronting attacks, making it difficult to effectively deal with complex scenarios and highly realistic forgery audio.

Method used

Using a method based on potential audio features comparison learning, a fake audio data set is generated by multi-dimensional data augmentation of the original audio data, and an audio detection model RawNet2-C is constructed that combines Sinc convolutional layer, residual block and feature scaling map is combined with the comparison learning module to perform two-stage training to improve detection accuracy.

Benefits of technology

It significantly improves the detection accuracy of the model for high-realistic falsified audio, enhances the adaptability to background noise and speech speed/tongue changes, improves the robustness and discrimination ability in complex scenarios, and is significantly better than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220731A_ABST
    Figure CN120220731A_ABST
Patent Text Reader

Abstract

The invention discloses an audio counterfeit detection method and system based on audio potential feature comparative learning, and belongs to the technical field of audio counterfeit detection.The method comprises the steps that data enhancement is conducted on original audio data, a counterfeit audio data set is generated, an audio detection model is constructed, and the audio detection model comprises a comparative learning model; then performing first-stage training on the audio detection model based on the forged audio data set, after the first-stage training is completed, performing second-stage training by using a contrast learning model, and finally performing forged detection on the audio based on the audio detection model completing the first-stage training and the second-stage training. And the detection precision and the generalization ability are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of audio forgery detection, and particularly relates to an audio forgery detection method and system based on contrastive learning of audio latent features. Background Art

[0002] With the rapid development of audio information services, the user scale has been continuously growing. Currently, the scale of online music users in China has reached 608 million. Especially with the application of new artificial intelligence technologies and new applications such as generative artificial intelligence (AIGC) in the audio field, the audio output by deep learning-based audio generation and cloning algorithms is increasingly approaching real audio, resulting in the further aggregation and amplification of some legal risks during the audio transmission process. Therefore, the legal use of audio data is an issue that society currently attaches importance to.

[0003] At present, the methods for audio forgery detection mainly include: forgery detection methods based on audio signal features, such as detection methods using audio features such as phase spectrum, Mel spectrogram, spectrogram, and improved time delay; forgery detection methods based on machine learning, such as methods using linear SVM, weighted K-nearest neighbor, and enhanced tree ensemble, etc. However, the current technologies still have defects such as insufficient detection accuracy and weak generalization ability. Specifically, for the methods based on audio signal features, features such as phase spectrum and Mel spectrogram are difficult to comprehensively cover the complex changes of audio forgery. When facing advanced forgery technologies, it is difficult to distinguish between genuine and fake, and when the audio is in a complex environment, environmental noise, etc. will seriously interfere with feature extraction, resulting in a decrease in accuracy. At the same time, such methods are less sensitive to newly emerging forgery patterns and are difficult to adapt in a timely manner. For the methods based on machine learning, the model highly depends on the quality and diversity of training data. Incomplete samples or annotation biases are likely to cause a large number of misjudgments, and the detection effect on forged audio in rare and special scenarios is poor. Moreover, its generalization ability is insufficient, it is difficult to cope with the continuously evolving new forgery technologies, the computational resource consumption is large, it is difficult to apply in resource-constrained scenarios, and it is also easily vulnerable to adversarial attacks, making the detection results invalid. Therefore, there is an urgent need for a method to solve the above problems. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes an audio forgery detection method and system based on contrastive learning of audio latent features to solve the problems existing in the above-mentioned prior art.

[0005] In the first aspect, to achieve the above object, the present invention provides an audio forgery detection method based on contrastive learning of audio latent features, including the following steps:

[0006] Perform data augmentation on the original audio data to generate a forged audio dataset;

[0007] Construct an audio detection model, and the audio detection model includes a contrastive learning model;

[0008] Perform the first - stage training on the audio detection model based on the forged audio dataset;

[0009] After completing the first - stage training, use the contrastive learning model for the second - stage training;

[0010] Perform forged detection on the audio based on the audio detection model that has completed the first - stage training and the second - stage training.

[0011] Optionally, for data augmentation of the original audio data, the process of generating the forged audio dataset includes:

[0012] Confirm the positive - negative sample distribution ratio of the dataset. If the positive - negative sample ratio is not equal to 1:1, adjust the data;

[0013] Perform data augmentation on the adjusted original audio data to generate the forged audio dataset.

[0014] Optionally, the process of performing data augmentation on the adjusted original audio data includes: performing Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the adjusted original audio data.

[0015] Optionally, construct the audio detection model, and the audio detection model further includes: a Sinc layer, a residual block, a GRU layer, and a fully - connected layer.

[0016] Optionally, perform the first - stage training on the audio detection model based on the forged audio dataset. The number of training epochs for the first - stage training is N. After the training epochs are completed, the first - stage training ends. The process of the first - stage training includes:

[0017] Train the model based on the cross - entropy loss function.

[0018] Optionally, during the process of using the contrastive learning model for the second - stage training after completing the first - stage training, it includes:

[0019] Perform the second - stage training based on the cross - entropy loss function and the loss function of contrastive learning.

[0020] In a second aspect, the present invention also provides an audio forgery detection system based on audio latent feature contrastive learning for implementing an audio forgery detection method based on audio latent feature contrastive learning. The system includes:

[0021] A data processing module for performing data augmentation on the original audio data to generate a forged audio dataset;

[0022] A model construction module for constructing an audio detection model, where the audio detection model includes a contrastive learning model, a Sinc layer, a residual block, a GRU layer, and a fully connected layer;

[0023] A model training module for performing the first-stage training on the audio detection model based on the forged audio dataset, and after completing the first-stage training, using the contrastive learning model to perform the second-stage training;

[0024] A detection module for performing forged detection on audio based on the audio detection model that has completed the first-stage training and the second-stage training.

[0025] Optionally, the data processing module includes:

[0026] A data augmentation unit for performing Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the original audio data.

[0027] In a third aspect, the present invention further provides a computer terminal device, including:

[0028] One or more processors;

[0029] A memory coupled to the processor for storing one or more programs;

[0030] When the one or more programs are executed by the one or more processors, the one or more processors implement an audio forgery detection method based on audio latent feature contrastive learning.

[0031] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, an audio forgery detection method based on audio latent feature contrastive learning is implemented.

[0032] Compared with the prior art, the present invention has the following advantages and technical effects:

[0033] An audio forgery detection method and system based on audio latent feature contrast learning provided by the present invention first generates a forged audio dataset covering complex scenarios by performing multi-dimensional data augmentation on the original data (including Gaussian noise addition, waveform displacement, stretching, and pitch correction); secondly, constructs an audio detection model RawNet2-C that integrates a Sinc convolutional layer, a residual block, and a feature scaling mapping, and integrates a contrast learning module; after the first-stage training of the model based on the augmented data, further jointly optimizes the classification and feature discrimination capabilities through a two-stage training strategy, and finally significantly improves the detection accuracy of the model for highly realistic forged audio. Through data augmentation and staged training, the model can effectively enhance its adaptability to background noise, speech rate / tonal changes, directly extract deep latent features from the original waveform, avoid the limitations of traditional artificial feature design, and strengthen the robustness and discrimination capabilities in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings forming a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0035] Figure 1 It is a schematic diagram of audio enhancement according to an embodiment of the present invention;

[0036] Figure 2 It is a schematic diagram of residual connection according to an embodiment of the present invention;

[0037] Figure 3 It is a schematic diagram of FMS feature scaling according to an embodiment of the present invention;

[0038] Figure 4 It is a flowchart of the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0040] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0041] Embodiment 1

[0042] As Figure 4 shown, in this embodiment, an audio forgery detection method based on audio latent feature contrast learning is provided, including:

[0043] Perform data augmentation on the original audio data to generate a forged audio dataset;

[0044] Construct an audio detection model, where the audio detection model includes a contrastive learning model;

[0045] Perform the first-stage training on the audio detection model based on the forged audio dataset;

[0046] After completing the first-stage training, use the contrastive learning model for the second-stage training;

[0047] Perform forged detection on the audio based on the audio detection model that has completed the first-stage training and the second-stage training.

[0048] Specifically, the above process includes:

[0049] Step 1: Generate enhanced data based on the original audio data and data augmentation methods, and construct a forged audio dataset.

[0050] Step 2: Construct the model RawNet2-C based on audio latent features combined with contrastive learning.

[0051] Step 3: Perform the first-stage training on RawNet2-C based on the enhanced dataset to initially achieve the detection of forged audio.

[0052] Step 4: After the first-stage training of the RawNet2-C model ends, fuse contrastive learning, and based on the action mechanism of contrastive learning, perform the second-stage training to optimize and improve the detection effect of the model.

[0053] As an implementation method in this embodiment, the process of performing data augmentation on the original audio data to generate a forged audio dataset includes:

[0054] Confirm the positive and negative sample distribution ratio of the dataset. If the positive and negative sample ratio is not equal to 1:1, adjust the data;

[0055] Perform data augmentation on the adjusted original audio data to generate a forged audio dataset, where the forged audio dataset contains real and forged audio.

[0056] Specifically, the process of Step 1 is as follows:

[0057] Confirm the positive and negative sample distribution ratio of the dataset. If the distribution ratio of the positive and negative samples is not equal to 1:1, adjust the data;

[0058] Use the augmentation method to perform data augmentation on the audio signal, and the audio augmentation result is as Figure 1 shown.

[0059] As an implementation method in this embodiment, the process of data augmentation for the adjusted original audio data includes: performing Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the adjusted original audio data.

[0060] Specifically, the augmentation methods include:

[0061] (1) Gaussian noise augmentation: Add Gaussian distribution noise with a mean of 0 and an adjustable standard deviation to the original audio signal to simulate background noise in the real environment.

[0062] (2) Waveform displacement: Translate the audio signal on the time axis without changing the frequency content of the signal to simulate small changes in the signal over time, such as small delays or advances of the speaker.

[0063] (3) Waveform stretching: Adjust the duration of the signal by changing the time interval between each sample point in the signal, change the playback speed of the audio signal without changing its pitch, and is used to simulate different speech rates.

[0064] (4) Pitch correction: Adjust the pitch by changing the fundamental frequency of the signal while keeping the relative positions of other frequency components unchanged, change the pitch of the audio signal without changing its playback speed, and is used to simulate speech with different tones.

[0065] As an implementation method in this embodiment, an audio detection model is constructed, and the audio detection model further includes: a Sinc layer, a residual block, a GRU layer, and a fully connected layer, where the GRU layer is a gated recurrent unit (GRU) layer.

[0066] Specifically, the process of step 2 is as follows:

[0067] First, the Sinc convolution directly processes the time series signal of the audio. The Sinc convolution is an interpretable convolution filter structure that can directly extract frame rate-related information from the original waveform in a deep learning framework through the characteristics of a parameterized band-pass filter and is located in the first layer of the network. This structure provides an alternative to traditional spectrum extraction methods, such as MFCC, perceptual linear prediction, filter bank extraction, etc.

[0068] For the audio signal, the Sinc convolution uses a predefined function g to extract audio signal features, where g contains only a few learnable variables , and the definition of the convolution operation based on the Sinc function is as follows:

[0069] (1)

[0070] where is the speech signal, n is the length of the speech data, is the output of the filter. The g function is defined in the form of a rectangular band-pass filter, and its frequency-domain characteristics are as follows:

[0071] (2)

[0072] where f represents different frequency components in the audio signal, and are the low and high cut-off frequencies respectively, and both are learnable parameters, is the rectangular function in the amplitude-frequency domain and is defined as follows:

[0073] (3)

[0074] After Fourier transform, it is converted into the time-domain expression form as follows:

[0075] (4)

[0076] where The function is defined as .

[0077] Finally, in order to smooth the truncation characteristics of the g function, a window function is multiplied on the basis of the g function. Here, the Hamming Window is used, and the form is as follows:

[0078] (5)

[0079] (6)

[0080] represents the new function obtained after multiplying by the Hamming window function, represents the Hamming window function, represents the length of the Hamming window function, that is, the number of samples covered by the window. The residual block is further used to process the input features to extract frame-level features. The core idea of the residual block is to introduce skip connections, allowing signals in the network to bypass one or more layers, as shown in Figure 2 .

[0081] In the most basic form, a residual block can be expressed as:

[0082] (7)

[0083] where, is the audio feature extracted after convolution operation, is the residual function, representing the residual mapping to be learned, represents the set of learnable parameters in the residual function, and c is the output feature.

[0084] The residual block enables the network to only learn the difference between the input and the output, rather than the complete output. This setting simplifies the learning task because learning the difference is usually easier than learning the complete output target. And the skip connection allows the input features to be directly passed to the subsequent layers, which not only helps the flow of gradients but also ensures the effective transmission of feature information in the network, avoiding the gradual decrease of features in deep networks.

[0085] Using the feature scaling mapping to enhance the expressive power of the output features of the residual block, first let be the output features of the residual block, that is , where is the sequence length in time, is the number of filters. First, perform global average pooling on the time axis, then perform feed-forward through a fully connected layer, and then activate through Sigmoid to derive the scale vector for performing the feature scaling mapping FMS. By representing the scale vector as , then a scaled feature is designed. The feature scaling is as shown in Figure 3 to perform the following additive feature scaling on the features:

[0086] (8)

[0087] where, represents the output features of the residual block, represents the features after additive feature scaling, represents the corresponding element in the scale vector for feature scaling, . The features can also be scaled multiplicatively:

[0088] (9)

[0089] The two can also be executed in an interactive order, which is expressed as follows:

[0090] (10)

[0091] (11)

[0092] Using contrastive learning to maximize the similarity between the features of real audio data and the features of forged audio data, while minimizing the similarity between the features of the same class of data, that is, the similarity between the features of real sample data or the similarity between the features of forged audio data.

[0093] Model construction: The RawNet2-C model first directly processes the temporal signal features of audio using Sinc convolution, then further processes the input features using residual blocks to extract frame-level features. Then, feature scaling mapping is used to enhance the expressive power of the output features of the residual blocks. Finally, contrastive learning is used to maximize the similarity between the features of real audio data and forged audio data to enhance the discriminative ability of the model. The structure of the model is shown in Table 1.

[0094] Table 1 Structure Table of RawNet2-C Model

[0095] Layer Input 64,000 samples Output shape SincNet layer Convolution (129, 1, 128), max pooling (3), normalization, and LeakyReLU (21290,128) Residual block × 2 Normalization and LeakyReLU, convolution (3, 1, 128), normalization and LeakyReLU, convolution (3, 1, 128), max pooling (3), FMS (2365,128) Residual block × 4 Normalization and LeakyReLU, convolution (3, 1, 512), normalization and LeakyReLU, convolution (3, 1, 512), max pooling (3), FMS (29,512) GRU GRU(1024) (1024) Fully connected layer 1024 (1024) Output 1024 4

[0096] In Table 1, the functions and the meanings of the function parameters at each level in the table are described as follows:

[0097] Regarding the convolution (129, 1, 128) in the SincNet convolutional layer:

[0098] This is a one-dimensional convolution operation. The size of the convolution kernel is 129 (i.e., the length on the time axis is 129), the number of input channels is 1 (assuming the input data is a one-dimensional signal, such as a speech signal, etc.), and the number of output channels is 128. It is mainly used to extract local features of the input signal, and feature calculation is performed by sliding the convolution kernel on the input signal.

[0099] Regarding the max pooling (3) in the SincNet convolutional layer:

[0100] Max pooling is a downsampling operation. Here, the size of the pooling window is 3, that is, a maximum value is taken every 3 units. Its role is to reduce the dimension of the data, while retaining the maximum feature value in the local area, which helps to reduce the computational complexity and can prevent overfitting to a certain extent.

[0101] Regarding the normalization and LeakyReLU in the SincNet convolutional layer:

[0102] Normalization (such as batch normalization, etc.) can accelerate the training process of the network. By adjusting and scaling the input of the neurons, the input of each layer has zero mean and unit variance, making the training of the network more stable.

[0103] LeakyReLU is an activation function. It solves the problem that the gradient of the traditional ReLU function is zero when the input is negative. When the input of LeakyReLU is negative, it outputs a very small value (determined by the negative slope parameter), so that the neuron can also have a small gradient when the input is negative, avoiding the "death" of the neuron.

[0104] Regarding the normalization and LeakyReLU in the residual block × 2:

[0105] The same operation is performed, and the input is normalized and activated at the beginning of the residual block.

[0106] For the convolution (3, 1, 128) in the residual block × 2:

[0107] The convolution kernel size is 3, the number of input channels is 1, and the number of output channels is 128. The features are continuously extracted and transformed.

[0108] For the normalization and LeakyReLU in the residual block × 2:

[0109] The features obtained after convolution are mapped and normalized.

[0110] For the convolution (3, 1, 128) in the residual block × 2:

[0111] Convolution operation is performed again to further extract deeper features.

[0112] For the max pooling (3) in the residual block × 2:

[0113] Downsampling operation is performed to reduce the data dimension and retain key features.

[0114] For the FMS in the residual block × 2:

[0115] Feature scaling is performed.

[0116] For the normalization and LeakyReLU in the residual block × 4:

[0117] The features after convolution are mapped and normalized.

[0118] For the convolution (3, 1, 512) in the residual block × 4:

[0119] The convolution kernel size is 3, the number of input channels is 1, and the number of output channels is 512, which is used to extract more complex features. As the number of channels increases, the feature expression ability is enhanced.

[0120] For the normalization and LeakyReLU in the residual block × 4:

[0121] The features after convolution are mapped and normalized.

[0122] For the convolution (3, 1, 512) in the residual block × 4:

[0123] Convolution operation is continued to deepen the feature extraction level.

[0124] For the max pooling (3) in the residual block × 4:

[0125] Reduce the data dimension.

[0126] For the FMS in the residual block × 4:

[0127] Perform feature scaling.

[0128] For the GRU(1024) in the GRU (Gated Recurrent Unit):

[0129] GRU is a recurrent neural network structure used to process sequential data. Here, 1024 represents the dimension of the GRU hidden units. It controls the flow of information through update gates and reset gates, and can better handle long-term dependencies in the sequence, passing the information from previous time steps to subsequent time steps.

[0130] For the 1024 in the fully connected layer:

[0131] The fully connected layer performs a linear transformation on the previously extracted features (with a dimension of 1024). Each neuron is connected to all neurons in the previous layer, and is used to map the features to a new space for further tasks such as classification or regression.

[0132] For the output layer:

[0133] The dimension of the output layer is 4. This may be a four-class classification problem. Each output unit corresponds to the probability of a class (if it is a classification task and uses the softmax activation function), or represents four different output features (if it is other tasks such as regression).

[0134] As an implementation manner in this embodiment, based on the forged audio dataset, the audio detection model is trained in the first stage. The number of training epochs in the first stage is N. After the training epochs are completed, the first stage of training ends. The process of the first stage of training includes:

[0135] Train the model based on the cross-entropy loss function.

[0136] Specifically, the process of step 3 includes:

[0137] First, without using contrastive learning, perform the first stage of training on RawNet2-C to enable the model to have preliminary audio anti-forgery capabilities. In the first stage, its loss function is as follows:

[0138] (12)

[0139] Where, is the cross-entropy loss function, y is the label. For real audio data, the label is 1, and for forged audio data, the label is 0. p is the predicted probability that the input audio data is real data.

[0140] As an implementation method of this embodiment, after completing the first stage of training, the process of using the contrastive learning model to perform the second stage of training includes:

[0141] The second stage of training is performed based on the cross entropy loss function and the contrastive learning loss function.

[0142] Specifically, the process of step 4 includes:

[0143] The model is trained in the second phase using contrastive learning to improve its discrimination ability. The second phase uses a speech-enhanced dataset for training. Since the model already has basic recognition capabilities after the first round of training, the number of training rounds in the second phase is relatively small, with 30 rounds trained here. The convergence process is relatively stable in the early stages, but will fluctuate later, and finally the model with the best effect on the validation set is selected.

[0144] The loss function of contrastive learning is:

[0145] (13)

[0146] in, Represents the loss function of contrastive learning. The value of k depends on whether the audio data is of the same type. If it is the same type of audio data, k is 1, otherwise k is 0. and Respectively Features of audio data of different categories and the same category, r is a constant, controlling the examples between similar samples, Represents the similarity function, which is used to measure the similarity between two features, and P represents the total number of samples during training.

[0147] Finally, in the second stage, the loss function of the model is as follows:

[0148] (14)

[0149] in and are weights, among which Mainly used for the model to learn forged features. It is used to narrow the distance between samples with the same features, for example, to make samples with forged features more aggregated in the feature space to facilitate forgery classification. Therefore, it is hoped that the model can achieve better forgery detection results through contrastive learning in the second stage. Smaller, Larger, used in experiments is 0.1, When it is 0.45, the effect is better. Represents the total loss function.

[0150] Data in actual experiments:

[0151] (1) Select experimental data:

[0152] In this experiment, ASVspoof2019 LA was used as the dataset. There were approximately 50,000 pieces of data in the training set, about 2,500 pieces of real audio data, and about 48,000 pieces of forged audio data. To meet the needs of contrastive learning, the real audio in this experiment was data-augmented. After augmentation, the real data was approximately 54,000 pieces, basically meeting the requirement of a sample ratio of 1:1.

[0153] (2) Experimental results:

[0154] This experiment was tested on the test set of ASVspoof2019 LA. The test set contained approximately 140,000 audio samples, including about 7,000 pieces of real audio data and about 133,000 pieces of forged audio data. The discrimination accuracy rate of the audio data after testing was 95.56%.

[0155] Based on this, an audio forgery detection method based on contrastive learning of audio latent features provided by an embodiment of the present invention first performs data augmentation on the original audio data to generate a forged audio dataset, and then constructs an audio detection model. The audio detection model includes a contrastive learning model. Then, the audio detection model is trained in the first stage based on the forged audio dataset. After completing the first stage of training, the contrastive learning model is used for the second stage of training. Finally, audio forgery detection is performed on the audio based on the audio detection model that has completed the first stage of training and the second stage of training. In the ASVspoof2019 LA test set of this application, the detection accuracy rate of the model for 140,000 pieces of audio reached 95.56%, which was significantly better than traditional methods; by simulating diverse noises and speech variations through data augmentation, the model could adapt to complex real scenarios (such as background noise, speed / pitch differences); by combining Sinc convolution and residual blocks, deep latent features were directly extracted from the original waveform, avoiding the limitations of artificial feature design; the two-stage training strategy effectively balanced the classification and feature discrimination tasks, enhancing the model's sensitivity to subtle forgery traces.

[0156] Embodiment 2

[0157] In this embodiment, a computer terminal device is provided, including:

[0158] One or more processors;

[0159] A memory, coupled to the processor, for storing one or more programs;

[0160] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.

[0161] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the methods in the above embodiments are implemented.

[0162] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the methods in the above embodiments.

[0163] The above program can run in a processor or can also be stored in a memory (or referred to as a computer-readable medium). The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0164] These computer programs can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 The steps corresponding to different steps can be implemented by different modules.

[0165] In this embodiment, such a device or system is provided. The system is called an audio forgery detection system based on audio latent feature contrast learning, including:

[0166] A data processing module for performing data augmentation on the original audio data to generate a forged audio dataset;

[0167] A model construction module for constructing an audio detection model, where the audio detection model includes a contrast learning model, a Sinc layer, a residual block, a GRU layer, and a fully connected layer;

[0168] A model training module, configured to perform a first-stage training on the audio detection model based on the forged audio dataset, and after completing the first-stage training, use a contrastive learning model for a second-stage training;

[0169] A detection module, configured to perform forged detection on an audio based on the audio detection model that has completed the first-stage training and the second-stage training.

[0170] As an implementation manner in this embodiment, the data processing module includes:

[0171] A data augmentation unit, configured to perform Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the original audio data;

[0172] A data distribution adjustment unit, configured to confirm the positive and negative sample distribution ratio of the dataset, and adjust the data when it is not equal to 1:1.

[0173] As an implementation manner in this embodiment, when the training module performs the first-stage training on the audio detection model based on the forged audio dataset, a cross-entropy loss function is used for training.

[0174] As an implementation manner in this embodiment, after completing the first-stage training, when the training module uses the contrastive learning model for the second-stage training, a cross-entropy loss function and a loss function of contrastive learning are used for training.

[0175] This system or device is used to implement the functions of the method in the above embodiment. Each module in this system or device corresponds to each step in the method. Those that have been described in the method will not be elaborated here.

[0176] Through the above implementation manner, the problem of audio forged detection based on contrastive learning of audio latent features in the related art is solved. By combining contrastive learning with a specific model architecture and a staged training strategy, the problem of insufficient discrimination ability of traditional audio forged detection methods in complex scenarios is solved.

[0177] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the technical field of the present application within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An audio forgery detection method based on contrastive learning of audio latent features, characterized in that It includes the following steps: Perform data augmentation on the original audio data to generate a forged audio dataset; Construct an audio detection model, where the audio detection model includes a contrastive learning model; Perform the first-stage training on the audio detection model based on the forged audio dataset; After completing the first-stage training, use the contrastive learning model for the second-stage training; Perform forged detection on the audio based on the audio detection model that has completed the first-stage training and the second-stage training.

2. The method according to claim 1, wherein The process of performing data augmentation on the original audio data to generate a forged audio dataset includes: Confirm the positive and negative sample distribution ratio of the dataset. If the positive and negative sample ratio is not equal to 1:1, adjust the data; Perform data augmentation on the adjusted original audio data to generate a forged audio dataset.

3. The method according to claim 1, characterized in that, The process of performing data augmentation on the adjusted original audio data includes: performing Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the adjusted original audio data.

4. The method according to claim 1, wherein Construct an audio detection model, and the audio detection model further includes: a Sinc layer, a residual block, a GRU layer, and a fully connected layer.

5. The method according to claim 1, characterized in that Perform the first-stage training on the audio detection model based on the forged audio dataset. The number of training epochs for the first-stage training is N. After the training epochs are completed, the first-stage training ends. The process of the first-stage training includes: Train the model based on the cross-entropy loss function.

6. The method according to claim 1, wherein After completing the first-stage training, use The process of using the contrastive learning model for the second-stage training includes: Perform the second-stage training based on the cross-entropy loss function and the loss function of contrastive learning.

7. An audio forgery detection system based on contrastive learning of audio latent features, characterized in that, The system includes: A data processing module for performing data augmentation on the original audio data to generate a forged audio dataset; A model construction module for constructing an audio detection model, where the audio detection model includes a contrastive learning model, a Sinc layer, a residual block, a GRU layer, and a fully connected layer; A model training module for performing the first-stage training on the audio detection model based on the forged audio dataset, and after completing the first-stage training, using the contrastive learning model for the second-stage training; A detection module for performing forged detection on the audio based on the audio detection model that has completed the first-stage training and the second-stage training.

8. The system according to claim 7, wherein The data processing module includes: A data augmentation unit for performing Gaussian noise augmentation, waveform displacement, waveform stretching, and pitch correction on the original audio data; A data distribution adjustment unit for confirming the positive and negative sample distribution ratio of the dataset and adjusting the data when it is not equal to 1:

1.

9. A computer terminal device, characterized in that, It includes: One or more processors; A memory coupled to the processor for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the audio forged detection method based on contrastive learning of audio latent features as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio forged detection method based on contrastive learning of audio latent features as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Generation method of training sample set and storage medium

    CN115862607A

  • Multi-task learning-based few-sample named entity recognition method and device and medium

    CN116644755A

  • Speech enhancement model joint training method based on frequency sub-band

    CN118248159A

  • Voice spoofing detection method based on feature-enhanced attention mechanism

    CN118298832A

  • Deep forgery detection method, device and equipment based on cross-modal mask modeling

    CN119202989A

Cited By

  • Multi-modal forged video detection method based on multi-head addition cross attention mechanism

    CN120635786A

  • Continuous forged voice detection method and system based on multi-expert confrontation disturbance

    CN121506150A