Audio authentic identification method, related device, equipment and storage medium

By extracting and classifying audio features and training them in combination with the annotation type of pseudo-audio features, the problem of insufficient generalization and accuracy of existing audio pseudo-audio verification techniques is solved, and more efficient audio pseudo-audio verification is achieved.

CN120236604APending Publication Date: 2025-07-01HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271119.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing audio pseudo-recognition technology is insufficient in generalization and accuracy, making it difficult to effectively distinguish between real audio and fake audio.

Method used

By extracting features based on target audio, audio features are obtained, and these features are classified and predicted using the audio pseudo-identification model to identify audio types. The audio pseudo-audio model is trained with the annotation type of pseudo-audio feature as the target. The annotation type is fake audio, and the pseudo-audio feature is sampled by several audio feature distributions representing different audio forgery types.

Benefits of technology

Improve the generalization and accuracy of audio pseudo-identification, prevent model overfitting, and force the model to learn potential representations of fake audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236604A_ABST
    Figure CN120236604A_ABST
Patent Text Reader

Abstract

The invention discloses an audio authentication method, a related device, equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction based on a target audio, and obtaining the audio features of the target audio; performing classification prediction on the audio features based on an audio authentic identification model to obtain an identification type of the target audio; wherein the identification type is real audio or forged audio, the audio authentic identification model is obtained by training by taking the labeling type of the pseudo audio features as a target, the labeling type is forged audio, and the pseudo audio features are obtained by sampling a plurality of audio feature distributions representing different audio forged types. According to the scheme, generalization and accuracy of audio authentic identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to an audio forgery detection method and related devices, equipment, and storage media. Background Art

[0002] Currently, the generation technology of audio forgery is getting better and better, even reaching the level of being indistinguishable from the real one by human senses.

[0003] Traditional audio forgery detection methods need to set different strategies and disposal methods for different scenarios respectively, with weak generalization ability, and it is difficult to guarantee the accuracy of audio forgery detection. In view of this, how to improve the generalization and accuracy of audio forgery detection has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide an audio forgery detection method and related devices, equipment, and storage media, which can improve the generalization and accuracy of audio forgery detection.

[0005] To solve the above technical problem, a first aspect of the present application provides an audio forgery detection method, including: extracting features based on a target audio to obtain audio features of the target audio; classifying and predicting the audio features based on an audio forgery detection model to obtain an identification type of the target audio; wherein, the identification type is a real audio or a forged audio, and the audio forgery detection model is trained with the annotation type of forged audio features as the target, the annotation type is a forged audio, and the forged audio features are sampled from a plurality of audio feature distributions representing different audio forgery types.

[0006] To solve the above technical problem, a second aspect of the present application provides an audio forgery detection device, including: a feature extraction module and a classification prediction module, the feature extraction module is configured to extract features based on a target audio to obtain audio features of the target audio; the classification prediction module is configured to classify and predict the audio features based on an audio forgery detection model to obtain an identification type of the target audio; wherein, the identification type is a real audio or a forged audio, and the audio forgery detection model is trained with the annotation type of forged audio features as the target, the annotation type is a forged audio, and the forged audio features are sampled from a plurality of audio feature distributions representing different audio forgery types.

[0007] To solve the above technical problem, a third aspect of the present application provides an electronic device, at least including a memory and a processor coupled to each other, the memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the audio forgery detection method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the audio anti-counterfeiting method in the first aspect above.

[0009] In the above solution, feature extraction is performed based on the target audio to obtain the audio features of the target audio, and the audio features are classified and predicted based on the audio anti-counterfeiting model to obtain the identification type of the target audio, and the identification type is a genuine audio or a forged audio. The audio anti-counterfeiting model is trained with the annotation type of the forged audio features as the target, and the annotation type is a forged audio. The forged audio features are sampled from several audio feature distributions representing different audio forgery types. Therefore, on the one hand, by sampling several audio feature distributions of different audio forgery types, it is possible to simulate as many as possible the audio forgery types that have appeared and those that have not appeared. Especially under the training samples of limited forgery types, it is possible to prevent the model from overfitting as much as possible, which helps to improve the generalization of the audio anti-counterfeiting model. On the other hand, by training with the annotation type (i.e., forged audio) of the forged audio features as the target, it is possible to force the audio anti-counterfeiting model to learn the potential representation of the forged audio, which helps to improve the accuracy of the audio anti-counterfeiting model. Therefore, the generalization and accuracy of audio anti-counterfeiting can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the audio anti-counterfeiting method of the present application;

[0011] Figure 2a is a schematic diagram of the process of an embodiment of training the audio anti-counterfeiting model of the present application;

[0012] Figure 2b is a schematic diagram of the process of another embodiment of training the audio anti-counterfeiting model of the present application;

[0013] Figure 3 is a schematic framework diagram of an embodiment of the audio anti-counterfeiting device of the present application;

[0014] Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0015] Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following will describe the solutions of the embodiments of the present application in detail with reference to the accompanying drawings of the specification.

[0017] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0018] The terms "system" and "network" are often used interchangeably in this document. The term " / or" in this document is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this document generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "plurality" in this document means two or more than two.

[0019] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the audio forgery detection method of this application.

[0020] Specifically, it may include the following steps:

[0021] Step S11: Extract features based on the target audio to obtain the audio features of the target audio.

[0022] In the embodiments of the present disclosure, the actual type of the target audio may be a real audio, that is, the target audio may be audio data actually collected; or, the actual type of the target audio may also be a forged audio, that is, the target audio may be audio data obtained by means such as audio synthesis. The actual type of the target audio is not limited herein.

[0023] In one implementation scenario, as a possible implementation example, the features of the target audio may be extracted based on manually defined parameters to obtain the audio features of the target audio. That is to say, the audio features of the target audio may be artificial audio features. Exemplarily, the audio features of the target audio may include but are not limited to: Mel-frequency cepstral coefficients, linear frequency cepstral coefficients, constant-Q cepstral coefficients, group delay cepstral coefficients based on phase spectrum coefficients, orthogonal phase information, etc. The specific type of the artificial audio features is not limited herein.

[0024] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, the features of the target audio may also be extracted based on a feature extraction model to obtain the audio features of the target audio. For the convenience of distinguishing from the foregoing artificial audio features, in other words, the audio features of the target audio may be model audio features. Exemplarily, the features of the target audio may be extracted by a feature extraction model such as wave2vector to obtain the model audio features. The specific type of the feature extraction model is not limited herein.

[0025] In yet another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, for the case where the input audio is the target audio, on the one hand, feature extraction can be performed on the input audio based on manually defined parameters to obtain artificial audio features, and on the other hand, feature extraction can be performed on the input audio based on a feature extraction model to obtain model audio features. Then, based on the artificial audio features and the model audio features, fused audio features can be obtained, so as to process the fused audio features based on a time-domain attention mechanism to obtain first attention features, and process the fused audio features based on a frequency-domain attention mechanism to obtain second attention features. Then, the output audio features, that is, the audio features of the target audio, can be obtained by fusing the first attention features and the second attention features. In the above manner, by combining artificial audio features and model audio features to obtain fused audio features, and processing the fused audio features through a time-frequency attention mechanism to obtain the audio features of the target audio, it is possible to use artificial audio features to make up for the possible performance deficiencies caused by the feature extraction model being affected by training data, and use the feature extraction model to be able to learn more essential feature information (especially forged information) to make up for the deficiency that artificial audio features are difficult to take into account different audio frequency distributions, so that the artificial audio features and the model audio features complement each other, and the time-frequency attention mechanism can further strengthen the representation learning to better obtain the essential forgery traces.

[0026] In a specific implementation scenario, for artificial audio features, specifically, operations such as frame splitting, windowing, and Fourier transform can be performed on the target audio to obtain artificial audio features. For the convenience of description, the artificial audio features can be denoted as where represents the d-dimensional feature representation of the t-th audio frame in the target audio, and T represents the total number of frames in the target audio. Similarly, the model audio features can be denoted as where represents the -dimensional feature representation of the t-th audio frame in the target audio. It should be noted that the specific values of the above feature dimensions d and are not limited herein.

[0027] In a specific implementation scenario, after obtaining the artificial audio features and the model audio features, the two can be fused to obtain fused audio features. Exemplarily, the artificial audio features and the model audio features can be concatenated to obtain fused audio features. Still taking the foregoing artificial audio features and model audio features as an example, the feature dimension of the fused audio features after concatenation is the sum of d and . Of course, the above example is only one possible example of fusing artificial audio features and model audio features, and other possible fusion methods are not exemplified one by one herein.

[0028] In a specific implementation scenario, after obtaining the fused audio features, the time-domain attention mechanism and the frequency-domain attention mechanism can be used to process the fused audio features respectively. Specifically, the time-domain attention mechanism can be used to process the fused audio features to obtain the first attention weight, and the first attention weight can be used to weight the fused audio features to obtain the first attention feature; similarly, the frequency-domain attention mechanism can be used to process the fused audio features to obtain the second attention weight, and the second attention weight can be used to weight the fused audio features to obtain the second attention feature. It should be noted that for the specific processes of the time-domain attention mechanism and the frequency-domain attention mechanism, the technical details of the time-domain attention mechanism and the frequency-domain attention mechanism can be referred to and will not be elaborated here.

[0029] In a specific implementation scenario, after obtaining the first attention feature and the second attention feature, operations such as addition, averaging, and weighting can be performed on the first attention feature and the second attention feature to fuse the first attention feature and the second attention feature to obtain the output audio features, that is, the audio features of the target audio when the input audio is the target audio.

[0030] Step S12: Classify and predict the audio features based on the audio forgery detection model to obtain the identification type of the target audio.

[0031] In the embodiments of the present disclosure, the identification type is a genuine audio or a forged audio. The audio forgery detection model is trained with the annotation type of the forged audio features as the target. The annotation type is a forged audio, and the forged audio features are sampled from several audio feature distributions representing different audio forgery types. It should be noted that the audio forgery detection model can include but is not limited to: convolutional layers, fully connected layers, softmax, etc. The network structure of the audio forgery detection model is not limited here. In addition, as a possible example, when the audio forgery detection model classifies and predicts the audio features, it can specifically obtain the probability value of the target audio belonging to a genuine audio or a forged audio, and then the identification type of the target audio can be determined accordingly. For example, when the probability value of the target audio belonging to a genuine audio is greater than the probability value of the target audio belonging to a forged audio, it can be determined that the identification type of the target audio is a genuine audio; conversely, when the probability value of the target audio belonging to a genuine audio is less than the probability value of the target audio belonging to a forged audio, it can be determined that the identification type of the target audio is a forged audio.

[0032] In an implementation scenario, as a possible implementation example, before the audio anti-counterfeiting model is trained with the annotation type of the fake audio feature as the target, the audio anti-counterfeiting model can also be trained with the first sample type annotated by the first sample audio as the target first, and the first sample type is a forged audio or a genuine audio. Of course, the above implementation is only a possible implementation in the actual application process, that is, it does not mean that before the audio anti-counterfeiting model is trained with the annotation type of the fake audio feature as the target, the audio anti-counterfeiting model must be trained with the first sample type annotated by the first sample audio as the target first; in other words, before the audio anti-counterfeiting model is trained with the annotation type of the fake audio feature as the target, the audio anti-counterfeiting model may not be trained with the first sample type annotated by the first sample audio as the target first. For example, the audio anti-counterfeiting model can be directly trained with the annotation type of the fake audio feature as the target, which is not limited here. In the above method, before the audio anti-counterfeiting model is trained with the annotation type of the fake audio feature as the target, the audio anti-counterfeiting model is trained with the first sample type annotated by the first sample audio as the target first, and the first sample type is a forged audio or a genuine audio, which can force the audio anti-counterfeiting model to perform learning and training in different stages. In the first stage, it is trained with the first sample type annotated by the first sample audio as the target, and in the second stage, it is trained with the annotation type of the fake audio feature as the target, which helps to improve the learning efficiency of the audio anti-counterfeiting model.

[0033] In a specific implementation scenario, please refer to Figure 2a , Figure 2a which is a schematic diagram of the process of an embodiment of training the audio anti-counterfeiting model of the present application. As Figure 2a shown, when the audio anti-counterfeiting model is trained with the first sample type annotated by the first sample audio as the target, it can specifically extract features based on the first sample audio to obtain the first sample audio feature of the first sample audio. The specific process can refer to the relevant description of feature extraction of the input audio above, that is, in this case, the output audio feature is the first sample audio feature of the first sample audio. On this basis, the audio anti-counterfeiting model can classify and predict the first sample audio feature to obtain the first prediction type of the first sample audio, and adjust the network parameters of the audio anti-counterfeiting model based on the difference between the first sample type and the first prediction type of the first sample audio. Exemplarily, when the audio anti-counterfeiting model classifies and predicts the first sample audio feature to obtain the prediction probability values of the first sample audio belonging to the genuine audio and the forged audio, the cross-entropy and other loss functions can be used to calculate the above prediction probability values in combination with the first sample type annotated by the first sample audio to obtain the training loss of the audio anti-counterfeiting model in the first stage, and the network parameters of the audio anti-counterfeiting model can be adjusted based on the training loss.

[0034] In a specific implementation scenario, the audio forgery detection model can be iteratively trained with the first sample type labeled for the first sample audio until the training converges. At this time, the audio forgery detection model can be regarded as a basic forged audio detection model.

[0035] In an implementation scenario, to train the audio forgery detection model with the labeled type of forged audio features as the target, several audio feature distributions of different audio forgery types can be obtained first. Specifically, several audio sets belonging to different audio forgery types can be obtained first. Each audio within the audio set of the same audio forgery type has the corresponding audio forgery type. Then, for each audio set of an audio forgery type, fitting is performed based on each audio within the audio set to obtain the audio feature distribution of the corresponding audio forgery type. In the above manner, by first obtaining the audio sets of various audio forgery types and then performing fitting accordingly to obtain the audio feature distributions of various audio forgery types, the potential audio feature distributions of various audio forgery types can be extracted.

[0036] In a specific implementation scenario, an audio forgery type can correspondingly refer to an audio forgery algorithm, an audio forgery system, etc., which is not limited here. For example, taking a total of audio forgery algorithm 1, audio forgery algorithm 2,..., audio forgery algorithm M as an example, each audio forgery algorithm corresponds to an audio forgery type, that is, there are a total of M audio forgery types. Based on this, an audio set can be obtained using audio forgery algorithm i. Each audio within this audio set belongs to the i-th audio forgery type, where i is any integer value from 1 to M. Of course, the above example is only a possible example of audio forgery types, and other possible situations of audio forgery types are not listed one by one here.

[0037] In a specific implementation scenario, after obtaining the audio sets of various audio forgery types, specifically, a Gaussian mixture model or the like can be used to perform feature fitting on each audio within the audio set of the same audio forgery type to obtain the audio feature distribution of this audio forgery type. For the sake of description, the audio feature distributions of the above M audio forgery types can be respectively denoted as: where μ i , respectively represent the mean and variance of the i-th audio forgery type.

[0038] In an implementation scenario, after obtaining the audio feature distributions of various audio forgery types, the audio feature distributions can be sampled to obtain pseudo-audio features. Specifically, at least one audio feature distribution of an audio forgery type can be selected as the target feature distribution during each training, and then sampling can be performed based on each target feature distribution respectively to obtain pseudo-audio features. Exemplarily, still taking the aforementioned M audio forgery types as an example, during a certain training, the audio feature distribution of one of the audio forgery types can be selected for sampling, and a sub-audio feature obtained can be used as the pseudo-audio feature for this training; or, during a certain training, the audio feature distributions of multiple audio forgery types can be selected for sampling respectively to obtain the pseudo-audio features corresponding to the audio forgery types. In the above manner, at least one audio feature distribution of an audio forgery type is selected as the target feature distribution during each training, and sampling is performed based on each target feature distribution respectively to obtain pseudo-audio features, which can expand as many possible pseudo-audio features that have appeared and not appeared as possible, and helps the audio forgery detection model to learn forgery information accordingly.

[0039] In an implementation scenario, as a possible implementation example, when the audio forgery detection model is trained with the annotation type of the pseudo-audio feature as the target, it can classify and predict the pseudo-audio feature based on the audio forgery detection model to obtain the predicted type of the pseudo-audio feature, and then adjust the network parameters of the audio forgery detection model based on the difference between the annotation type and the predicted type of the pseudo-audio feature. Similar to the loss metric in the aforementioned first stage, the audio forgery detection model can classify and predict the pseudo-audio feature to obtain the predicted probability values of the pseudo-audio feature belonging to the real audio feature and the forged audio feature, and then the loss function such as cross entropy can be used to calculate the above predicted probability values in combination with the annotation type of the pseudo-audio feature to obtain the training loss when the audio forgery detection model is trained with the annotation type of the pseudo-audio feature as the target, so as to adjust the network parameters of the audio forgery detection model based on the training loss.

[0040] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, when the audio anti-forgery model is trained with the annotation type of the forged audio features as the target, it can also be trained with the target of reconstructing the sample artificial audio features and the sample model audio features based on the second sample audio features of the second sample audio, and the second sample type annotated by the second sample audio is real audio or forged audio. It should be noted that the second sample audio features are obtained by fusing the sample artificial audio features and the sample model audio features. The sample artificial audio features are obtained by extracting features from the second sample audio based on artificially defined parameters, and the sample model audio features are obtained by extracting features from the second sample audio based on the feature extraction model. Specifically, reference can be made to the foregoing related description of feature extraction of the input audio. In other words, when the input audio is the second sample audio, the output audio features obtained by extracting features from the input audio are the second sample audio features. In the above manner, by combining feature reconstruction to constrain the training process of the audio anti-forgery model, the audio anti-forgery model can be forced to fully learn the extraction of anti-forgery information representation.

[0041] In a specific implementation scenario, please refer to Figure 2b , Figure 2b which is a schematic diagram of the process of another embodiment of training the audio anti-forgery model of the present application. As Figure 2b shown, when reconstructing features with the sample artificial audio features as the target, specifically, the second sample audio features can be reconstructed based on the first reconstruction model to obtain the reconstructed artificial audio features, and the first reconstruction model can at least include a deconvolution network; similarly, when reconstructing features with the sample model audio features as the target, specifically, the second sample audio features can be reconstructed based on the second reconstruction model to obtain the reconstructed model audio features, and the second reconstruction model at least includes a deconvolution network. Of course, the above examples are only one possible example of feature reconstruction, and other possible ways of feature reconstruction will not be exemplified one by one here.

[0042] In a specific implementation scenario, when the audio anti-forgery model is trained based on the forged audio features and the second sample audio, it can specifically first extract features from the second sample audio to obtain the second sample audio features. Then, the audio anti-forgery model can be used to classify and predict the second sample audio features to obtain the second prediction type, and classify and predict the forged audio features to obtain the prediction type of the forged audio features. In addition, feature reconstruction is performed based on the second sample audio features to obtain the reconstructed artificial audio features and the reconstructed model audio features. Furthermore, the network parameters of the audio anti-forgery model can be adjusted based on the differences between the second sample type and the second prediction type, the differences between the labeled type and the prediction type of the forged audio features, the differences between the sample artificial audio features and the reconstructed artificial audio features, and the differences between the sample model audio features and the reconstructed model audio features. It should be noted that the differences between the second sample type and the second prediction type and the differences between the labeled type and the prediction type of the forged audio features can be measured based on loss functions such as cross-entropy. In addition, the differences between the sample artificial audio features and the reconstructed artificial audio features and the differences between the sample model audio features and the reconstructed model audio features can be measured based on loss functions such as mean square error. For ease of description, the training loss of the audio anti-forgery model can be characterized as:

[0043] L = L ce + αL re1 + βL re2

[0044] In the above formula, L ce represents the differences between the second sample type and the second prediction type and the differences between the labeled type and the prediction type of the forged audio features measured based on loss functions such as cross-entropy, and L re1 represents the differences between the sample artificial audio features and the reconstructed artificial audio features measured based on loss functions such as mean square error, and L re2 represents the differences between the sample model audio features and the reconstructed model audio features measured based on loss functions such as mean square error. In addition, α and β are the constraint coefficients of the training loss. Through the above training process, the original genuine and forged audio classification boundary can be further adjusted, enhancing the model's ability to identify data with unknown forgery algorithms and the model's generalization ability.

[0045] Based on the target audio, the above solution extracts features to obtain the audio features of the target audio, and classifies and predicts the audio features based on the audio anti-forgery model to obtain the identification type of the target audio, where the identification type is real audio or forged audio. The audio anti-forgery model is trained with the labeled type of the pseudo-audio features as the target, and the labeled type is forged audio. The pseudo-audio features are sampled from several audio feature distributions representing different audio forgery types. Therefore, on the one hand, by sampling several audio feature distributions of different audio forgery types, it is possible to simulate as much as possible the audio forgery types that have appeared and those that have not appeared. Especially under the limited training samples of forgery types, it is possible to prevent the model from overfitting as much as possible, which helps to improve the generalization of the audio anti-forgery model. On the other hand, by training with the labeled type of the pseudo-audio features (i.e., forged audio) as the target, it is possible to force the audio anti-forgery model to learn the potential representation of forged audio, which helps to improve the accuracy of the audio anti-forgery model. Therefore, the generalization and accuracy of audio anti-forgery can be improved.

[0046] Please refer to Figure 3 , Figure 3 which is a schematic framework diagram of an embodiment of the audio anti-forgery device of the present application. The audio anti-forgery device 30 includes: a feature extraction module 31 and a classification and prediction module 32. The feature extraction module 31 is configured to extract features based on the target audio to obtain the audio features of the target audio; the classification and prediction module 32 is configured to classify and predict the audio features based on the audio anti-forgery model to obtain the identification type of the target audio; wherein, the identification type is real audio or forged audio, and the audio anti-forgery model is trained with the labeled type of the pseudo-audio features as the target, the labeled type is forged audio, and the pseudo-audio features are sampled from several audio feature distributions representing different audio forgery types.

[0047] Based on the target audio, the audio anti-forgery device 30 in the above solution extracts features to obtain the audio features of the target audio, and classifies and predicts the audio features based on the audio anti-forgery model to obtain the identification type of the target audio, where the identification type is real audio or forged audio. The audio anti-forgery model is trained with the labeled type of the pseudo-audio features as the target, the labeled type is forged audio, and the pseudo-audio features are sampled from several audio feature distributions representing different audio forgery types. Therefore, on the one hand, by sampling several audio feature distributions of different audio forgery types, it is possible to simulate as much as possible the audio forgery types that have appeared and those that have not appeared. Especially under the limited training samples of forgery types, it is possible to prevent the model from overfitting as much as possible, which helps to improve the generalization of the audio anti-forgery model. On the other hand, by training with the labeled type of the pseudo-audio features (i.e., forged audio) as the target, it is possible to force the audio anti-forgery model to learn the potential representation of forged audio, which helps to improve the accuracy of the audio anti-forgery model. Therefore, the generalization and accuracy of audio anti-forgery can be improved.

[0048] In some disclosed embodiments, the audio anti-forgery device 30 includes a set acquisition module for acquiring a plurality of audio sets respectively belonging to different audio forgery types; wherein, each audio within an audio set of the same audio forgery type has the corresponding audio forgery type; the audio anti-forgery device 30 includes a feature fitting module for, for each audio set of an audio forgery type, performing fitting based on each audio within the audio set to obtain the audio feature distribution of the corresponding audio forgery type.

[0049] In some disclosed embodiments, the audio anti-forgery device 30 includes a distribution selection module for selecting, during each training, the audio feature distributions of at least one audio forgery type as the target feature distributions respectively; the audio anti-forgery device 30 includes a feature sampling module for performing sampling respectively based on each target feature distribution to obtain pseudo-audio features.

[0050] In some disclosed embodiments, before the audio anti-forgery model is trained with the annotation type of the pseudo-audio features as the target, the audio anti-forgery model is first trained with the first sample type annotated by the first sample audio as the target, and the first sample type is a forged audio or a genuine audio.

[0051] In some disclosed embodiments, the audio anti-forgery device 30 includes a first extraction module for extracting features based on the first sample audio to obtain the first sample audio features of the first sample audio; the audio anti-forgery device 30 includes a first prediction module for classifying and predicting the first sample audio features based on the audio anti-forgery model to obtain the first prediction type of the first sample audio; the audio anti-forgery device 30 includes a first adjustment module for adjusting the network parameters of the audio anti-forgery model based on the difference between the first sample type and the first prediction type of the first sample audio.

[0052] In some disclosed embodiments, when the audio anti-forgery model is trained based on the pseudo-audio features, it is also trained with the goal of reconstructing the sample artificial audio features and the sample model audio features from the second sample audio features of the second sample audio, and the second sample type annotated by the second sample audio is a genuine audio or a forged audio; wherein, the second sample audio features are obtained by fusing the sample artificial audio features and the sample model audio features, the sample artificial audio features are obtained by extracting features from the second sample audio based on artificially defined parameters, and the sample model audio features are obtained by extracting features from the second sample audio based on a feature extraction model.

[0053] In some disclosed embodiments, the audio anti-forgery device 30 includes a second extraction module configured to extract features from the second sample audio to obtain second sample audio features; the audio anti-forgery device 30 includes a second prediction module configured to perform classification prediction on the second sample audio features based on the audio anti-forgery model to obtain a second prediction type, and perform classification prediction on the forged audio features based on the audio anti-forgery model to obtain the prediction type of the forged audio features; the audio anti-forgery device 30 includes a feature reconstruction module configured to reconstruct features based on the second sample audio features to obtain reconstructed artificial audio features and reconstructed model audio features; the audio anti-forgery device 30 includes a second adjustment module configured to adjust the network parameters of the audio anti-forgery model based on the difference between the second sample type and the second prediction type, the difference between the labeled type and the prediction type of the forged audio features, the difference between the sample artificial audio features and the reconstructed artificial audio features, and the difference between the sample model audio features and the reconstructed model audio features.

[0054] In some disclosed embodiments, the feature reconstruction module is specifically configured to reconstruct features of the second sample audio based on a first reconstruction model to obtain reconstructed artificial audio features; wherein, the first reconstruction model includes at least a deconvolution network.

[0055] In some disclosed embodiments, the feature reconstruction module is specifically configured to reconstruct features of the second sample audio based on a second reconstruction model to obtain reconstructed model audio features; wherein, the second reconstruction model includes at least a deconvolution network.

[0056] In some disclosed embodiments, the frequency anti-forgery device 30 includes an artificial extraction module configured to extract features from the input audio based on artificially defined parameters to obtain artificial audio features, the frequency anti-forgery device 30 includes a model extraction module configured to extract features from the input audio based on a feature extraction model to obtain model audio features; the frequency anti-forgery device 30 includes a first fusion module configured to obtain fused audio features based on the artificial audio features and the model audio features; the frequency anti-forgery device 30 includes a time-frequency attention module configured to process the fused audio features based on a time-domain attention mechanism to obtain a first attention feature, and process the fused audio features based on a frequency-domain attention mechanism to obtain a second attention feature; the frequency anti-forgery device 30 includes a second fusion module configured to fuse the first attention feature and the second attention feature to obtain output audio features; wherein, when the input audio is a target audio, the output audio features are the audio features of the target audio, when the input audio is a first sample audio, the output audio features are the first sample audio features of the first sample audio, and when the input audio is a second sample audio, the output audio features are the second sample audio features of the second sample audio.

[0057] Please refer to Figure 4 , Figure 4It is a schematic diagram of the framework of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-mentioned embodiments of the audio anti-counterfeiting method. For details, reference can be made to the foregoing disclosed embodiments, which will not be elaborated here. As a possible example, the electronic device 40 may include, but is not limited to, a smart phone, a tablet computer, a server, etc. The specific type of the electronic device 40 is not limited herein.

[0058] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above-mentioned embodiments of the audio anti-counterfeiting method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.

[0059] In the above solution, the electronic device 40 extracts features based on the target audio to obtain the audio features of the target audio, and classifies and predicts the audio features based on the audio anti-counterfeiting model to obtain the identification type of the target audio, and the identification type is a genuine audio or a forged audio. The audio anti-counterfeiting model is trained with the labeled type of the forged audio features as the target. The labeled type is a forged audio, and the forged audio features are sampled from several audio feature distributions representing different audio forgery types. Therefore, on the one hand, by sampling several audio feature distributions of different audio forgery types, it is possible to simulate as much as possible the audio forgery types that have appeared and those that have not appeared. Especially under the training samples of limited forgery types, it is possible to prevent the model from overfitting as much as possible, which helps to improve the generalization of the audio anti-counterfeiting model. On the other hand, by training with the labeled type (i.e., forged audio) of the forged audio features as the target, it is possible to force the audio anti-counterfeiting model to learn the potential representation of the forged audio, which helps to improve the accuracy of the audio anti-counterfeiting model. Therefore, the generalization and accuracy of audio anti-counterfeiting can be improved.

[0060] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above-described embodiments of the audio anti-counterfeiting method.

[0061] In the above solution, the computer-readable storage medium 50 extracts features based on the target audio to obtain the audio features of the target audio, and classifies and predicts the audio features based on the audio anti-counterfeiting model to obtain the identification type of the target audio, and the identification type is a genuine audio or a forged audio. The audio anti-counterfeiting model is trained with the annotation type of the forged audio features as the target, and the annotation type is a forged audio. The forged audio features are sampled from several audio feature distributions representing different audio forgery types. Therefore, on the one hand, by sampling several audio feature distributions of different audio forgery types, it is possible to simulate as much as possible the audio forgery types that have appeared and those that have not appeared. Especially under the training samples of limited forgery types, it is possible to prevent the model from overfitting as much as possible, which helps to improve the generalization of the audio anti-counterfeiting model. On the other hand, by training with the annotation type (i.e., forged audio) of the forged audio features as the target, it is possible to force the audio anti-counterfeiting model to learn the potential representation of the forged audio, which helps to improve the accuracy of the audio anti-counterfeiting model. Therefore, the generalization and accuracy of audio anti-counterfeiting can be improved.

[0062] In some embodiments, the functions or modules included in the device provided by the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0063] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0064] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in the form of electricity, machinery or others.

[0065] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0066] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0067] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0068] If the technical solution of the present application involves personal information, before the product applying the technical solution of the present application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of the present application involves sensitive personal information, before the product applying the technical solution of the present application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is considered consent to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. An audio authentication method, characterized in that: include: Perform feature extraction based on the target audio to obtain audio features of the target audio; The audio features are classified and predicted based on an audio authentication model to obtain an authentication type of the target audio; wherein the authentication type is real audio or forged audio, the audio authentication model is trained with the annotated type of the pseudo audio features as the target, the annotated type is forged audio, and the pseudo audio features are sampled from a number of audio feature distributions representing different audio forgery types.

2. The method according to claim 1, characterized in that The step of acquiring the distribution of the plurality of audio features comprises: Acquire a plurality of audio sets respectively belonging to different audio forgery types; wherein each audio in the audio set of the same audio forgery type has a corresponding audio forgery type; For each audio set of the audio forgery type, fitting is performed based on the individual audios in the audio set to obtain an audio feature distribution corresponding to the audio forgery type.

3. The method according to claim 1, characterized in that The step of acquiring the pseudo audio feature comprises: During each training, at least one audio feature distribution of the audio forgery type is selected as the target feature distribution; Sampling is performed based on each of the target feature distributions to obtain the pseudo audio features.

4. The method according to claim 1, characterized in that Before the audio authentication model is trained with the annotated type of the pseudo audio feature as a target, the audio authentication model is first trained with a first sample type annotated by a first sample audio as a target, and the first sample type is forged audio or real audio.

5. The method according to claim 4, characterized in that The step of training the audio authentication model with the first sample type marked by the first sample audio as a target includes: Perform feature extraction based on the first sample audio to obtain a first sample audio feature of the first sample audio; Performing classification prediction on the first sample audio feature based on the audio authentication model to obtain a first prediction type of the first sample audio; Based on the difference between the first sample type and the first predicted type of the first sample audio, the network parameters of the audio authentication model are adjusted.

6. The method according to claim 1, characterized in that When the audio forgery detection model is trained based on the pseudo audio feature, the model is also trained with the goal of reconstructing sample artificial audio features and sample model audio features based on the second sample audio feature of the second sample audio, and the second sample type marked by the second sample audio is real audio or forged audio; Among them, the second sample audio feature is obtained by fusing the sample artificial audio feature and the sample model audio feature, the sample artificial audio feature is obtained by extracting features from the second sample audio based on artificially defined parameters, and the sample model audio feature is obtained by extracting features from the second sample audio based on a feature extraction model.

7. The method according to claim 6, characterized in that The step of training the audio authentication model based on the pseudo audio feature and the second sample audio comprises: Perform feature extraction based on the second sample audio to obtain the second sample audio feature; Classify and predict the second sample audio feature based on the audio authentication model to obtain a second prediction type, classify and predict the pseudo audio feature based on the audio authentication model to obtain a prediction type of the pseudo audio feature, and reconstruct features based on the second sample audio feature to obtain reconstructed artificial audio features and reconstructed model audio features; Based on the difference between the second sample type and the second predicted type, the difference between the labeled type and the predicted type of the pseudo audio feature, the difference between the sample artificial audio feature and the reconstructed artificial audio feature, and the difference between the sample model audio feature and the reconstructed model audio feature, the network parameters of the audio fake identification model are adjusted.

8. The method according to claim 6, characterized in that The step of reconstructing features with the sample artificial audio features as the target includes: Reconstructing the second sample audio feature based on a first reconstruction model to obtain a reconstructed artificial audio feature; wherein the first reconstruction model at least includes a deconvolution network; And / or, the step of reconstructing features with the sample model audio features as a target includes: The second sample audio feature is reconstructed based on a second reconstruction model to obtain a reconstructed model audio feature; wherein the second reconstruction model at least includes a deconvolution network.

9. The method according to any one of claims 1 to 8, characterized in that: The feature extraction step comprises: Extracting features from the input audio based on manually defined parameters to obtain artificial audio features, and extracting features from the input audio based on a feature extraction model to obtain model audio features; Based on the artificial audio features and the model audio features, obtaining fused audio features; Processing the fused audio features based on a time-domain attention mechanism to obtain a first attention feature, and processing the fused audio features based on a frequency-domain attention mechanism to obtain a second attention feature; Obtaining an output audio feature based on fusing the first attention feature and the second attention feature; Among them, when the input audio is the target audio, the output audio feature is the audio feature of the target audio, when the input audio is the first sample audio, the output audio feature is the first sample audio feature of the first sample audio, and when the input audio is the second sample audio, the output audio feature is the second sample audio feature of the second sample audio.

10. An audio authentication device, characterized in that: include: A feature extraction module, used to extract features based on the target audio to obtain audio features of the target audio; A classification prediction module is used to classify and predict the audio features based on an audio authentication model to obtain an authentication type of the target audio; wherein the authentication type is real audio or forged audio, the audio authentication model is trained with the annotated type of the pseudo audio features as the target, the annotated type is forged audio, and the pseudo audio features are sampled from a plurality of audio feature distributions representing different audio forgery types.

11. An electronic device, characterized in that: The device at least comprises a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the audio authentication method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the audio authentication method according to any one of claims 1 to 9.