Audio Copy Move Deepfake Detection Method

By combining the weighted fusion of Mel spectrogram and Hilbert yellow spectrogram, an audio classification network is built, which solves the problems of insufficient accuracy of audio copy mobile forgery detection and difficulty in music detection in the prior art, and realizes high-precision detection of audio forgery.

CN118865987BActive Publication Date: 2025-07-11XIHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410867697.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-07-11
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The existing audio copy mobile forgery detection methods have insufficient detection accuracy and cannot accurately detect music audio. The real-time voice activity detection framework is prone to false alarms or missed reports, and the threshold selection is highly subjective.

Method used

The convolutional neural network model is used to combine the weighted fusion of Mel spectrogram and Hilbert yellow spectrogram to build an audio classification convolutional neural network. By extracting and predicting the fusion spectrum map, the detection of audio copy mobile forgery is achieved.

Benefits of technology

Improves detection accuracy for short and long audio, and can accurately detect copy-move forgery in songs and pure music, eliminates the limitations of real-time voice activity detection and provides higher detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118865987B_ABST
    Figure CN118865987B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio copy-move deepfake detection method, belonging to the field of digital audio technology, which includes the following steps: obtaining the audio to be detected, generating a Mel spectrogram and a Hilbert-Huang spectrogram based on the audio to be detected; performing weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram to obtain a fused spectrogram; constructing an audio classification convolutional neural network and setting network preset parameters; training the audio classification convolutional neural network to obtain a trained audio classification convolutional neural network; inputting the fused spectrogram into the trained audio classification convolutional neural network for audio copy-move deepfake detection to obtain an audio copy-move deepfake detection result. The present invention solves the problems of insufficient detection accuracy of existing audio copy-move forgery detection and the inability to detect music audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital audio, and particularly relates to an audio copy-move deepfake detection method. Background Art

[0002] With the development and transformation of digital audio technology, music, sounds, and voices are increasingly showing digital and networked characteristics. With the wide spread of digital audio in multimedia, the act of deliberately tampering with and forging audio has become increasingly serious, leading to a series of information security problems. Audio forgery and audio tampering refer to modifying or editing audio content through digital technology to deceive, mislead, and distort facts.

[0003] Audio copy-move forgery (ACMF) is a common forgery technique that modifies the semantics of the original audio by copying and pasting the voiced segments of the same audio source, and it is easy to implement. Due to the simplicity of the tampering operation and the tampering of the same audio source, copy-move forgery has become one of the most difficult tampering methods to detect. Therefore, it is necessary to use audio forensics technology to detect ACMF to protect the authenticity and credibility of audio content.

[0004] In the existing ACMF detection methods, if monitoring is based on a real-time voice activity detection (VAD) framework, it may mislabel non-speech parts as speech or speech parts as non-speech, resulting in false alarms or missed detections, making the detection results inaccurate and the performance degraded. Secondly, VAD is mainly used to detect human speech signals and cannot detect non-speech signals such as music, and VAD tends to detect short-term speech activities and may not be able to accurately segment long speeches. In addition, almost all ACMF detection methods have a threshold selection for determining whether it is a forgery category. Threshold selection is a relatively subjective behavior, and the accuracy is still insufficient. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, an audio copy-move deepfake detection method provided by the present invention solves the problems of insufficient accuracy in detecting existing audio copy-move forgery and inability to detect music audio by combining a convolutional neural network model and spectrogram image fusion.

[0006] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0007] An audio copy-move deepfake detection method provided by the present invention includes the following steps:

[0008] S1. Obtain the audio to be detected, and generate a Mel spectrogram and a Hilbert-Huang spectrogram based on the audio to be detected;

[0009] S2. Weightedly fuse the Mel spectrogram and the Hilbert-Huang spectrogram to obtain a fused spectrogram;

[0010] S3. Construct an audio classification convolutional neural network and set the network preset parameters;

[0011] S4. Train the audio classification convolutional neural network to obtain a trained audio classification convolutional neural network;

[0012] S5. Input the fused spectrogram into the trained audio classification convolutional neural network for audio copy-move deepfake detection to obtain the audio copy-move deepfake detection result.

[0013] The beneficial effects of the present invention are as follows: An audio copy-move deepfake detection method provided by the present invention obtains a fused spectrogram with prominent energy features by weightedly fusing the Mel spectrogram and the Hilbert-Huang spectrogram generated from the audio to be detected, and constructs a spectrogram classification convolutional neural network for classifying and predicting the fused spectrogram to extract the features in the fused spectrogram and predict and classify the attack situation of the audio to be detected being attacked by audio copy-move deepfakes; through the fused spectrogram and the spectrogram classification convolutional neural network, the present invention can achieve accurate detection effects for both short audio and long audio, eliminates the limitations of the VAD method, and can accurately detect copy-move forgeries even for songs and pure music, and can be well applied to audio forensics.

[0014] Further, the S1 includes the following steps:

[0015] S11. Extract the frequency and several intrinsic modes of the audio to be detected;

[0016] The calculation expression of the intrinsic mode of the audio to be detected is as follows:

[0017]

[0018] where x(t) represents the signal to be detected, C i (t) represents the i-th intrinsic mode of the signal to be detected, R n (t) represents the residual term after the signal to be detected is decomposed, where i = 1, 2,..., n;

[0019] S12. Generate a Mel spectrogram according to the frequency of the audio to be detected;

[0020] The calculation expression of the Mel spectrogram is as follows:

[0021]

[0022] Among them, Mel(f) represents the Mel spectrogram, lg(·) represents the logarithmic calculation function with base 10, and f represents the frequency of the audio to be detected;

[0023] S13. Perform Hilbert transform on the intrinsic mode of the audio to be detected to obtain the Hilbert transform result of the intrinsic mode;

[0024] S14. Generate a modal transform signal based on the Hilbert transform result of the intrinsic mode;

[0025] The calculation expression of the modal transform signal is as follows:

[0026]

[0027] Among them, Z i (t) represents the modal transform signal, a i (t) represents the i-th intrinsic mode correlation signal, represents the transformation factor, d i (t) represents the Hilbert transform result of the i-th intrinsic mode of the signal to be detected;

[0028] S15. Use the modal transform signal to draw the Hilbert-Huang spectrogram.

[0029] The beneficial effect of adopting the above further solution is that the present invention provides a specific method for generating a Mel spectrogram and a Hilbert-Huang spectrogram based on the audio to be detected. By generating a Mel spectrogram and a Hilbert-Huang spectrogram for the audio to be detected, different audio information can be obtained, providing a basis for spectrogram fusion and target audio information extraction.

[0030] Further, the calculation expression of the fused spectrogram in S2 is as follows:

[0031]

[0032] Among them, F fusion (x, y) represents the pixel at the abscissa x and ordinate y on the fused spectrogram, ω1 represents the first fusion weight coefficient, F1(x, y) represents the pixel at the abscissa x and ordinate y on the Mel spectrogram, ω2 represents the second fusion weight coefficient, and F2(x, y) represents the pixel at the abscissa x and ordinate y on the Hilbert-Huang spectrogram.

[0033] The beneficial effect of adopting the above further solution is that the present invention provides a calculation method for the fused spectrogram. Through the first fusion weight coefficient and the second fusion weight coefficient, spectrogram fusion with adjustable emphasis on audio information can be realized, and effective audio information can be provided when dealing with known audio copy-move forgery attack types or unknown audio copy forgery attack types.

[0034] Furthermore, when the attack mode corresponding to audio copy-move forgery detection is additive noise, let the first fusion weight coefficient be 0.7 and the second fusion weight coefficient be 0.3, and perform weighted fusion of the Mel spectrogram and the Hilbert-Huang spectrogram;

[0035] When the attack mode corresponding to audio copy-move forgery detection is median filtering, let the first fusion weight coefficient be 0.8 and the second fusion weight coefficient be 0.2, and perform weighted fusion of the Mel spectrogram and the Hilbert-Huang spectrogram;

[0036] When the attack mode corresponding to audio copy-move forgery detection is compressed audio attack, let the first fusion weight coefficient be 0.9 and the second fusion weight coefficient be 0.1, and perform weighted fusion of the Mel spectrogram and the Hilbert-Huang spectrogram;

[0037] When the attack mode corresponding to audio copy-move forgery detection is unknown, let the first fusion weight coefficient be 0.8 and the second fusion weight coefficient be 0.2, and perform weighted fusion of the Mel spectrogram and the Hilbert-Huang spectrogram.

[0038] The beneficial effect of adopting the above further scheme is that the present invention respectively provides different attack types of known audio copy-move forgery and the attack type of unknown audio copy-move forgery, and when performing image fusion, the optimal combination of fusion weight coefficients can extract the corresponding most effective audio information for audio copy-move deep forgery detection according to different situations.

[0039] Furthermore, the audio classification convolutional neural network includes an image input layer, a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fully connected layer, and a softmax layer connected in sequence;

[0040] The first convolutional module includes a first convolutional layer, a batch normalization layer, a first max pooling layer, and a first activation layer connected in sequence;

[0041] The second convolutional module includes a second convolutional layer, a second max pooling layer, and a second activation layer connected in sequence;

[0042] The third convolutional module includes a third convolutional layer, a third activation layer, and a first average pooling layer connected in sequence;

[0043] The fourth convolutional module includes a fourth convolutional layer, a fourth activation layer, and a second average pooling layer connected in sequence.

[0044] The beneficial effects of adopting the above further solution are as follows: The present invention provides a network structure of an audio classification convolutional neural network. The effective extraction of global features is realized through the first convolutional module and the second convolutional module in the convolutional neural network. The local feature extraction of the fused audio information is realized through the third convolutional module and the fourth convolutional module. And the fully connected layer and the softmax layer provide a basis for feature integration and classification prediction.

[0045] Further, the S4 includes the following steps:

[0046] S41. Obtain a fused spectrogram training dataset;

[0047] S42. Input each fused spectrogram in the fused spectrogram training dataset into the image input layer;

[0048] S43. Use the first convolutional module to extract features from the fused spectrogram in the image input layer to obtain a first feature;

[0049] The calculation expression of the first feature is as follows:

[0050] F out1 =σ(Maxpooling(BN(C j1 (F fusion ))))

[0051] Where, F out1 represents the first feature, σ(·) represents the Relu activation function, Maxpooling(·) represents the max pooling function, BN(·) represents the batch normalization function, C j1 (·) represents the first convolutional function, F fusion represents the fused spectrogram. Among them, the convolutional kernel size in the first convolutional function is the first scale, and the first scale is 10×10;

[0052] S44. Use the second convolutional module to extract features from the texture features to obtain a second feature;

[0053] The calculation expression of the second feature is as follows:

[0054] F out2 =σ(Maxpooling(C j2 (F out1 )))

[0055] Where, F out2 represents the second feature, C j2 (·) represents the second convolutional function. Among them, the convolutional kernel size in the second convolutional function is the second scale, and the second scale is 7×7;

[0056] S45. Extract features from the second feature using the third convolutional module to obtain the third feature;

[0057] The calculation expression of the third feature is as follows:

[0058] F out3 = δ(Avgpooling(C j3 (F out2 )))

[0059] where F out3 represents the third feature, δ(·) represents the LeakyRelu activation function, Avgpooling(·) represents the average pooling function, and C j3 (·) represents the third convolutional function, where the size of the convolutional kernel in the third convolutional function is the third scale, and the third scale is 5×5;

[0060] S46. Extract features from the third feature using the fourth convolutional module to obtain the fourth feature;

[0061] The calculation expression of the fourth feature is as follows:

[0062] F out4 = δ(Avgpooling(C j4 (F out3 )))

[0063] where F out4 represents the fourth feature, and C j4 (·) represents the fourth convolutional function, where the size of the convolutional kernel in the fourth convolutional function is the fourth scale, and the fourth scale is 3×3;

[0064] S47. Use the fully connected layer and the softmax layer to perform feature integration and prediction classification on the fourth feature to obtain the classification result corresponding to the fused spectrogram;

[0065] S48. Adjust the preset parameters of the network until the consistency rate between the classification result corresponding to the fused spectrogram and its own category reaches the preset classification threshold, save the preset parameters of the network, complete the training of the audio classification convolutional neural network, and obtain the trained audio classification convolutional neural network.

[0066] The beneficial effects of adopting the above further solution are as follows: The present invention provides a method for training an audio classification convolutional neural network. By setting the convolutional kernel size from large to small, the receptive field of the convolutional layer changes from large to small, from global to local. The larger convolutional kernel can ignore the changes brought about by audio processing. After passing through the convolutional layer with a larger convolutional kernel, it is possible to better capture the tampered fragments in the spectrogram, providing a basis for correctly classifying the attacked and tampered audio. The larger convolutional kernel can also reduce the parameters of the entire audio classification convolutional neural network. While the convolutional kernel size changes from large to small, the number of filters is correspondingly increased to generate more texture-related feature maps, facilitating the correct classification of tampered audio after feature integration.

[0067] Other advantages of the present invention will be analyzed in more detail in the subsequent embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0069] Figure 1 It is a flowchart of the steps of a method for detecting audio copy-move deep fakes in an embodiment of the present invention.

[0070] Figure 2 It is a structural block diagram of an audio classification convolutional neural network in an embodiment of the present invention.

[0071] Figure 3(a) is a schematic diagram of the attack detection results after adding 20 dB Gaussian white noise under different weight combinations in an embodiment of the present invention.

[0072] Figure 3(b) is a schematic diagram of the attack detection results after compressing the audio at 32 kbps under different weight combinations in an embodiment of the present invention.

[0073] Figure 3(c) is a schematic diagram of the attack detection results after median filtering the audio under different weight combinations in an embodiment of the present invention.

[0074] Figure 3(d) is a schematic diagram of the detection results without attack under different weight combinations in an embodiment of the present invention.

[0075] Figure 4 It is a schematic diagram of the experimental results of the scheme comparison under different weight combinations and different post-processing attacks in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0077] As Figure 1 shown, in an embodiment of the present invention, the present invention provides an audio copy mobile deepfake detection method, including the following steps:

[0078] S1. Obtain the audio to be detected, and generate a Mel spectrogram and a Hilbert-Huang spectrogram based on the audio to be detected;

[0079] The S1 includes the following steps:

[0080] S11. Extract the frequency and several intrinsic modes of the audio to be detected;

[0081] The calculation expression of the intrinsic mode of the audio to be detected is as follows:

[0082]

[0083] where x(t) represents the signal to be detected, C i (t) represents the i-th intrinsic mode of the signal to be detected, and R n (t) represents the residual term after the decomposition of the signal to be detected, where i = 1, 2,..., n;

[0084] S12. Generate a Mel spectrogram according to the frequency of the audio to be detected;

[0085] The calculation expression of the Mel spectrogram is as follows:

[0086]

[0087] where Mel(f) represents the Mel spectrogram, lg(·) represents the logarithmic calculation function with base 10, and f represents the frequency of the audio to be detected;

[0088] S13. Perform a Hilbert transform on the intrinsic mode of the audio to be detected to obtain the Hilbert transform result of the intrinsic mode;

[0089] S14. Generate a modal transform signal based on the Hilbert transform result of the intrinsic mode;

[0090] The calculation expression of the modal transformation signal is as follows:

[0091]

[0092] Where Z i (t) represents the modal transformation signal, and a i (t) represents the i-th intrinsic mode correlation signal, represents the transformation factor, and d i (t) represents the Hilbert transform result of the i-th intrinsic mode of the signal to be detected;

[0093] S15. Draw the Hilbert-Huang spectrogram using the modal transformation signal.

[0094] S2. Perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram to obtain the fused spectrogram;

[0095] The calculation expression of the fused spectrogram in S2 is as follows:

[0096]

[0097] Where F fusion (x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the fused spectrogram, ω1 represents the first fusion weight coefficient, F1(x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the Mel spectrogram, ω2 represents the second fusion weight coefficient, and F2(x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the Hilbert-Huang spectrogram.

[0098] If the attack mode corresponding to audio copy-move forgery detection is additive noise, set the first fusion weight coefficient to 0.7 and the second fusion weight coefficient to 0.3, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram;

[0099] If the attack mode corresponding to audio copy-move forgery detection is median filtering, set the first fusion weight coefficient to 0.8 and the second fusion weight coefficient to 0.2, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram;

[0100] If the attack mode corresponding to audio copy-move forgery detection is compressed audio attack, set the first fusion weight coefficient to 0.9 and the second fusion weight coefficient to 0.1, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram;

[0101] If the attack mode corresponding to audio copy-move forgery detection is unknown, set the first fusion weight coefficient to 0.8 and the second fusion weight coefficient to 0.2, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram.

[0102] After the weighted fusion of the Mel spectrogram and the Hilbert-Huang spectrogram, the resulting fused spectrogram has the characteristics of highlighting the energy part and remaining unchanged in the non-energy part. The highlighted part in the fused spectrogram is the key point that the audio classification convolutional neural network adopted in this solution needs to extract features and classify correctly.

[0103] S3. Construct an audio classification convolutional neural network and set the network preset parameters;

[0104] As Figure 2 shown, the audio classification convolutional neural network includes an image input layer, a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a fully connected layer, and a softmax layer connected in sequence;

[0105] The first convolution module includes a first convolution layer, a batch normalization layer, a first max pooling layer, and a first activation layer connected in sequence;

[0106] The second convolution module includes a second convolution layer, a second max pooling layer, and a second activation layer; In this embodiment, both the first convolution module and the second convolution module adopt max pooling layers to facilitate better extraction of global features.

[0107] The third convolution module includes a third convolution layer, a third activation layer, and a first average pooling layer connected in sequence;

[0108] The fourth convolution module includes a fourth convolution layer, a fourth activation layer, and a second average pooling layer. In this embodiment, both the third convolution module and the fourth convolution module adopt average pooling layers to facilitate better extraction of local features.

[0109] In this embodiment, the ReLU activation function is used for both the first activation layer and the second activation layer, and the LeakyReLU activation function is used for both the third activation layer and the fourth activation layer. The sizes of the convolution kernels of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer decrease in sequence. As a preferred solution, the size of the convolution kernel of the first convolutional layer is 10×10, the size of the convolution kernel of the second convolutional layer is 7×7, the size of the convolution kernel of the third convolutional layer is 5×5, and the size of the convolution kernel of the fourth convolutional layer is 3×3. The parameters of the hidden layers in the fully connected layer are 32 and 2 in sequence. The audio classification convolutional neural network adopts the Adam optimizer, the learning rate is set to 0.001, and the loss function adopts the binary cross-entropy loss function. In this embodiment, the sizes of the convolution kernels are set to decrease in sequence, so that the receptive fields of the convolutional layers increase from small to large, from global to local. As the sizes of the convolution kernels decrease in sequence, the number of filters also gradually increases to generate more texture-related feature maps. Through the first convolutional layer, a roughly complete texture feature map can be extracted. The second convolutional layer further extracts a global feature map based on the first convolutional module. However, directly classifying based on the features extracted by the second convolutional module will result in a large number of classification errors. Therefore, in this solution, the third convolutional layer and the fourth convolutional layer are used to further extract local feature maps from the extracted global feature map, so that the convolutional neural network can classify more accurately. The detailed features obtained in the local feature map are the prominent parts when the Mel spectrogram and the Hilbert-Huang spectrogram are weighted and fused. This solution is based on the fused spectrogram, highlights the energy parts, and then correctly classifies based on the features extracted from the energy parts, and is applicable to the detection of music.

[0110] S4. Train the audio classification convolutional neural network to obtain a trained audio classification convolutional neural network;

[0111] The S4 includes the following steps:

[0112] S41. Obtain a fused spectrogram training dataset;

[0113] S42. Input each fused spectrogram in the fused spectrogram training dataset into the image input layer;

[0114] S43. Use the first convolutional module to extract features from the fused spectrogram in the image input layer to obtain a first feature;

[0115] The calculation expression of the first feature is as follows:

[0116] F out1 =σ(Maxpooling(BN(C j1 (F fusion ))))

[0117] where F out1represents the first feature, σ(·) represents the Relu activation function, Maxpooling(·) represents the max pooling function, BN(·) represents the batch normalization function, C j1 (·) represents the first convolution function, F fusion represents the fused spectrogram, where the convolution kernel size in the first convolution function is the first scale; in this embodiment, the convolution kernel size of the first scale is 10×10;

[0118] S44. Use the second convolution module to extract features from the texture features to obtain the second feature;

[0119] The calculation expression of the second feature is as follows:

[0120] F out2 = σ(Maxpooling(C j2 (F out1 )))

[0121] where, F out2 represents the second feature, C j2 (·) represents the second convolution function, where the convolution kernel size in the second convolution function is the second scale; in this embodiment, the convolution kernel size of the second scale is 7×7;

[0122] S45. Use the third convolution module to extract features from the second feature to obtain the third feature;

[0123] The calculation expression of the third feature is as follows:

[0124] F out3 = δ(Avgpooling(C j3 (F out2 )))

[0125] where, F out3 represents the third feature, δ(·) represents the LeakyRelu activation function, Avgpooling(·) represents the average pooling function, C j3 (·) represents the third convolution function, where the convolution kernel size in the third convolution function is the third scale; in this embodiment, the convolution kernel size of the third convolution scale is 5×5;

[0126] S46. Use the fourth convolution module to extract features from the third feature to obtain the fourth feature;

[0127] The calculation expression of the fourth feature is as follows:

[0128] F out4 = δ(Avgpooling(C j4 (F out3 )))

[0129] Among them, F out4 represents the fourth feature, and C j4 (·) represents the fourth convolution function. In the fourth convolution function, the size of the convolution kernel is the fourth scale; in this embodiment, the size of the convolution kernel of the fourth scale is 3×3.

[0130] S47. Use the fully connected layer and the softmax layer to perform feature integration and prediction classification on the fourth feature to obtain the classification result corresponding to the fused spectrogram;

[0131] S48. Adjust the preset parameters of the network until the consistency rate between the classification result corresponding to the fused spectrogram and its own category reaches the preset classification threshold, save the preset parameters of the network, complete the training of the audio classification convolutional neural network, and obtain the trained audio classification convolutional neural network. In this embodiment, the preset classification threshold is 98%.

[0132] S5. Input the fused spectrogram into the trained audio classification convolutional neural network for audio copy-move deepfake detection to obtain the audio copy-move deepfake detection result.

[0133] In a practical example of the present invention, the method of the present invention was experimentally verified based on the SHORT-ACMF dataset. The speech duration in the SHORT-ACMF dataset is 3 to 6 seconds, and it contains 4,596 speech files. It is constructed by performing different operations such as noise addition, median filtering, and MP3 compression on the TIMIT dataset. To further test the robustness of the audio classification convolutional neural network, three hybrid post-processing attacks, namely: (MP3(32Kbps)+NOISE(10dB), MP3+MEDIAN, and MP3+MEDIAN+NOISE), were added to the SHORT-ACMF dataset to obtain the HYBRID dataset, which contains 2,298 speech files. To test the performance of the method of the present invention in detecting music ACMF, the MUSIC-ACMF dataset was also created, which includes a total of 350 forged music files.

[0134] ACMF detection is a binary classification problem, including two classification categories: forged and non-forged. The classification results are as follows:

[0135] True positive (TP, True Positive): The true category is forged audio, and the predicted category is forged audio;

[0136] False positive (FP, False Positive): The true category is non-forged audio, and the predicted category is forged audio;

[0137] False Negative (FN): The true class is forged audio, and the predicted class is non-forged audio;

[0138] True Negative (TN): The true class is non-forged audio, and the predicted class is non-forged audio;

[0139] In this practical example, four metrics, namely precision, recall, F1-score, and accuracy, are used to evaluate the performance of the method of the present invention and other ACMF detection methods:

[0140] The calculation expressions for the above-mentioned precision, recall, F1-score, and accuracy are as follows:

[0141]

[0142] Among them, Precision represents precision, Recall represents recall, F1score represents the F1-score, and Accuracy represents accuracy. Among them, the accuracy is the correct rate of the audio copy-move deepfake detection result, the precision is the proportion of the detected forged audio samples that are actually forged audio, the recall is the proportion of the samples correctly identified as forged audio among all the forged audio samples, and the F1-score comprehensively considers the precision and recall.

[0143] When performing weighted fusion of Mel spectrogram and Hilbert-Huang spectrogram, different weights show different performances under different ACMF attacks. In this solution, the trained audio classification convolutional neural network was used to classify the audio data in the SHORT-ACMF dataset; as shown in Figures 3(a), 3(b), 3(c), and 3(d), the larger the first fusion weight coefficient corresponding to the Mel spectrogram, the better the performance of the audio classification convolutional neural network. However, when the first fusion weight coefficient corresponding to the Mel spectrogram is 1 and the second fusion weight coefficient corresponding to the Hilbert-Huang spectrogram is 0, the performance of the audio classification convolutional neural network will decline, which indicates that the spectrogram weighted fusion method provided by the present invention is effective. In Figure 3(a), when the attack method of audio copy-move deepfake is additive noise, when using the weight combination Threshold with the first fusion weight coefficient of 0.7 and the second fusion weight coefficient of 0.3, an F1 score of 94% and an accuracy of 93% can be obtained; in Figure 3(b), when the attack method of audio copy-move deepfake is median filtering, when using the weight combination with the first fusion weight coefficient of 0.8 and the second fusion weight coefficient of 0.2, an F1 score of 90% and an accuracy of 91% can be obtained; in Figure 3(c), when the attack method of audio copy-move deepfake is compressed audio attack, when using the weight combination with the first fusion weight coefficient of 0.9 and the second fusion weight coefficient of 0.1, an F1 score of 90% and an accuracy of 91% can be obtained; in Figure 3(d), when the audio is not attacked after processing, when using the weight combination with the first fusion weight coefficient of 0.8 and the second fusion weight coefficient of 0.2, the recall rate can reach 81% and the other evaluation indicators can all reach more than 90%. When it is expected to detect a large number of attacks in a certain way, a more appropriate weight combination can be selected according to the attack method. If the attack method is uncertain, the weight combination with the first fusion weight coefficient of 0.8 and the second fusion weight coefficient of 0.2 is adopted, which is more comprehensive. Therefore, in this embodiment, it is recommended to use the weight combination with the first fusion weight coefficient of 0.8 and the second fusion weight coefficient of 0.2 to perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram, and use the corresponding fused spectrogram for the detection of audio copy-move deepfake.

[0144] Based on the SHORT-ACMF dataset, the present invention tests the performance of copy-move forgery detection for short speech, and selects representative advanced methods in the existing ACMF detection framework for comparative experiments. In the VAD-based framework, the LBP algorithm and the feature fusion algorithm (PITCH+FORMANT) are selected. In the framework based on key point matching, the SIFT key point matching algorithm is adopted. In the framework based on neural network, the CNN-based method is selected, and comparative experiments are carried out. The single post-processing attack methods in the comparative experiments include: 20dB additive noise attack, 30dB additive noise attack, 32kbps MP3 compression, 64kbps MP3 compression, median filtering, and no attack added; the comparative experiment results of the single post-processing attack are shown in Table 1:

[0145] Table 1

[0146]

[0147] As can be seen from Table 1, in terms of accuracy and F1 measure, the single post-processing attack has little impact on the present scheme. The present scheme has achieved an accuracy and F1 score of more than 89% in all cases except under the 64kbps MP3 compression attack. Even for the 20dB additive noise, the F1 score corresponding to the present scheme can reach 89%. This is because the spectrum fusion and the large convolutional kernel can ignore the changes brought by the attack. The scheme of the present invention will not suffer information loss like the SIFT, SURF, and VAD methods.

[0148] To test the robustness of the present scheme, in this embodiment, comparative experiments under hybrid post-processing attack methods are carried out based on the HYBRID dataset. The hybrid post-processing attack methods in the comparative experiments include: MP3 and Median (first 32kbps MP3 compression, then median filtering), MP3, Median and Noise (first 32kbps MP3 compression, then median filtering, and finally add 20dB additive noise), and MP3 and Noise (first 32kbps MP3 compression, then add 20dB additive noise). The comparative experiment results of the hybrid post-processing attack are shown in Table 2:

[0149] Table 2

[0150]

[0151] As can be seen from Table 2, when there are more post-processings in the hybrid post-processing attack, the present scheme processes the audio spectrogram better. This also reveals the fact that the detection accuracy of non-forged attack audio is lower than that of forged attack audio.

[0152] Such as Figure 4As shown, when the weight combinations of the method of the present invention are [0.7, 0.3], [0.8, 0.2] and [0.9, 0.1] respectively, the effects are compared with those of the LBP algorithm, the SIFT algorithm, the feature fusion algorithm, and the CNN-based method under different post-processing. It can be found that even under the attack mode of hybrid post-processing, this solution can still maintain a stable accuracy rate. This is because after the Mel spectrogram and the Hilbert-Huang spectrogram are fused, the audio information can be effectively retained, making this solution robust to strong post-processing attacks.

[0153] In order to test the detection effect of this solution on forged music, in this embodiment, audio copy-move depth position forgery detection after spectrogram fusion under different weight combinations is carried out based on the MUSIC-ACMF dataset. The results are shown in Table 3:

[0154] Table 3

[0155] Music [0.7,0.3] [0.8,0.2] [0.9,0.1] Accuracy 0.68 0.77 0.80 Recall 0.72 0.88 1 Precision 0.66 0.71 0.71 F1 Score 0.69 0.79 0.83

[0156] As can be seen from Table 3, when the first fusion weight coefficient takes the value of 0.7 and the second fusion weight coefficient takes the value of 0.3, when the first fusion weight coefficient takes the value of 0.8 and the second fusion weight coefficient takes the value of 0.2, and when the first fusion weight coefficient takes the value of 0.9 and the second fusion weight coefficient takes the value of 0.1, this solution can well detect the tampered music.

[0157] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. An audio copy mobile deepfake detection method, characterized in that, It includes the following steps: S1. Obtain the audio to be detected, and generate a Mel spectrogram and a Hilbert-Huang spectrogram based on the audio to be detected; S2. Perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram to obtain a fused spectrogram; The calculation expression of the fused spectrogram in S2 is as follows: Among them, F fusion (x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the fused spectrogram, ω1 represents the first fusion weight coefficient, F1(x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the Mel spectrogram, ω2 represents the second fusion weight coefficient, and F2(x, y) represents the pixel at the position where the abscissa is x and the ordinate is y on the Hilbert-Huang spectrogram; When the attack mode corresponding to audio copy-move forgery detection is additive noise, let the first fusion weight coefficient be 0.7 and the second fusion weight coefficient be 0.3, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram; When the attack mode corresponding to audio copy-move forgery detection is median filtering, let the first fusion weight coefficient be 0.8 and the second fusion weight coefficient be 0.2, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram; When the attack mode corresponding to audio copy-move forgery detection is compressed audio attack, let the first fusion weight coefficient be 0.9 and the second fusion weight coefficient be 0.1, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram; When the attack mode corresponding to audio copy-move forgery detection is unknown, let the first fusion weight coefficient be 0.8 and the second fusion weight coefficient be 0.2, and perform weighted fusion on the Mel spectrogram and the Hilbert-Huang spectrogram; S3. Construct an audio classification convolutional neural network and set the network preset parameters; The audio classification convolutional neural network includes an image input layer, a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a fully connected layer, and a softmax layer connected in sequence; The first convolutional module includes a first convolutional layer, a batch normalization layer, a first max pooling layer, and a first activation layer connected in sequence; The second convolutional module includes a second convolutional layer, a second max pooling layer, and a second activation layer connected in sequence; The third convolutional module includes a third convolutional layer, a third activation layer, and a first average pooling layer connected in sequence; The fourth convolutional module includes a fourth convolutional layer, a fourth activation layer, and a second average pooling layer connected in sequence; S4. Train the audio classification convolutional neural network to obtain a trained audio classification convolutional neural network; S5. Input the fused spectrogram into the trained audio classification convolutional neural network for audio copy-move deep forgery detection to obtain an audio copy-move deep forgery detection result.

2. The audio copy mobile deepfake detection method according to claim 1, wherein, S1 includes the following steps: S11. Extract the frequency and several intrinsic modes of the audio to be detected; The calculation expression of the intrinsic mode of the audio to be detected is as follows: where x(t) represents the signal to be detected, and C i (t) represents the i-th intrinsic mode of the signal to be detected, and R n (t) represents the residual term after the decomposition of the signal to be detected, where i = 1, 2, …, n; S12. Generate a Mel spectrogram according to the frequency of the audio to be detected; The calculation expression of the Mel spectrogram is as follows: where Mel(f) represents the Mel spectrogram, lg(·) represents the logarithmic calculation function with base 10, and f represents the frequency of the audio to be detected; S13. Perform Hilbert transform on the intrinsic mode of the audio to be detected to obtain the Hilbert transform result of the intrinsic mode; S14. Generate a modal transform signal based on the Hilbert transform result of the intrinsic mode; The calculation expression of the modal transform signal is as follows: Among them, Z i (t) represents the modal transform signal, and a i (t) represents the i-th intrinsic mode correlation signal, represents the transformation factor, and d i (t) represents the Hilbert transform result of the i-th intrinsic mode of the signal to be detected; S15. Use the modal transform signal to draw a Hilbert-Huang spectrogram.

3. The audio copy mobile deepfake detection method according to claim 1, characterized in that, The said S4 includes the following steps: S41. Obtain the fused spectrogram training dataset; S42. Input each fused spectrogram in the fused spectrogram training dataset into the image input layer; S43. Use the first convolutional module to extract features from the fused spectrogram in the image input layer to obtain the first feature; The calculation expression of the said first feature is as follows: F out1 = σ(Maxpooling(BN(C j1 (F fusion )))) Among them, F out1 represents the first feature, σ(·) represents the Relu activation function, Maxpooling(·) represents the max pooling function, BN(·) represents the batch normalization function, and C j1 (·) represents the first convolution function, and F fusion represents the fused spectrogram, where the convolution kernel size in the first convolution function is the first scale; S44. Use the second convolutional module to extract features from the texture feature to obtain the second feature; The calculation expression of the said second feature is as follows: F out2 = σ(Maxpooling(C j2 (F out1 ))) Among them, F out2 represents the second feature, and C j2 (·) represents the second convolution function, where the convolution kernel size in the second convolution function is the second scale; S45. Use the third convolutional module to extract features from the second feature to obtain the third feature; The calculation expression of the said third feature is as follows: F out3 = δ(Avgpooling(C j3 (F out2 ))) Among them, F out3 represents the third feature, δ(·) represents the LeakyRelu activation function, Avgpooling(·) represents the average pooling function, C j3 (·) represents the third convolution function, where the size of the convolution kernel in the third convolution function is the third scale; S46. Use the fourth convolutional module to extract features from the third feature to obtain the fourth feature; The calculation expression of the said fourth feature is as follows: F out4 = δ(Avgpooling(C j4 (F out3 ))) Among them, F out4 represents the fourth feature, and C j4 (·) represents the fourth convolution function, where the size of the convolution kernel in the fourth convolution function is the fourth scale; S47. Use the fully connected layer and the softmax layer to perform feature integration and prediction classification on the fourth feature to obtain the classification result corresponding to the fused spectrogram; S48. Adjust the preset parameters of the network until the consistency rate between the classification result corresponding to the fused spectrogram and its own category reaches the preset classification threshold, save the preset parameters of the network, complete the training of the audio classification convolutional neural network, and obtain the trained audio classification convolutional neural network.

Citation Information

Patent Citations

  • Text classification method and device based on improved textCNN model and storage medium

    CN109918497A

  • Multi-feature fusion braking noise classification and identification method

    CN115081473A