An audio deepfake detection method fusing multi-source features and cross-scale modeling
By integrating multi-source features and cross-scale modeling, an audio deepfake detection method is developed, which addresses the shortcomings of traditional detection methods in terms of detection accuracy and robustness, and achieves efficient identification and accurate detection of deepfake audio.
Patent Information
- Application Number
- CN202511294644.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-11
AI Technical Summary
In existing technologies, traditional audio deepfake detection methods lack explicit attention to underlying acoustic and physical features, making it difficult to capture deviations in the physical laws of speech generation, resulting in insufficient detection accuracy or poor robustness, and failing to fully cover all forgery patterns.
We employ a method that integrates multi-source features and cross-scale modeling. We acquire multi-level deep audio features and physical acoustic features through a dual-branch data augmentation strategy. We then combine channel attention and gating strategies to perform feature fusion, constructing a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities. We integrate local and global perception information using grouped convolution and multi-branch attention paths.
It significantly improves the ability to identify deepfake audio, enhances generalization and robustness, improves detection accuracy and the ability to capture subtle changes in features across time and frequency scales in fake audio, while maintaining the naturalness and auditory acceptability of the speech.
Smart Images

Figure CN120783799B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio detection, and in particular to an audio deepfake detection method that integrates multi-source features and cross-scale modeling. Background Technology
[0002] With the development of generative artificial intelligence, audio language models that combine speech processing and natural language processing have made significant progress, making speech cloning, speech conversion and text-to-speech technologies more mature. The threshold for creating fake audio has been greatly reduced. Deep fake audio can be highly realistic in semantics and intonation, which has led to an increasing risk of false dissemination and identity fraud. Therefore, deep fake audio detection has become an essential part.
[0003] In existing technologies, traditional detection methods generally rely on high-level speech representations extracted by pre-trained models, lacking explicit attention to underlying acoustic and physical features. This makes it difficult to capture deviations in the physical laws of speech generation. Meanwhile, forgery features often exist as weak artifacts at multiple spatiotemporal scales and abstraction levels, resulting in existing traditional detection methods being unable to fully cover all forgery patterns and easily leading to insufficient detection accuracy or poor robustness.
[0004] Therefore, designing a deepfake audio detection method to improve the accurate identification of various deepfake audio types has become an urgent problem to be solved. Summary of the Invention
[0005] Based on this, this invention proposes an audio deepfake detection method that integrates multi-source features and cross-scale modeling. Through a dual-branch data augmentation strategy, it effectively improves the ability to identify hidden forgery patterns in synthetic samples. While maintaining the naturalness and auditory acceptability of the speech, it introduces a representative forgery perturbation space, thereby significantly improving generalization and robustness under various types of deepfake attacks. Furthermore, it acquires multi-level deep audio features and physical acoustic features separately and performs feature fusion, combining channel attention and gating strategies to enhance the expressive power of forgery-related information and improve sensitivity to physical layer artifacts, further improving detection accuracy. It also constructs a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities through multi-scale attention enhancement, which can efficiently capture subtle changes across time and frequency scales in forged audio. Through the structural design of grouped convolution and multi-branch attention paths, it effectively integrates local and global perceptual information, providing a highly expressive intermediate representation for the forgery detection task. This invention improves the accuracy, robustness, and generalization ability of audio deepfake detection.
[0006] This invention proposes an audio deepfake detection method that integrates multi-source features and cross-scale modeling, comprising:
[0007] The raw audio data is acquired and data augmentation processing is performed. The data augmentation processing is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch.
[0008] Feature extraction is performed on the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features, respectively. The multi-level deep audio features are based on a cross-language speech representation learning model.
[0009] Feature fusion processing is performed based on the multi-level deep audio features and physical acoustic features to obtain fused features. The feature fusion processing is based on channel attention mechanism and gating mechanism.
[0010] The fused features are subjected to multi-scale attention enhancement processing to obtain multi-scale attention enhancement features, wherein the multi-scale attention enhancement processing is based on two parallel convolutional branches;
[0011] Classification is performed based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection results.
[0012] In summary, this audio deepfake detection method, which integrates multi-source features and cross-scale modeling, effectively improves the ability to identify hidden forgery patterns in synthetic samples through a dual-branch data augmentation strategy. While maintaining the naturalness and auditory acceptability of the speech, it introduces a representative forgery perturbation space, significantly improving generalization and robustness under various types of deepfake attacks. Furthermore, it acquires multi-level deep audio features and physical acoustic features separately, and performs feature fusion, combining channel attention and gating strategies to enhance the expressive power of forgery-related information and improve sensitivity to physical layer artifacts, further enhancing detection accuracy. It also constructs a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities through multi-scale attention enhancement, efficiently capturing subtle changes across time and frequency scales in forged audio. Through the structural design of grouped convolution and multi-branch attention paths, it effectively integrates local and global perceptual information, providing a highly expressive intermediate representation for forgery detection tasks. This invention improves the accuracy, robustness, and generalization ability of audio deepfake detection. Specifically, the process involves acquiring raw audio data and performing data augmentation. This data augmentation is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptually invariant perturbation branch. This effectively improves the ability to identify hidden forgery patterns in synthetic samples. While maintaining the naturalness and auditory acceptability of the speech, it introduces a representative forgery perturbation space, thereby significantly improving generalization and robustness against various types of deepfake attacks. Feature extraction is then performed on the augmented raw audio data to obtain multi-level deep audio features and physical acoustic features. These multi-level deep audio features are based on a cross-lingual speech representation learning model. Feature fusion processing is then performed based on these multi-level deep audio features and physical acoustic features to obtain fused features. This feature fusion processing is based on a channel attention machine. By implementing control and gating mechanisms, the expressive power of forgery-related information is enhanced, the sensitivity to physical layer artifacts is improved, and the detection accuracy is further improved. Multi-scale attention enhancement processing is applied to the fused features to obtain multi-scale attention enhancement features. This multi-scale attention enhancement processing is based on two parallel convolutional branches, constructing a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities. It can efficiently capture subtle changes across time and frequency scales in forged audio. Through the structural design of grouped convolution and multi-branch attention paths, local and global perceptual information is effectively integrated, providing a highly expressive intermediate representation for the forgery detection task. Classification is performed based on the multi-scale attention enhancement features to obtain the final audio deep forgery detection result. This invention improves the accuracy, robustness, and generalization ability of audio deep forgery detection.
[0013] Furthermore, the step of acquiring the raw audio data and performing data enhancement processing specifically includes:
[0014] Obtain the raw audio data and input it into the semantic artifact injection branch and the perceptual invariant perturbation branch respectively;
[0015] The semantic artifact injection branch constructs class-forged enhanced samples based on the original audio data, and the perceptual invariant perturbation branch performs hidden generation artifact capture enhancement based on the original audio data. The hidden generation artifact capture enhancement is based on a perturbation generation strategy constrained by perceptual distance.
[0016] The forgery-like enhancement sample construction performs local band masking, band-limiting inversion, and temporal nonlinear stretching on the Mel spectrum to simulate nonlinear distortion of deep synthesized audio. Then, based on the speech style transfer strategy, pseudo-speech segments are synthesized according to the pre-trained TTS model. The pseudo-speech segments and the original audio data are the same text. The pseudo-speech segments and the original audio data are fused with low weights to obtain synthesized timbre artifacts. Then, based on speech rate modification processing, local silence segment addition, and background audio superposition, temporal mismatch artifact simulation is performed to obtain forgery-like enhancement samples.
[0017] The discrimination boundary is optimized based on the aforementioned forged enhanced samples and hidden artifact capture enhanced samples, and the discrimination boundary optimization is based on a random sampling mechanism.
[0018] Furthermore, the step of enhancing the hidden artifact capture based on the original audio data by the perceptually invariant perturbation branch specifically includes:
[0019] Based on the perturbation generation strategy constrained by perceptual distance, perturbation segments are generated for each audio segment in the original audio data. The perturbation generation strategy constrained by perceptual distance is as follows:
[0020] ,
[0021] ,
[0022] in, Indicates a perturbation segment. This represents the original audio segment in the original audio data. This represents the disturbance coefficient. This represents the intermediate representation of the deep feature extraction model. Indicates the perceived speech quality index. This represents the human ear perception threshold control term. This indicates the upper limit of the disturbance amplitude.
[0023] Furthermore, the step of extracting features from the augmented original audio data to obtain multi-level deep audio features and physical acoustic features specifically includes:
[0024] Multi-level deep audio feature extraction is performed on the original audio data based on a cross-language speech representation learning model. The specific algorithm for multi-level deep audio feature extraction is as follows:
[0025] ,
[0026] in, Represents multi-layered deep audio features. Indicates feature connection, This represents the output vector of each transformer layer. Indicates the number of layers;
[0027] Then, physical acoustic features are obtained based on the original audio data. The physical acoustic features include time domain features, frequency domain features, sub-band energy features, and frequency trajectory features. The time domain features include short-time energy and zero-crossing rate. The frequency domain features include spectral centroid and spectral roll-off index. The sub-band energy features are based on the division of frequency bands by the Mel filter bank and the extraction of the energy of each sub-band. The frequency trajectory features are based on the instantaneous dominant frequency and the first-order difference standard deviation of the instantaneous dominant frequency.
[0028] Furthermore, the step of performing feature fusion processing based on the multi-level deep audio features and physical acoustic features to obtain fused features specifically includes:
[0029] Time-dimensional alignment of multi-layered deep audio features with physical acoustic features;
[0030] An adaptive fusion algorithm is used to fuse multi-level deep audio features with physical acoustic features to obtain fused features. The adaptive fusion algorithm includes a channel attention mechanism and a gating mechanism. Specifically, the adaptive fusion algorithm includes:
[0031] ,
[0032] in, Indicates fusion features, Represents the spatial gating tensor. , These represent the channel attention vectors for multi-layered deep audio features and physical acoustic features, respectively. Represents multi-layered deep audio features. Representing physical acoustic features, the channel attention vector is generated based on global pooling and a shared multilayer perceptron.
[0033] Furthermore, the step of performing multi-scale attention enhancement processing on the fused features to obtain multi-scale attention-enhanced features specifically includes:
[0034] The fused features are subjected to multi-scale attention enhancement processing, which is based on two parallel convolutional branches, the two parallel convolutional branches including Convolutional branches and Convolutional branches;
[0035] The fused features are divided into channels to obtain multiple feature sub-tensors, and the feature sub-tensors are input respectively. Convolutional branches and Convolutional branches;
[0036] The The convolutional branch includes a one-dimensional global average pooling along the horizontal dimension and a one-dimensional global average pooling along the vertical dimension to obtain horizontal and vertical positional information, respectively. The specific algorithm for the convolution branch is as follows:
[0037] ,
[0038] ,
[0039] in, Indicates horizontal position information. Indicates vertical dimension location information. , These represent the time dimension and the frequency dimension, respectively. , Indexes representing the frequency and time dimensions. Represents the fused feature vector;
[0040] The The convolutional branch is based on a cross-spatial information aggregation mechanism, the Convolution branches include Convolution branch paths and Convolutional branch path, the The convolutional branch path encodes global spatial information based on two-dimensional global average pooling. The specific algorithm for the two-dimensional global average pooling is as follows:
[0041] ,
[0042] in, This represents the output of a two-dimensional global average pooling method.
[0043] The Convolutional branches and the The output of the convolutional branch undergoes sub-tensor merging processing. The specific algorithm for this sub-tensor merging processing is as follows:
[0044] ,
[0045] in, This indicates multi-scale attention enhancement features. This indicates tensor merging and reorganization. This indicates multi-scale attention enhancement processing. Represents the characteristic tensor. Indicates the number of channel divisions.
[0046] Furthermore, the step of classifying based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection result specifically includes:
[0047] Multi-scale attention-enhanced features are added together in dimension 1 to reshape the dimension;
[0048] Then, feature extraction is performed on the multi-scale attention-enhanced features after reshaping the dimensions using two-dimensional max pooling to obtain the main features;
[0049] The final audio depth forgery detection result is obtained by classifying based on the main features. The specific algorithm for classification is as follows:
[0050] ,
[0051] in, This indicates the final audio deepfake detection result. , This represents the weights of the fully connected layers in the first and third layers. , This indicates the bias of the fully connected layers 1 and 3. Indicates the main features, This indicates flattening. This represents two-dimensional max pooling.
[0052] This invention proposes an audio deepfake detection system that integrates multi-source features and cross-scale modeling, comprising:
[0053] The data augmentation module is used to acquire raw audio data and perform data augmentation processing. The data augmentation processing is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch.
[0054] The feature extraction module is used to extract features from the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features, respectively. The multi-level deep audio features are based on a cross-language speech representation learning model.
[0055] The feature fusion module is used to perform feature fusion processing based on the multi-level deep audio features and physical acoustic features to obtain fused features. The feature fusion processing is based on channel attention mechanism and gating mechanism.
[0056] A multi-scale attention enhancement module is used to perform multi-scale attention enhancement processing on the fused features to obtain multi-scale attention enhancement features. The multi-scale attention enhancement processing is based on two parallel convolutional branches.
[0057] A classification module is used to classify based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection result.
[0058] The present invention also provides a storage medium storing one or more programs that, when executed by a processor, implement the audio depth forgery detection method described above, which integrates multi-source features and cross-scale modeling.
[0059] The present invention also provides a computer device, the computer device including a memory and a processor, wherein:
[0060] The memory is used to store computer programs;
[0061] When the processor executes the computer program stored in the memory, it implements the audio depth forgery detection method described above, which integrates multi-source features and cross-scale modeling. Attached Figure Description
[0062] Figure 1 This is a flowchart of the audio deep forgery detection method that integrates multi-source features and cross-scale modeling proposed in the first embodiment of the present invention.
[0063] Figure 2 This is a schematic diagram of the structure of the audio deep forgery detection system that integrates multi-source features and cross-scale modeling proposed in the second embodiment of the present invention.
[0064] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0065] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0066] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0068] Please see Figure 1 The diagram shows a flowchart of the audio deepfake detection method integrating multi-source features and cross-scale modeling proposed in the first embodiment of the present invention. This audio deepfake detection method integrating multi-source features and cross-scale modeling includes steps S01 to S05, wherein:
[0069] Step S01: Acquire raw audio data and perform data enhancement processing;
[0070] It should be noted that in this embodiment, the data augmentation process is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch. The original audio data is acquired and input into the semantic artifact injection branch and the perceptual invariant perturbation branch respectively.
[0071] The semantic artifact injection branch constructs class-forged enhanced samples based on the original audio data, and the perceptual invariant perturbation branch performs hidden generation artifact capture enhancement based on the original audio data. The hidden generation artifact capture enhancement is based on a perturbation generation strategy constrained by perceptual distance.
[0072] The forgery-like enhancement sample construction performs local band masking, band-limiting inversion, and temporal nonlinear stretching on the Mel spectrum to simulate nonlinear distortion of deep synthesized audio. Then, based on the speech style transfer strategy, pseudo-speech segments are synthesized according to the pre-trained TTS model. The pseudo-speech segments and the original audio data are the same text. The pseudo-speech segments and the original audio data are fused with low weights to obtain synthesized timbre artifacts. Then, based on speech rate modification processing, local silence segment addition, and background audio superposition, temporal mismatch artifact simulation is performed to obtain forgery-like enhancement samples.
[0073] The discrimination boundary is optimized based on the aforementioned forged enhanced samples and hidden artifact capture enhanced samples, and the discrimination boundary optimization is based on a random sampling mechanism.
[0074] Based on the perturbation generation strategy constrained by perceptual distance, perturbation segments are generated for each audio segment in the original audio data. The perturbation generation strategy constrained by perceptual distance is as follows:
[0075] ,
[0076] ,
[0077] in, Indicates a perturbation segment. This represents the original audio segment in the original audio data. This represents the disturbance coefficient. This represents the intermediate representation of the deep feature extraction model. Indicates the perceived speech quality index. This represents the human ear perception threshold control term. This indicates the upper limit of the disturbance amplitude.
[0078] Step S02: Perform feature extraction on the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features respectively;
[0079] It should be noted that in this embodiment, the multi-level deep audio features are based on a cross-lingual speech representation learning model. The large-scale cross-lingual speech representation learning model in this embodiment originates from wav2vec 2.0 and has been trained on 128 languages. The original audio signal is first processed by a feature encoder comprising multiple convolutional neural networks (CNNs). This feature encoder extracts a 1024-dimensional vector representation every 20ms through a receptive field of 25ms to obtain a latent representation. These encoder features are then fed into a 24-layer transformer network to derive a contextual representation. The output vectors from all 24 transformer layers are concatenated into a composite representation. In this embodiment, the outputs of all 24 transformer hidden layers are retained, and the hidden layer features of all layers are concatenated. Multi-level deep audio features are extracted from the original audio data according to the cross-lingual speech representation learning model. The specific algorithm for multi-level deep audio feature extraction is as follows:
[0080] ,
[0081] in, Represents multi-layered deep audio features. Indicates feature connection, This represents the output vector of each transformer layer. Indicates the number of layers;
[0082] Then, physical acoustic features are obtained based on the original audio data. The physical acoustic features include time domain features, frequency domain features, sub-band energy features, and frequency trajectory features. The time domain features include short-time energy and zero-crossing rate. The frequency domain features include spectral centroid and spectral roll-off index. The sub-band energy features are based on the division of frequency bands by the Mel filter bank and the extraction of the energy of each sub-band. The frequency trajectory features are based on the instantaneous dominant frequency and the first-order difference standard deviation of the instantaneous dominant frequency.
[0083] Step S03: Perform feature fusion processing based on multi-level deep audio features and physical acoustic features to obtain fused features;
[0084] It should be noted that in this embodiment, the feature fusion processing is based on channel attention mechanism and gating mechanism to align multi-level deep audio features with physical acoustic features in the time dimension.
[0085] An adaptive fusion algorithm is used to fuse multi-level deep audio features with physical acoustic features to obtain fused features. The adaptive fusion algorithm includes a channel attention mechanism and a gating mechanism. Specifically, the adaptive fusion algorithm includes:
[0086] ,
[0087] in, Indicates fusion features, Represents the spatial gating tensor. , These represent the channel attention vectors for multi-layered deep audio features and physical acoustic features, respectively. Represents multi-layered deep audio features. Representing physical acoustic features, the channel attention vector is generated based on global pooling and a shared multilayer perceptron.
[0088] Step S04: Perform multi-scale attention enhancement processing on the fused features to obtain multi-scale attention-enhanced features;
[0089] It should be noted that in this embodiment, the multi-scale attention enhancement processing is based on two parallel convolutional branches to perform multi-scale attention enhancement processing on the fused features. The multi-scale attention enhancement processing is based on two parallel convolutional branches, and these two parallel convolutional branches include... Convolutional branches and Convolutional branches;
[0090] The fused features are divided into channels to obtain multiple feature sub-tensors, and the feature sub-tensors are input respectively. Convolutional branches and Convolutional branches;
[0091] The The convolutional branch includes a one-dimensional global average pooling along the horizontal dimension and a one-dimensional global average pooling along the vertical dimension to obtain horizontal and vertical positional information, respectively. The specific algorithm for the convolution branch is as follows:
[0092] ,
[0093] ,
[0094] in, Indicates horizontal position information. Indicates vertical dimension location information. , These represent the time dimension and the frequency dimension, respectively. , Indexes representing the frequency and time dimensions. Represents the fused feature vector;
[0095] The The convolutional branch is based on a cross-spatial information aggregation mechanism, the Convolution branches include Convolution branch paths and Convolutional branch path, the The convolutional branch path encodes global spatial information based on two-dimensional global average pooling. The specific algorithm for the two-dimensional global average pooling is as follows:
[0096] ,
[0097] in, This represents the output of a two-dimensional global average pooling method.
[0098] The Convolutional branches and the The output of the convolutional branch undergoes sub-tensor merging processing. The specific algorithm for this sub-tensor merging processing is as follows:
[0099] ,
[0100] in, This indicates multi-scale attention enhancement features. This indicates tensor merging and reorganization. This indicates multi-scale attention enhancement processing. Represents the characteristic tensor. Indicates the number of channel divisions.
[0101] Step S05: Classify based on multi-scale attention enhancement features to obtain the final audio depth forgery detection results;
[0102] It should be noted that in this embodiment, the multi-scale attention enhancement features are added together in dimension 1 to reshape the dimension;
[0103] Then, feature extraction is performed on the multi-scale attention-enhanced features after reshaping the dimensions using two-dimensional max pooling to obtain the main features;
[0104] The final audio depth forgery detection result is obtained by classifying based on the main features. The specific algorithm for classification is as follows:
[0105] ,
[0106] in, This indicates the final audio deepfake detection result. , This represents the weights of the fully connected layers in the first and third layers. , This indicates the bias of the fully connected layers 1 and 3. Indicates the main features, This indicates flattening. This represents two-dimensional max pooling.
[0107] In summary, this audio deepfake detection method, which integrates multi-source features and cross-scale modeling, effectively improves the ability to identify hidden forgery patterns in synthetic samples through a dual-branch data augmentation strategy. While maintaining the naturalness and auditory acceptability of the speech, it introduces a representative forgery perturbation space, significantly improving generalization and robustness under various types of deepfake attacks. Furthermore, it acquires multi-level deep audio features and physical acoustic features separately, and performs feature fusion, combining channel attention and gating strategies to enhance the expressive power of forgery-related information and improve sensitivity to physical layer artifacts, further enhancing detection accuracy. It also constructs a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities through multi-scale attention enhancement, efficiently capturing subtle changes across time and frequency scales in forged audio. Through the structural design of grouped convolution and multi-branch attention paths, it effectively integrates local and global perceptual information, providing a highly expressive intermediate representation for forgery detection tasks. This invention improves the accuracy, robustness, and generalization ability of audio deepfake detection. Specifically, the process involves acquiring raw audio data and performing data augmentation. This data augmentation is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptually invariant perturbation branch. This effectively improves the ability to identify hidden forgery patterns in synthetic samples. While maintaining the naturalness and auditory acceptability of the speech, it introduces a representative forgery perturbation space, thereby significantly improving generalization and robustness against various types of deepfake attacks. Feature extraction is then performed on the augmented raw audio data to obtain multi-level deep audio features and physical acoustic features. These multi-level deep audio features are based on a cross-lingual speech representation learning model. Feature fusion processing is then performed based on these multi-level deep audio features and physical acoustic features to obtain fused features. This feature fusion processing is based on a channel attention machine. By implementing control and gating mechanisms, the expressive power of forgery-related information is enhanced, the sensitivity to physical layer artifacts is improved, and the detection accuracy is further improved. Multi-scale attention enhancement processing is applied to the fused features to obtain multi-scale attention enhancement features. This multi-scale attention enhancement processing is based on two parallel convolutional branches, constructing a forgery pattern recognition framework with multi-scale and spatial context modeling capabilities. It can efficiently capture subtle changes across time and frequency scales in forged audio. Through the structural design of grouped convolution and multi-branch attention paths, local and global perceptual information is effectively integrated, providing a highly expressive intermediate representation for the forgery detection task. Classification is performed based on the multi-scale attention enhancement features to obtain the final audio deep forgery detection result. This invention improves the accuracy, robustness, and generalization ability of audio deep forgery detection.
[0108] Please see Figure 2The figure shows a schematic diagram of the audio deepfake detection system that integrates multi-source features and cross-scale modeling proposed in the second embodiment of the present invention. The system includes:
[0109] Data augmentation module 10 is used to acquire raw audio data and perform data augmentation processing. The data augmentation processing is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch.
[0110] The feature extraction module 20 is used to extract features from the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features, respectively. The multi-level deep audio features are based on a cross-language speech representation learning model.
[0111] Feature fusion module 30 is used to perform feature fusion processing based on the multi-level deep audio features and physical acoustic features to obtain fused features. The feature fusion processing is based on channel attention mechanism and gating mechanism.
[0112] The multi-scale attention enhancement module 40 is used to perform multi-scale attention enhancement processing on the fused features to obtain multi-scale attention enhancement features. The multi-scale attention enhancement processing is based on two parallel convolutional branches.
[0113] The classification module 50 is used to classify based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection result.
[0114] The present invention also proposes a computer storage medium storing one or more programs that, when executed by a processor, implement the above-described audio depth forgery detection method that integrates multi-source features and cross-scale modeling.
[0115] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to realize the above-mentioned audio depth forgery detection method that integrates multi-source features and cross-scale modeling.
[0116] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0117] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0118] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0120] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for detecting deep audio forgery that integrates multi-source features and cross-scale modeling, characterized in that, include: The raw audio data is acquired and data augmentation processing is performed. The data augmentation processing is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch. The steps of acquiring raw audio data and performing data enhancement processing specifically include: Obtain the raw audio data and input it into the semantic artifact injection branch and the perceptual invariant perturbation branch respectively; The semantic artifact injection branch constructs class-forged enhanced samples based on the original audio data, and the perceptual invariant perturbation branch performs hidden generation artifact capture enhancement based on the original audio data. The hidden generation artifact capture enhancement is based on a perturbation generation strategy constrained by perceptual distance. The forgery-like enhancement sample construction performs local band masking, band-limiting inversion, and temporal nonlinear stretching on the Mel spectrum to simulate nonlinear distortion of deep synthesized audio. Then, based on the speech style transfer strategy, pseudo-speech segments are synthesized according to the pre-trained TTS model. The pseudo-speech segments and the original audio data are the same text. The pseudo-speech segments and the original audio data are fused with low weights to obtain synthesized timbre artifacts. Then, based on speech rate modification processing, local silence segment addition, and background audio superposition, temporal mismatch artifact simulation is performed to obtain forgery-like enhancement samples. The discrimination boundary is optimized based on the aforementioned forgery enhancement samples and hidden artifact capture enhancement samples, and the discrimination boundary optimization is based on a random sampling mechanism. The step of enhancing artifact capture by the perceptual invariant perturbation branch based on the original audio data specifically includes: Based on the perturbation generation strategy constrained by perceptual distance, perturbation segments are generated for each audio segment in the original audio data. The perturbation generation strategy constrained by perceptual distance is as follows: , , in, Indicates a perturbation segment. This represents the original audio segment in the original audio data. This represents the disturbance coefficient. This represents the intermediate representation of the deep feature extraction model. Indicates the perceived speech quality index. This represents the human ear perception threshold control term. Indicates the upper limit of the disturbance amplitude; Feature extraction is performed on the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features, respectively. The multi-level deep audio features are based on a cross-language speech representation learning model. Feature fusion processing is performed based on the multi-level deep audio features and physical acoustic features to obtain fused features. The feature fusion processing is based on channel attention mechanism and gating mechanism. The fused features are subjected to multi-scale attention enhancement processing to obtain multi-scale attention enhancement features, wherein the multi-scale attention enhancement processing is based on two parallel convolutional branches; Classification is performed based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection results.
2. The audio deepfake detection method integrating multi-source features and cross-scale modeling according to claim 1, characterized in that, The step of extracting features from the augmented original audio data to obtain multi-level deep audio features and physical acoustic features specifically includes: Multi-level deep audio feature extraction is performed on the original audio data based on a cross-language speech representation learning model. The specific algorithm for multi-level deep audio feature extraction is as follows: , in, Represents multi-layered deep audio features. Indicates feature connection, This represents the output vector of each transformer layer. Indicates the number of floors; Then, physical acoustic features are obtained based on the original audio data. The physical acoustic features include time-domain features, frequency-domain features, sub-band energy features, and frequency trajectory features. The time-domain features include short-time energy and zero-crossing rate. The frequency-domain features include spectral centroid and spectral roll-off index. The sub-band energy features are based on the Mel filter bank to divide the frequency bands and extract the energy of each sub-band. The frequency trajectory features are based on the instantaneous dominant frequency and the first-order difference standard deviation of the instantaneous dominant frequency.
3. The audio depth forgery detection method fusing multi-source features and cross-scale modeling according to claim 1, characterized in that, The step of performing feature fusion processing based on the multi-level deep audio features and physical acoustic features to obtain fused features specifically includes: Time-dimensional alignment of multi-layered deep audio features with physical acoustic features; An adaptive fusion algorithm is used to fuse multi-level deep audio features with physical acoustic features to obtain fused features. The adaptive fusion algorithm includes a channel attention mechanism and a gating mechanism. Specifically, the adaptive fusion algorithm includes: , in, Indicates fusion characteristics, Represents the spatial gating tensor. , These represent the channel attention vectors for multi-layered deep audio features and physical acoustic features, respectively. Represents multi-layered deep audio features. Representing physical acoustic features, the channel attention vector is generated based on global pooling and a shared multilayer perceptron.
4. The audio deepfake detection method integrating multi-source features and cross-scale modeling according to claim 1, characterized in that, The step of performing multi-scale attention enhancement processing on the fused features to obtain multi-scale attention-enhanced features specifically includes: The fused features are subjected to multi-scale attention enhancement processing, which is based on two parallel convolutional branches, the two parallel convolutional branches including Convolutional branches and Convolutional branches; The fused features are divided into channels to obtain multiple feature sub-tensors, and the feature sub-tensors are input respectively. Convolutional branches and Convolutional branches; The The convolutional branch includes a one-dimensional global average pooling along the horizontal dimension and a one-dimensional global average pooling along the vertical dimension to obtain horizontal and vertical positional information, respectively. The specific algorithm for the convolution branch is as follows: , , in, Indicates horizontal position information. Indicates vertical dimension location information. , These represent the time dimension and the frequency dimension, respectively. , Indexes representing the frequency and time dimensions. Represents the fused feature vector; The The convolutional branch is based on a cross-spatial information aggregation mechanism, the Convolution branches include Convolution branch paths and Convolutional branch path, the The convolutional branch path encodes global spatial information based on two-dimensional global average pooling. The specific algorithm for the two-dimensional global average pooling is as follows: , in, This represents the output of a two-dimensional global average pooling method. The Convolutional branches and the The output of the convolutional branch undergoes sub-tensor merging processing, and the specific algorithm for this sub-tensor merging processing is as follows: , in, This indicates multi-scale attention enhancement features. This indicates tensor merging and reorganization. This indicates multi-scale attention enhancement processing. Represents the characteristic tensor. Indicates the number of channel divisions.
5. The audio deepfake detection method integrating multi-source features and cross-scale modeling according to claim 1, characterized in that, The step of classifying based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection result specifically includes: Multi-scale attention-enhanced features are added together along the layer dimension to reshape the dimension; Then, feature extraction is performed on the multi-scale attention-enhanced features after reshaping the dimensions using two-dimensional max pooling to obtain the main features; The final audio depth forgery detection result is obtained by classifying based on the main features. The specific algorithm for classification is as follows: , in, This indicates the final audio deepfake detection result. , This represents the weights of the fully connected layers in the first and third layers. , This indicates the bias of the fully connected layers 1 and 3. Indicates the main features, This indicates flattening. This represents two-dimensional max pooling.
6. An audio deepfake detection system that integrates multi-source features and cross-scale modeling, characterized in that, include: The data augmentation module is used to acquire raw audio data and perform data augmentation processing. The data augmentation processing is based on a dual-branch data augmentation strategy, which includes a semantic artifact injection branch and a perceptual invariant perturbation branch. The steps of acquiring raw audio data and performing data enhancement processing specifically include: Obtain the raw audio data and input it into the semantic artifact injection branch and the perceptual invariant perturbation branch respectively; The semantic artifact injection branch constructs class-forged enhanced samples based on the original audio data, and the perceptual invariant perturbation branch performs hidden generation artifact capture enhancement based on the original audio data. The hidden generation artifact capture enhancement is based on a perturbation generation strategy constrained by perceptual distance. The forgery-like enhancement sample construction performs local band masking, band-limiting inversion, and temporal nonlinear stretching on the Mel spectrum to simulate nonlinear distortion of deep synthesized audio. Then, based on the speech style transfer strategy, pseudo-speech segments are synthesized according to the pre-trained TTS model. The pseudo-speech segments and the original audio data are the same text. The pseudo-speech segments and the original audio data are fused with low weights to obtain synthesized timbre artifacts. Then, based on speech rate modification processing, local silence segment addition, and background audio superposition, temporal mismatch artifact simulation is performed to obtain forgery-like enhancement samples. The discrimination boundary is optimized based on the aforementioned forgery enhancement samples and hidden artifact capture enhancement samples, and the discrimination boundary optimization is based on a random sampling mechanism. The step of enhancing artifact capture by the perceptual invariant perturbation branch based on the original audio data specifically includes: Based on the perturbation generation strategy constrained by perceptual distance, perturbation segments are generated for each audio segment in the original audio data. The perturbation generation strategy constrained by perceptual distance is as follows: , , in, Indicates a perturbation segment. This represents the original audio segment in the original audio data. This represents the disturbance coefficient. This represents the intermediate representation of the deep feature extraction model. Indicates the perceived speech quality index. This represents the human ear perception threshold control term. Indicates the upper limit of the disturbance amplitude; The feature extraction module is used to extract features from the data-enhanced original audio data to obtain multi-level deep audio features and physical acoustic features, respectively. The multi-level deep audio features are based on a cross-language speech representation learning model. The feature fusion module is used to perform feature fusion processing based on the multi-level deep audio features and physical acoustic features to obtain fused features. The feature fusion processing is based on channel attention mechanism and gating mechanism. A multi-scale attention enhancement module is used to perform multi-scale attention enhancement processing on the fused features to obtain multi-scale attention enhancement features. The multi-scale attention enhancement processing is based on two parallel convolutional branches. A classification module is used to classify based on the multi-scale attention enhancement features to obtain the final audio depth forgery detection result.
7. A storage medium, characterized in that, The storage medium stores one or more programs that, when executed by a processor, implement the audio depth forgery detection method as described in any one of claims 1-5, which integrates multi-source features and cross-scale modeling.
8. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the audio depth forgery detection method according to any one of claims 1-5, which integrates multi-source features and cross-scale modeling.
Citation Information
Patent Citations
Multi-scale feature fusion depth forgery detection method based on reconstruction learning
CN119068318A
Face depth forgery detection method based on enhanced double-branch fusion model
CN119723638A