Audio authenticity identification method and device based on dual-path Transform
By introducing a dual-path Transformer structure and attention mechanism in audio signal processing, the problem of difficult to decouple semantic features and acoustic features in the prior art is solved, and more effective audio signal feature extraction and identification are achieved.
Patent Information
- Application Number
- CN202510222544.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively capture long-term dependencies and complex time-frequency modes in feature extraction of audio signals, and it is difficult to effectively decouple semantic features and acoustic features.
The audio authenticity and false identification method based on dual-path Transformer is adopted, and the frequency and time domain information is processed separately through the dual-path Transformer structure, and combined with the attention mechanism, the efficient decoupling of semantic information and acoustic information is achieved.
It better captures the long-term dependency relationship and time-frequency information in the audio signal, improves the feature extraction and utilization ability of the audio signal, and enhances the model's understanding and identification ability of the audio signal.
Smart Images

Figure CN120220726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing and analysis, and particularly relates to a method and device for authenticating the authenticity of audio based on a dual-path Transformer. Background Art
[0002] With the development of artificial intelligence technology, audio processing has become an active research field. The audio signal contains rich information, including both semantic-level content, such as the information conveyed by the speaker, and acoustic-level features, such as the frequency and intensity of the sound. In many application scenarios, such as speech recognition, speech synthesis, and speech enhancement, it is necessary to effectively extract and process this information.
[0003] Traditional audio feature extraction methods have limitations in capturing long-term dependencies and complex time-frequency patterns in audio signals, and it is difficult to effectively decouple semantic features and acoustic features. For example, Mel-Frequency Cepstral Coefficients (MFCC) and Low-Frequency Cepstral Coefficients (LFCC). Usually based on signal processing technology, the audio signal is converted into a set of compact feature vectors. MFCC mainly focuses on the local information of the spectrum and ignores the interaction between different frequency bands. Although LFCC considers vocal tract information, it still has deficiencies in dealing with complex speech variations.
[0004] Deep learning models, especially Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN), have been applied to audio feature extraction and have solved these problems to a certain extent. However, RNN is limited by its sequential processing method and has low training efficiency. CNN usually requires multiple layers of convolution to capture long-term dependencies, resulting in a high model complexity. The Transformer model has been widely used in the field of audio processing due to its advantages in processing sequence data. This model can capture long-distance dependencies, which is particularly useful for understanding the complex structure in audio signals. However, existing Transformer-based methods still face challenges in effectively fusing frequency-domain and time-domain information and separating acoustic features and semantic features in audio signals. Summary of the Invention
[0005] The object of the present invention is to provide a method and device for audio authenticity identification based on a dual-path Transformer in view of the above-mentioned existing technical problems, introducing a dual-path Transformer structure to better capture long-term dependencies and time-frequency information in audio signals; processing frequency-domain and time-domain information respectively through the dual-path Transformer structure and combining with the attention mechanism to achieve efficient decoupling of semantic information and acoustic information, so as to better extract and utilize audio information at different levels.
[0006] The technical solution of the present invention is as follows:
[0007] A method for audio authenticity identification based on a dual-path Transformer includes the following steps:
[0008] Preprocessing: Convert the original audio signal into a time-frequency representation, and extract features through LFCC or STFT to obtain the time-domain features and frequency-domain features of the audio signal; after preprocessing, obtain a time-frequency representation feature = {B, F, T} containing time-domain features and frequency-domain features, where B is the batch size, F is the number of frequency channels equal to the number of LFCC linear filters, and T is the number of time frames;
[0009] Convolutional encoder: Encode the input time-frequency representation through multiple 2D convolutional layers to extract local features;
[0010] Process the frequency-domain and time-domain information respectively through the dual-path Transformer structure to form semantic-acoustic information decoupling, including:
[0011] Frequency-domain path processing: Reshape the tensor output by the encoder along the time dimension into [B, F′, C, T′] and input it into the frequency-domain Transformer. The frequency-domain Transformer models the features of each frequency at different time steps, captures the spectral dynamic changes of the audio signal, and extracts acoustic-related features;
[0012] Time-domain path processing: Reshape the tensor output by the encoder along the frequency dimension into [B, T', C, F'], and input it into the time-domain Transformer. The time-domain Transformer models the features of each time step at different frequencies, captures the time-domain envelope changes of the audio signal, and extracts semantic-related features;
[0013] Attention mechanism fusion: Combine the features of the frequency domain and the time domain through the attention mechanism to enhance the model's understanding of the audio signal;
[0014] Post-processing transformation: Adjust the dimension order of the output sequence through aggregation information, linear transformation, and sequence reconstruction operations to obtain a feature that separates semantic information and only contains acoustic information with a shape of [B, C, F′, T′].
[0015] Furthermore, the specific content of the attention mechanism fusion includes the following:
[0016] Construct Query, Key, and Value: Query comes from the output of the frequency-domain Transformer and represents the frequency-domain features:
[0017] Query = x freq ,
[0018] where x freq has a shape of [B, T′, d model , B is the batch size, T′ is the time dimension size, and d model is the model dimension;
[0019] Key and Value come from the output of the time-domain Transformer and represent the time-domain features:
[0020]
[0021] Value = x temp ,
[0022] where x temp has a shape of [B, F′, d model , F′ is the frequency dimension size, represents the transpose of x temp in the frequency and model dimensions;
[0023] Calculate the attention scores: The attention scores represent the attention weights of each time step to each frequency-domain point, and are calculated as the dot product of Query and Key, and then divided by the square root of the model dimension:
[0024]
[0025] The resulting shape is [B, T′, F′];
[0026] Calculate the attention weights: The attention weights are obtained by normalizing the attention scores through the softmax function, so that the sum of the weights of each time step to all frequency-domain points is 1:
[0027] Attention Weights = Softmax(Attention Scores),
[0028] The resulting shape remains [B, T′, F′], which are the attention weights of each time step to each frequency-domain point;
[0029] Applying attention weights: Use attention weights to perform weighted summation on Value to obtain the weighted feature representation at each time step:
[0030] x attended = Attention Weights × Value,
[0031] The resulting shape is [B, T′, d model , and a feature representation weighted based on frequency domain information is obtained for each time step.
[0032] Furthermore, the post - processing transformation specifically includes the following steps:
[0033] Aggregating information: Take the average of the feature x fused through the attention mechanism in the time dimension to aggregate the time - domain information: attended x
[0034] x combined = mean(x attended , dim = 1),
[0035] where mean represents the operation of taking the average in the time dimension;
[0036] Linear transformation: Apply a linear transformation to further transform the aggregated features:
[0037] x final = Linear(x combined ),
[0038] where Linear represents the linear transformation layer, which is used to map the features to the output dimension;
[0039] Sequence reconstruction: Stack the final representations of all frames and apply a final linear transformation to reconstruct the entire sequence to form a complete feature sequence:
[0040]
[0041] where, is the final feature vector of the i - th channel, C is the number of all channels, and the shape of each channel is [B, F′, T′]; stack represents the stacking operation;
[0042] Adjust the dimension order of the output sequence to obtain a feature that separates semantic information and only contains acoustic information with a shape of [B, C, F′, T′].
[0043] Furthermore, the calculation formula for the frequency - domain path processing to extract acoustic - related features is:
[0044] x freq = FreqTransformer(xencoder )
[0045] where x encoder is the output of the encoder; x freq is the output of the frequency-domain Transformer, with shape [B, C, F′, T′]; FreqTransformer represents the frequency-domain Transformer operation.
[0046] Furthermore, the calculation formula for the semantic-related features proposed by the time-domain path processing is:
[0047] x temp = TempTransformer(x encoder )
[0048] where x encoder is the output of the encoder, and x temp is the output of the time-domain Transformer, with shape [B, C, F′, T′]; TempTransformer represents the time-domain Transformer operation.
[0049] Furthermore, the specific steps for extracting the LFCC features include:
[0050] Pre-emphasis: The pre-emphasis process enhances the high-frequency part: y[n] = x[n] - αx[n - 1],
[0051] where x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal;
[0052] Framing: The pre-emphasized signal is segmented into short frames: y p [n] = y[n p , N + n],
[0053] where n p is the frame index, N is the number of samples per frame, and the frame shift is n, obtaining speech frames;
[0054] Windowing: Applying a window function to each speech frame: y w [n] = y p [n] · w[n],
[0055] where w[n] is the window function;
[0056] Frequency-domain conversion: Performing a fast Fourier transform on the windowed signal:
[0057]
[0058] where Y[k] is the spectrum obtained after FFT, and k is the frequency index;
[0059] The fast Fourier transform result is processed by a linear filter to obtain a linear spectrum:
[0060]
[0061] where h[k] is the response of the filter bank and K is the number of filters;
[0062] Logarithmic compression: Take the logarithm of the output of the filter bank to compress the dynamic range:
[0063] Y log [k] = log(Y fb [k]),
[0064] Perform a discrete cosine transform on the logarithmically compressed spectrum to obtain LFCC features:
[0065]
[0066] where m is the coefficient index of LFCC.
[0067] Furthermore, the specific steps for extracting features by the STFT include:
[0068] Pre-emphasis: Pre-emphasis processing enhances the high-frequency part: y[n] = x[n] - αx[n - 1], where x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal; Frame segmentation: Segment the pre-emphasized signal into short frames: y p [n] = y[n p , N + n],
[0069] where n p is the frame index, N is the number of samples per frame, and the frame shift is n, obtaining speech frames;
[0070] Windowing: Apply a window function to each speech frame: y w [n] = y p [n] · w[n],
[0071] where w[n] is the window function;
[0072] Frequency domain conversion: Perform a fast Fourier transform on the windowed signal:
[0073]
[0074] where Y[k] is the spectrum obtained after FFT and k is the frequency index;
[0075] Constructing the spectrogram: Combine the spectra of all frames to form the time-frequency representation of the signal, i.e., the spectrogram. Take the squared modulus of the spectrum as the energy spectrum and take the logarithmic form to compress the dynamic range:
[0076] S[k] = log(|Y[k]| 2 + ∈),
[0077] where Y[k] is the modulus of the spectrum; S[k] is the value at frequency k of the spectrogram; ∈ is a constant used to avoid division-by-zero errors in logarithmic operations.
[0078] Furthermore, the convolutional encoder extracting local features includes:
[0079] Convolutional layer processing: Pass the input time-frequency representation through a 2D convolutional layer to extract local features at different scales:
[0080] x conv = Conv2D(x, W_S, b, k, s, p),
[0081] where x is the input time-frequency representation, W and b are the weights and biases of the convolutional kernel respectively, and k, s, and p are the convolutional kernel size, convolutional stride, and convolutional padding parameters of the convolutional operation;
[0082] Layer normalization processing: Perform layer normalization on the output of each convolutional layer to stabilize the learning process and improve the generalization ability of the model:
[0083] x in = LayerNorm(x conv ),
[0084] where LayerNorm is the layer normalization function;
[0085] Apply the Snake activation function to enhance the model's ability to model complex non-linear relationships:
[0086]
[0087] where α is a learnable parameter used to control the parameters of the activation function;
[0088] Apply Dropout regularization to the activated feature map to prevent overfitting:
[0089] x encoder = Dropout(x act , p),
[0090] where p is the Dropout probability, representing the proportion of features randomly set to zero, and the encoder obtains x encoderThe feature has the shape of [B, VC, F′, T′], where B is the batch size, C is the number of channels output by the encoder, and F′ and T′ are the sizes of the encoded frequency and time dimensions respectively.
[0091] Furthermore, the method further includes the following steps: using a transposed convolution decoder to decode the encoded features to reconstruct a time-frequency representation that separates semantic information and only contains acoustic information;
[0092] Transposed convolution layer processing: Processing the decoupled features through a transposed convolution layer to gradually restore the time-frequency structure of the signal:
[0093] xdeconv = ConvTranspose2D(x, W, b, k, s, p),
[0094] where x is the input time-frequency representation, W and b are the weights and biases of the convolution kernel respectively, and k, s, and p are the convolution kernel size, convolution stride, and convolution padding parameters of the convolution operation;
[0095] Layer normalization processing: Performing layer normalization processing on the output of each convolution layer to stabilize the learning process and improve the generalization ability of the model:
[0096] x in = LayerNorm(x conv )
[0097] where LayerNorm is the layer normalization function;
[0098] Applying the Snake activation function to enhance the model's ability to model complex non-linear relationships:
[0099]
[0100] where α is a learnable parameter used to control the parameters of the activation function;
[0101] Applying Dropout regularization to the activated feature map to prevent overfitting:
[0102] x encoder = Dropout(x act , p)
[0103] where p is the Dropout probability, representing the proportion of features randomly set to zero.
[0104] This application also includes an audio authenticity discrimination device based on a dual-path Transformer, which adopts an audio authenticity discrimination method based on a dual-path Transformer, including:
[0105] A preprocessing module that converts the original audio signal into a time-frequency representation;
[0106] The 2D convolutional encoder module encodes the input time-frequency representation using multiple 2D convolutional layers to extract local features;
[0107] The frequency-domain Transformer module and the time-domain Transformer module process frequency-domain and time-domain information respectively to form decoupling of semantic-acoustic information;
[0108] The attention mechanism fusion module combines the features in the frequency domain and the time domain through the attention mechanism;
[0109] The transposed convolutional decoder module decodes the encoded features using a transposed convolutional decoder to reconstruct the time-frequency representation that only contains acoustic information with the semantic information separated.
[0110] The beneficial effects of the present invention compared with the existing technologies are:
[0111] 1. A method and device for audio authenticity identification based on a dual-path Transformer. The dual-path Transformer structure of the present application realizes the decoupling of semantic and acoustic information in the code by applying different processing and attention encoding layers to frequency-domain and time-domain data respectively. This decoupling is reflected in the device as x freq and x temp two independent processing flows, corresponding to the feature extraction in the frequency domain and the time domain respectively;
[0112] 2. A method and device for audio authenticity identification based on a dual-path Transformer. The present application realizes the fusion of frequency-domain and time-domain features through custom attention calculation, improving the model's understanding and utilization of audio features; the present application introduces the Snake activation function to enhance the model's non-linear modeling ability; in practical applications, the activation function can be realized by self-learning to modify the activation function parameters to enhance the model's ability to capture complex audio features
[0113] 3. A method and device for audio authenticity identification based on a dual-path Transformer. The modular design of the present application is reflected in the code as different class member variables and methods, such as PreTransform and PostTransform, etc., which are responsible for different processing tasks respectively. The optionality proposed by the present application can be realized at the code level by adding or deleting specific processing modules (such as transposed convolutional layers). This design enables the model to be flexibly adjusted according to different downstream tasks to adapt to audio reconstruction, enhancement or other tasks. Through the attention mechanism fusion and subsequent processing steps in the code, a comprehensive feature representation of the audio signal is realized. This representation includes not only the information in the frequency domain and the time domain, but also the fused features, providing a rich feature basis for audio processing. Description of the Drawings
[0114] Figure 1 This is the model structure diagram of the present application.
[0115] Figure 2 This is the LFCC extraction flow chart of the present application.
[0116] Figure 3 These are the effect diagrams of different α for the activation function.
[0117] Figure 4 This is the process flow of the modeling method for the experiment of the present application.
[0118] Figure 5 This is the schematic diagram of the FCN design for the experiment of the present application.
[0119] Figure 6 This is the training and verification flow chart for the experiment of the present application.
[0120] Figure 7 This is the end-to-end model flow chart for the experiment of the present application. Detailed implementation manners
[0121] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0122] The existing audio feature extraction methods and authenticity identification technologies have the following disadvantages and deficiencies:
[0123] 1. The traditional audio feature extraction methods MFCC and LFCC are difficult to capture the long-term dependence relationship and complex time-frequency patterns in audio signals, which limits the comprehensive understanding of audio information and is difficult to effectively capture the long-term dependence relationship in audio signals.
[0124] 2. The existing deep learning models, including RNN, CNN and some Transformer models, have deficiencies in separating semantic information and acoustic information in audio signals.
[0125] 3. The training efficiency of traditional RNN models is relatively low, and the complexity of CNN models is relatively high.
[0126] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.
[0127] Please refer to Figure 1-7 , an audio authenticity identification method based on a dual-path Transformer, as Figure 1 shown, including the following steps:
[0128] Preprocessing: Convert the original audio signal into a time-frequency representation, and extract features through LFCC or Short-Time Fourier Transform (STFT) to obtain the time-domain and frequency-domain features of the audio signal; after preprocessing, obtain the time-frequency representation feature = {B, F, T} containing time-domain and frequency-domain features, where B is the batch size, F is the number of frequency channels equal to the number of LFCC linear filters, and T is the number of time frames;
[0129] Convolutional encoder: Encode the input time-frequency representation through multiple 2D convolutional layers to extract local features;
[0130] Dual-path Transformer: Process the frequency-domain and time-domain information respectively through the dual-path Transformer structure to form semantic-acoustic information decoupling;
[0131] Transposed convolutional decoder: Use the transposed convolutional decoder to decode the encoded features to reconstruct the time-frequency representation that only contains acoustic information after separating the semantic information;
[0132] Transposed convolutional layer processing: Process the decoupled features through the transposed convolutional layer to gradually restore the time-frequency structure of the signal:
[0133] xdeconv = ConvTranspose2D(x, W, b, k, s, p),
[0134] where x is the input time-frequency representation, W and b are the weights and biases of the convolutional kernel respectively, and k, s, and p are the convolutional kernel size, convolutional stride, and convolutional padding parameters of the convolutional operation;
[0135] Layer normalization processing: Perform layer normalization processing on the output of each convolutional layer to stabilize the learning process and improve the generalization ability of the model:
[0136] x in = LayerNorm(x conv ),
[0137] where LayerNorm is the layer normalization function;
[0138] Apply the Snake activation function to enhance the model's ability to model complex non-linear relationships:
[0139]
[0140] Among them, α is a learnable parameter used to control the parameters of the activation function;
[0141] Apply Dropout regularization to the activated feature map to prevent overfitting:
[0142] x encoder = Dropput(x act , p),
[0143] Among them, p is the Dropout probability, representing the proportion of features randomly set to zero.
[0144] Apply the Transformer network in the frequency domain and time domain respectively, and then use the attention mechanism to fuse the outputs of the two paths, so as to effectively separate the acoustic features and semantic features of the audio signal and learn more discriminative audio representations.
[0145] As Figure 2 shown, the specific steps of LFCC feature extraction include:
[0146] Pre-emphasis: The pre-emphasis process enhances the high-frequency part: y[n] = x[n] - αx[n - 1],
[0147] Among them, x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal;
[0148] Framing: Divide the pre-emphasized signal into short frames: y p [n] = y[n p , N + n],
[0149] Among them, n p is the frame index, N is the number of samples per frame, and the frame shift is n, obtaining speech frames;
[0150] Windowing: Apply a window function, such as the Hamming window, to each speech frame: y w [n] = y p [n] · w[n],
[0151] Among them, w[n] is the window function;
[0152] Frequency domain conversion: Perform a fast Fourier transform (FFT) on the windowed signal:
[0153]
[0154] Among them, Y[k] is the spectrum obtained after FFT, and k is the frequency index;
[0155] The fast Fourier transform result is processed by a linear filter to obtain a linear spectrum:
[0156]
[0157] where h[k] is the response of the filter bank and K is the number of filters;
[0158] Logarithmic compression: Take the logarithm of the output of the filter bank to compress the dynamic range:
[0159] Y lig [k]=log(Y fb [k]),
[0160] Perform a discrete cosine transform (DCT) on the logarithmically compressed spectrum to obtain LFCC features:
[0161]
[0162] where m is the coefficient index of LFCC.
[0163] The specific steps for STFT feature extraction include:
[0164] Pre-emphasis: Pre-emphasis processing enhances the high-frequency part: y[n]=x[n]-αx[n - 1], where x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal; Frame segmentation: Segment the pre-emphasized signal into short frames: y p [n]=y[n p ,N + n],
[0165] where n p is the frame index, N is the number of samples per frame, the frame shift is n, and number of speech frames are obtained;
[0166] Windowing: Apply a window function to each speech frame: t w [n]=y p [n]·w[n],
[0167] where w[n] is the window function;
[0168] Frequency domain conversion: Perform a fast Fourier transform on the windowed signal:
[0169]
[0170] where Y[k] is the spectrum obtained after FFT and k is the frequency index;
[0171] Constructing the spectrogram: Combine the spectra of all frames to form the time-frequency representation of the signal, i.e., the spectrogram. Take the square of the modulus (amplitude) of the spectrum as the energy spectrum and take the logarithmic form to compress the dynamic range:
[0172] S[k] = log(|Y[k]| 2 + ∈),
[0173] where Y[k] is the modulus (amplitude) of the spectrum; S[k] is the value at frequency k of the spectrogram; ∈ is a constant used to avoid division-by-zero errors in logarithmic operations.
[0174] The convolutional encoder extracts local features including:
[0175] Convolutional layer processing: Pass the input time-frequency representation through a series of 2D convolutional layers. The kernel size, stride, and padding of each convolutional layer are set as needed to extract local features of different scales:
[0176] x conv = Conv2D(x, W, b, k, s, p),
[0177] where x is the input time-frequency representation, W and b are the weights and biases of the convolutional kernel respectively, and k, s, and p are the convolutional kernel size, convolutional stride, and convolutional padding parameters of the convolutional operation;
[0178] Layer normalization processing: Perform layer normalization on the output of each convolutional layer to stabilize the learning process and improve the generalization ability of the model:
[0179] x in = LayerNorm(x conv ),
[0180] where LayerNorm is the layer normalization function;
[0181] Apply the Snake activation function to enhance the model's ability to model complex non-linear relationships. The Snake activation function combines the periodicity of the sine function and the direct transmission characteristics of the linear function:
[0182]
[0183] where α is a learnable parameter used to control the parameters of the activation function; for the learned parameters, different activation effects are produced, as Figure 3 shown.
[0184] Apply Dropout regularization to the activated feature map to prevent overfitting:
[0185] x encoder = Dropout(x act , p),
[0186] Among them, p is the Dropout probability, representing the proportion of features randomly set to zero. The encoder obtains x encoder features, with the shape of [V, C, F′, T′], where B is the batch size, C is the number of channels output by the encoder, and F′ and T′ are the sizes of the encoded frequency and time dimensions respectively.
[0187] Through the above steps, the convolutional encoder can effectively extract rich local features from the input time-frequency representation, and enhance the expressive power and generalization performance of the model through layer normalization, the Snake activation function, and Dropout regularization.
[0188] The dual-path Transformer includes:
[0189] Frequency-domain path processing: Reshape the tensor output by the encoder along the time dimension to [B, F′, C, T′] and input it into the frequency-domain Transformer. The frequency-domain Transformer models the features of each frequency at different time steps, captures the spectral dynamic changes of the audio signal, and extracts acoustic-related features;
[0190] The calculation formula for the acoustic-related features extracted by the frequency-domain path processing is:
[0191] x freq = FreqTransformer(x encoder ),
[0192] where x encoder is the output of the encoder; x freq is the output of the frequency-domain Transformer, with the shape of [B, C, F′, T′]; FreqTransformer represents the frequency-domain Transformer operation.
[0193] Time-domain path processing: Reshape the tensor output by the encoder along the frequency dimension to [B, T', C, F'], and input it into the time-domain Transformer. The time-domain Transformer models the features of each time step at different frequencies, captures the temporal envelope changes of the audio signal, and extracts semantic-related features;
[0194] The calculation formula for the semantic-related features proposed by the time-domain path processing is:
[0195] x temp = TempTransformer(x encoder ),
[0196] where x encoder is the output of the encoder, x tempis the output of the time-domain Transformer, with the shape of [B, C, F′, T′]; TempTransformer represents the time-domain Transformer operation.
[0197] Attention mechanism fusion: Combine the frequency-domain and time-domain features through the attention mechanism to enhance the model's understanding of audio signals;
[0198] Post-processing transformation: Adjust the dimension order of the output sequence through operations such as aggregating information, linear transformation, and sequence reconstruction to obtain features that separate semantic information and only contain acoustic information, with the shape of [B, C, F′, T′].
[0199] The attention mechanism fusion specifically includes the following:
[0200] Construct Query (query), Key (key), and Value (value): Query comes from the output of the frequency-domain Transformer and represents the frequency-domain features:
[0201] Query = x freq ,
[0202] where x freq has the shape of [B, T′, d model , B is the batch size, T′ is the time dimension size, and d model is the model dimension;
[0203] Key and Value come from the output of the time-domain Transformer and represent the time-domain features:
[0204]
[0205] Value = x temp ,
[0206] where x temp has the shape of [B, F′, d model , F′ is the frequency dimension size, represents the transpose of x temp on the frequency and model dimensions;
[0207] Calculate the attention scores: The attention scores represent the attention weights of each time step to each frequency-domain point, and are calculated as the dot product of Query and Key, divided by the square root of the model dimension (scaled dot product):
[0208]
[0209] The resulting shape is [B, T′, F′];
[0210] Calculating Attention Weights: The attention weights are obtained by normalizing the attention scores using the softmax function, such that the sum of the weights for all frequency domain points at each time step is 1:
[0211] Attention Weights = Softmax(Attention Scores),
[0212] resulting in a shape of [B, T′, F′], which are the attention weights for each frequency domain point at each time step;
[0213] Applying Attention Weights: The attention weights are used to perform a weighted sum on the Value to obtain the weighted feature representation for each time step:
[0214] x attended = Attention Weights × Value,
[0215] resulting in a shape of [B, T′, d model , where a feature representation weighted based on frequency domain information is obtained for each time step.
[0216] The post - processing transformation specifically includes the following steps:
[0217] Aggregating Information: The feature x fused through the attention mechanism attended is averaged in the time dimension to aggregate the time domain information:
[0218] x combined = mean(x attended , dim = 1),
[0219] where mean represents the operation of taking the average in the time dimension;
[0220] Linear Transformation: A linear transformation is applied to the aggregated feature for further feature transformation:
[0221] x final = Linear(x combined ),
[0222] where Linear represents the linear transformation layer used to map the feature to the output dimension;
[0223] Sequence Reconstruction: The final representations of all frames are stacked, and a final linear transformation is applied to reconstruct the entire sequence, forming a complete feature sequence:
[0224]
[0225] where, is the final feature vector of the i-th channel, C is the number of all channels, and the shape of each channel is [B, F′, T′]; stack represents the stacking operation;
[0226] Adjust the dimension order of the output sequence to obtain features that separate semantic information and only contain acoustic information with the shape of [B, C, F′, T′].
[0227] The attention mechanism fusion step allows the model to weight the time-domain information according to the frequency-domain information, enabling it to more flexibly combine the frequency-domain and time-domain features. This fusion strategy not only enhances the model's understanding of audio signals but also improves the effectiveness of feature fusion, providing a cleaner feature representation for subsequent downstream tasks.
[0228] This application also includes an audio authenticity discrimination device based on a dual-path Transformer, which adopts an audio authenticity discrimination method based on a dual-path Transformer, including:
[0229] A preprocessing module that converts the original audio signal into a time-frequency representation;
[0230] A 2D convolutional encoder module that uses multiple 2D convolutional layers to encode the input time-frequency representation and extract local features;
[0231] A frequency-domain Transformer module and a time-domain Transformer module that respectively process the frequency-domain and time-domain information to form semantic-acoustic information decoupling;
[0232] An attention mechanism fusion module that combines the frequency-domain and time-domain features through the attention mechanism;
[0233] A transposed convolutional decoder module that uses a transposed convolutional decoder to decode the encoded features to reconstruct the time-frequency representation that separates semantic information and only contains acoustic information.
[0234] Experimental verification: For the audio authenticity discrimination task, to evaluate the performance of the invention compared with different (Pipeline, pipeline model) and (End-to-End, end-to-end model). In this method, a fully connected network (FCN) is used downstream, such as Figure 4 and Figure 5 Authenticity discrimination experiments were conducted, that is, an FCN was used on the extracted deep representations.
[0235] For all the modeling methods, experiments were conducted using the PyTorch library. The number of epochs (epoch_num) was kept at 20, and the batch size (batch_size) was 32. Cross-entropy was used as the loss function, and Adam was used as the optimizer. The Equal Error Rate (EER) and Accuracy (ACC) were used to evaluate the model performance. The training and validation processes are as Figure 6 shown.
[0236] Four authenticity discrimination benchmark datasets were selected for experiments. The dataset information is shown in Table 1 below.
[0237] Table 1: Experimental Dataset Information
[0238]
[0239] The experimental results of the pipeline method are shown in Table 2 below. HUERT and WAVLM are pre-trained models, LFCC and CQCC are feature extraction methods, CNN-LSTM and GMM are classifier models, and MLP is a multi-layer perceptron.
[0240] Table 2: Experimental Performance of FCN Models with Different PTM Representations
[0241]
[0242] For the end-to-end (E2E) authenticity discrimination method, this application started the experiment with random values, and the experimental process is as Figure 7 , and the E2E model of this application maintained the previous structure. The experimental results are shown in Table 3 below.
[0243]
[0244] Table 3: Experimental Performance of Different E2E Models.
[0245] The above-described embodiments only represent the specific implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application.
Claims
1. An audio authenticity identification method based on dual-path Transformer, characterized in that: The following steps are involved: Preprocessing: Convert the original audio signal into time-frequency representation, and extract features through LFCC or STFT to obtain the time domain features and frequency domain features of the audio signal; After preprocessing, a time-frequency representation containing time domain features and frequency domain features is obtained, feature = {B, F, T}, where B is the batch size, F is the number of frequency channels equal to the number of LFCC linear filters, and T is the number of time frames; Convolutional encoder: encodes the input time-frequency representation through multiple 2D convolutional layers to extract local features; The dual-path Transformer structure processes frequency domain and time domain information respectively to form semantic-acoustic information decoupling, including: Frequency domain path processing: The tensor output by the encoder is reshaped into [B, F′, C, T′] along the time dimension and input into the frequency domain Transformer. The frequency domain Transformer models the characteristics of each frequency at different time steps, captures the dynamic changes of the spectrum of the audio signal, and extracts acoustic related features. Time domain path processing: The tensor output by the encoder is reshaped to [B, T', C, F'] along the frequency dimension and input into the time domain Transformer. The time domain Transformer models the features of each time step at different frequencies, captures the time domain envelope changes of the audio signal, and extracts semantically related features. Attention mechanism fusion: The features of frequency domain and time domain are combined through the attention mechanism; Post-processing transformation: The dimensional order of the output sequence is adjusted by aggregating information, linear transformation, and sequence reconstruction operations to obtain a feature of the shape [B, C, F′, T′] that separates the semantic information and only contains acoustic information.
2. According to claim 1, the audio authenticity identification method based on dual-path Transformer is characterized in that: The attention mechanism fusion specifically includes the following contents: Construct Query, Key and Value: Query comes from the output of frequency domain Transformer, representing frequency domain features: Qurayy=x freq , Among them, x freq The shape is [B,T′,d model ], B is the batch size, T′ is the time dimension size, d model is the model dimension; Key and Value come from the output of the time domain Transformer, representing the time domain features: Value=x temp , Among them, x temp The shape is [B,F′,d model ], F′ is the frequency dimension, Represents x temp transposition in frequency and model dimensions; Calculate the attention score: The attention score represents the attention weight of each time step for each frequency domain point. It is calculated as the dot product of the query and the key, divided by the square root of the model dimension: The resulting shape is [B, T′, F′]; Calculate the attention weight: The attention weight is obtained by normalizing the attention score through the softmax function so that the sum of the weights of all frequency domain points at each time step is 1: Attention Weights=Softmax(Attention Scores), The resulting shape remains [B, T′, F′], with the attention weight of each frequency domain point at each time step; Apply attention weights: Use attention weights to perform weighted summation on Value to obtain a weighted feature representation for each time step: x attended =Attention Weights×Value, The resulting shape is [B, T′, d model ], each time step gets a feature representation weighted by frequency domain information.
3. According to claim 2, the audio authenticity identification method based on dual-path Transformer is characterized in that: The post-processing transformation specifically includes the following steps: Aggregate information: The features x fused through the attention mechanism attended Perform averaging in the time dimension to aggregate time domain information: x combined =mean(x attended ,dim=1), Among them, mean represents the operation of taking the average value in the time dimension; Linear transformation: Apply linear transformation to further transform the aggregated features: x final =Linear(x combined ), Among them, Linear represents the linear transformation layer, which is used to map features to the output dimension; Sequence reconstruction: stack the final representations of all frames and apply the final linear transformation to reconstruct the entire sequence to form a complete feature sequence: in, is the final feature vector of the i-th channel, C is the number of all channels, and the shape of each channel is [B, F′, T′]; stack represents the stacking operation; Adjust the order of the dimensions of the output sequence to obtain a feature of shape [B, C, F′, T′] that separates the semantic information and only contains acoustic information.
4. According to claim 3, the audio authenticity identification method based on dual-path Transformer is characterized in that: The calculation formula for extracting acoustic related features by frequency domain path processing is: x freq =FreqTransTormer(x encoder ), Among them, x encoder is the output of the encoder; x freq It is the output of the frequency domain Transformer, with a shape of [B, C, F′, T′]; FreqTransformer represents the frequency domain Transformer operation.
5. According to claim 3, the audio authenticity identification method based on dual-path Transformer is characterized in that: The calculation formula of the semantic related features proposed by the time domain path processing is: x temp =TempTransformer(x encoder ), Among them, x encoder is the output of the encoder, x temp is the output of the temporal Transformer, with a shape of [B, C, F′, T′]; TempTransformer represents the temporal Transformer operation.
6. The audio authenticity identification method based on dual-path Transformer according to claim 1 is characterized in that: The specific steps of extracting features from LFCC include: Pre-emphasis: Pre-emphasis processing enhances the high frequency part: y[n] = x[n] - αx[n-1], Where x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal; Framing: Divide the pre-emphasized signal into short frames: p [n]=y[n p ,N+n], Among them, n p is the frame index, N is the number of samples per frame, and the frame shift is n, so we get Speech frames; Windowing: Apply a window function to each speech frame: y w [n] = y p [n]·w[n], Where w[n] is the window function; Frequency domain conversion: Perform fast Fourier transform on the windowed signal: Where Y[k] is the spectrum obtained after FFT, and k is the frequency index; The fast Fourier transform result is processed by a linear filter to obtain a linear spectrum: Where h[k] is the response of the filter bank and K is the number of filters; Logarithmic compression: Taking the logarithm of the output of the filter bank compresses the dynamic range: AND log [k]=log(Y fb [k]), Perform discrete cosine transform on the logarithmically compressed spectrum to obtain the LFCC features: Where m is the coefficient index of LFCC.
7. The audio authenticity identification method based on dual-path Transformer according to claim 1 is characterized in that: The specific steps of the STFT feature extraction include: Pre-emphasis: Pre-emphasis processing enhances the high frequency part: y[n] = x[n] - ax[n-1], Where x[n] is the original audio signal, α is the pre-emphasis coefficient, and y[n] is the pre-emphasized signal; Framing: Divide the pre-emphasized signal into short frames: p [n]=y[n p ,N+n], Among them, n p is the frame index, N is the number of samples per frame, and the frame shift is n, so we get Speech frames; Windowing: Apply a window function to each speech frame: y w [n] = y p [n]·w[n], Where w[n] is the window function; Frequency domain conversion: Perform fast Fourier transform on the windowed signal: Where Y[k] is the spectrum obtained after FFT, and k is the frequency index; Construct a spectrogram: combine the spectra of all frames to form a time-frequency representation of the signal, i.e., a spectrogram. Take the square of the modulus of the spectrum as the energy spectrum, and take a logarithmic form to compress the dynamic range: S[k]=log(|Y[k]| 2 +∈), Where Y[k] is the modulus of the spectrum; S[k] is the value at frequency k in the spectrum graph; ∈ is a constant used to avoid division by zero errors in logarithmic operations.
8. The audio authenticity identification method based on dual-path Transformer according to claim 1 is characterized in that: The convolution encoder extracts local features including: Convolutional layer processing: The input time-frequency representation is passed through a 2D convolutional layer to extract local features of different scales: x conv =Conv2D(x,W,b,k,s,p), Where x is the time-frequency representation of the input, W and b are the weight and bias of the convolution kernel, respectively, and k, s, and p are the convolution kernel size, convolution stride, and convolution padding parameters of the convolution operation; Layer Normalization: The output of each convolutional layer is normalized to stabilize the learning process and improve the generalization ability of the model: x in =LayerNorm(x conv ), Among them, LayerNorm is the layer normalization function; Applying the Snake activation function enhances the model's ability to model complex nonlinear relationships: Among them, α is a learnable parameter used to control the parameters of the activation function; Apply Dropout regularization to the activated feature map to prevent overfitting: x encoder =Dropout(x act ,p), Among them, p is the Dropout probability, which indicates the proportion of features randomly set to zero. The encoder obtains x encoder Features, shape is [B, C, F′, T′], where B is the batch size, C is the number of channels output by the encoder, and F′ and T′ are the frequency and time dimensions after encoding, respectively.
9. The audio authenticity identification method based on dual-path Transformer according to claim 1 is characterized in that: The method further includes the following steps: decoding the encoded features using a transposed convolutional decoder to reconstruct a time-frequency representation containing only acoustic information from which semantic information is separated; Transposed convolution layer processing: The decoupled features are processed through the transposed convolution layer to restore the time-frequency structure of the signal: xdeconv=ConvTranspose2D(x,W,b,k,s,p), Where x is the time-frequency representation of the input, W and b are the weight and bias of the convolution kernel, respectively, and k, s, and p are the convolution kernel size, convolution stride, and convolution padding parameters of the convolution operation; Layer normalization: Perform layer normalization on the output of each convolutional layer: x in =LayerNorm(x conv ), Among them, LayerNorm is the layer normalization function; Applying the Snake activation function enhances the model's ability to model complex nonlinear relationships: Among them, α is a learnable parameter used to control the parameters of the activation function; Apply Dropout regularization to the activated feature map to prevent overfitting: x encoder =Dropout(x act ,p), Among them, p is the Dropout probability, which indicates the proportion of features randomly set to zero.
10. An audio authenticity identification device based on dual-path Transformer, characterized in that: The method for identifying the authenticity of audio based on a dual-path Transformer as claimed in any one of claims 1 to 9 comprises: A preprocessing module that converts the raw audio signal into a time-frequency representation; 2D convolutional encoder module, which uses multiple 2D convolutional layers to encode the input time-frequency representation and extract local features; The frequency domain Transformer module and the time domain Transformer module process the frequency domain and time domain information respectively to form semantic-acoustic information decoupling; The attention mechanism fusion module combines the frequency domain and time domain features through the attention mechanism; The transposed convolutional decoder module uses a transposed convolutional decoder to decode the encoded features to reconstruct the time-frequency representation that only contains acoustic information and separates the semantic information.
Citation Information
Cited By
Method and system for identifying forged audio
CN122090874A