Speech enhancement method based on double-branch double-flow attention mechanism
By employing a dual-branch, dual-stream attention mechanism for speech enhancement, the amplitude spectrum and complex spectrum processing of speech are optimized, solving the problems of insufficient fidelity and intelligibility in speech enhancement and achieving efficient speech enhancement results.
Patent Information
- Application Number
- CN202511275601.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-11
AI Technical Summary
Existing speech enhancement methods suffer from poor speech fidelity and intelligibility when dealing with complex background noise, and consume a lot of computational resources. Traditional methods cannot effectively coordinate amplitude and phase compensation.
A speech enhancement method based on a dual-branch dual-stream attention mechanism is adopted. By processing the amplitude spectrum branch and the residual spectrum branch in parallel, and combining the time-frequency dual-stream Transformer module and depthwise separable convolution, the amplitude spectrum and complex spectrum of the speech are optimized. The dual-stream attention mechanism is used to capture local details and global dependencies simultaneously, thereby reducing computational complexity.
It significantly improves the fidelity and clarity of speech enhancement, reduces computational complexity, and solves the problems of inconsistent amplitude and phase compensation and high computational resource consumption in traditional methods.
Smart Images

Figure CN120932670A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a speech enhancement method based on a dual-branch dual-stream attention mechanism. Background Technology
[0002] Currently, in everyday listening environments, speech signals are often interfered with by background noise. This distortion severely reduces speech intelligibility and quality, making many speech-related tasks, such as automatic speech recognition, more complex. Speech enhancement, aiming to recover clean speech from noisy speech, is a core technology for improving communication quality and the performance of hearing aids.
[0003] Traditional speech enhancement methods primarily rely on signal processing techniques such as spectral subtraction and Wiener filtering. However, these methods have limitations in handling complex background noise and preserving speech details. Background noise is diverse, including traffic noise, wind noise, machine noise, and human conversation, and traditional methods are often only applicable to specific types of noise, with limited effectiveness for other types. Traditional methods also frequently ignore the phase information of the speech signal, which can lead to a loss of naturalness and clarity in enhanced speech under high-noise environments.
[0004] It is evident that there is an urgent need for a speech enhancement method based on a dual-branch, dual-stream attention mechanism that offers high fidelity and clarity in speech enhancement. Summary of the Invention
[0005] In view of this, the present disclosure provides a speech enhancement method based on a dual-branch dual-stream attention mechanism, which at least partially solves the problems of inconsistent amplitude and phase compensation, high computational resource consumption, and poor speech fidelity and clarity in the prior art.
[0006] This disclosure provides a speech enhancement method based on a dual-branch, dual-stream attention mechanism, including:
[0007] Step 1: Estimate the amplitude spectrum of the clean speech signal through the amplitude spectrum branch, wherein the amplitude spectrum branch includes two layers of encoders, four layers of time-frequency dual-stream Transformer modules, a fusion module and two layers of decoders connected in sequence.
[0008] Step 2: Fine-tune the amplitude spectrum and phase spectrum through the residual spectrum branch. The residual spectrum branch structure is symmetrical with the amplitude spectrum branch, and the same input signal is processed in parallel.
[0009] Step 3: Add the amplitude spectrum enhanced by the amplitude spectrum branch to the complex spectrum enhanced by the residual spectrum branch to obtain the final enhanced speech signal.
[0010] According to a specific implementation of an embodiment of this disclosure, the encoder includes a two-dimensional convolutional layer, a layer normalization layer, and a parameterized activation function connected in sequence.
[0011] The decoder is a mirror image of the encoder, with all convolutional blocks in the decoder replaced by deconvolutional blocks.
[0012] According to a specific implementation of this disclosure, the processing flow of the time-frequency dual-stream Transformer module includes:
[0013] The input sequence is divided into a complete middle sequence and two side-part sequences.
[0014] Local features are extracted through a convolution module, and the intermediate sequence is further processed by scaling and offset operation blocks to generate independent scaling and offset parameters.
[0015] The processed intermediate sequence and the two side block sequences are input into the two-stream attention block to calculate local attention and global attention respectively;
[0016] The local attention outputs are concatenated along the time dimension and then added to the global attention output to obtain the joint attention result;
[0017] After dimensionality reduction via gating and convolutional modules, the result is added to the residual input, and then sequentially passed through layer normalization, feedforward network, activation function, linear layer, residual connection, and layer normalization output.
[0018] According to one specific implementation of this disclosure, the convolution module uses linear projection and depthwise two-dimensional separable convolution processing sequences to extract time-frequency local features and reduce computational complexity.
[0019] According to one specific implementation of this disclosure, the scaling and offset operation block generates independent scaling and offset parameters for the attention mechanism through learnable parameters gamma and beta, so as to enhance the expressive power of feature interaction.
[0020] According to a specific implementation of this disclosure, the processing flow of the dual-flow attention block includes:
[0021] A linear attention approach is adopted, skipping the Softmax operation, and global attention is calculated accordingly.
[0022] A block-based strategy is adopted to divide the sequence into multiple blocks. The quadratic attention is calculated independently for each block, and the squared ReLU is used instead of Softmax to calculate the local attention.
[0023] According to one specific implementation of this disclosure, the first fully connected layer of the feedforward network is a bidirectional gated loop unit.
[0024] According to a specific implementation of this disclosure, the convolution module of the residual spectrum branch uses an expansion factor of 2 in the linear layer to increase the feature dimension from N to 2N.
[0025] According to one specific implementation of this disclosure, the fusion module is used to integrate the intermediate features of the two branches and perform weighted fusion through an attention mechanism.
[0026] The speech enhancement scheme based on a dual-branch dual-stream attention mechanism in this embodiment includes: Step 1, estimating the amplitude spectrum of a clean speech signal through an amplitude spectrum branch, wherein the amplitude spectrum branch includes two layers of encoders, four layers of time-frequency dual-stream Transformer modules, a fusion module, and two layers of decoders connected in sequence; Step 2, finely adjusting the amplitude spectrum and phase spectrum through a residual spectrum branch, wherein the residual spectrum branch structure is symmetrical with the amplitude spectrum branch and processes the same input signal in parallel; Step 3, adding the amplitude spectrum enhanced by the amplitude spectrum branch to the complex spectrum enhanced by the residual spectrum branch to obtain the final enhanced speech signal.
[0027] The beneficial effects of the embodiments of this disclosure are as follows: By using the scheme of this disclosure, the amplitude spectrum and complex spectrum of speech are optimized respectively through a dual-branch parallel processing architecture. Combined with a time-frequency dual-stream attention mechanism, local details and global dependencies are captured simultaneously. In the encoder-decoder structure, depthwise separable convolution, parameterized scaling offset and GRU temporal modeling are introduced, which significantly improves the fidelity and clarity of speech enhancement. At the same time, by adopting a linear attention and block computing strategy, the computational complexity is greatly reduced while maintaining high performance, effectively solving the problems of inconsistent amplitude and phase compensation and high computational resource consumption in traditional methods. Attached Figure Description
[0028] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 A schematic flowchart illustrating a speech enhancement method based on a dual-branch dual-stream attention mechanism provided in this embodiment of the disclosure;
[0030] Figure 2 A model architecture diagram corresponding to a speech enhancement method based on a dual-branch dual-stream attention mechanism provided in this disclosure embodiment;
[0031] Figure 3 A dual-stream attention block framework diagram provided for embodiments of this disclosure;
[0032] Figure 4 A diagram of a convolution module is provided for an embodiment of this disclosure. Detailed Implementation
[0033] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0034] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0035] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0036] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The illustrations only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0037] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0038] This disclosure provides a speech enhancement method based on a dual-branch dual-stream attention mechanism, which can be applied to speech enhancement processes in scenarios such as communication and data processing.
[0039] See Figure 1This is a flowchart illustrating a speech enhancement method based on a dual-branch, dual-stream attention mechanism provided in an embodiment of this disclosure. Figure 1 As shown, the method mainly includes the following steps:
[0040] Step 1: Estimate the amplitude spectrum of the clean speech signal through the amplitude spectrum branch, wherein the amplitude spectrum branch includes two layers of encoders, four layers of time-frequency dual-stream Transformer modules, a fusion module and two layers of decoders connected in sequence.
[0041] Furthermore, the encoder includes a two-dimensional convolutional layer, a layer normalization function, and a parameterized activation function connected in sequence;
[0042] The decoder is a mirror image of the encoder, with all convolutional blocks in the decoder replaced by deconvolutional blocks.
[0043] Furthermore, the processing flow of the time-frequency dual-stream Transformer module includes:
[0044] The input sequence is divided into a complete middle sequence and two side-part sequences.
[0045] Local features are extracted through a convolution module, and the intermediate sequence is further processed by scaling and offset operation blocks to generate independent scaling and offset parameters.
[0046] The processed intermediate sequence and the two side block sequences are input into the two-stream attention block to calculate local attention and global attention respectively;
[0047] The local attention outputs are concatenated along the time dimension and then added to the global attention output to obtain the joint attention result;
[0048] After dimensionality reduction via gating and convolutional modules, the result is added to the residual input, and then sequentially passed through layer normalization, feedforward network, activation function, linear layer, residual connection, and layer normalization output.
[0049] Furthermore, the convolution module uses linear projection and depthwise two-dimensional separable convolution to process sequences in order to extract time-frequency local features and reduce computational complexity.
[0050] Furthermore, the scaling and offset operation block generates independent scaling and offset parameters for the attention mechanism through learnable parameters gamma and beta, thereby enhancing the expressive power of feature interactions.
[0051] Furthermore, the processing flow of the dual-flow attention block includes:
[0052] A linear attention approach is adopted, skipping the Softmax operation, and global attention is calculated accordingly.
[0053] A block-based strategy is adopted to divide the sequence into multiple blocks. The quadratic attention is calculated independently for each block, and the squared ReLU is used instead of Softmax to calculate the local attention.
[0054] Optionally, the first fully connected layer of the feedforward network is a bidirectional gated recurrent unit.
[0055] Optionally, the convolutional module of the residual spectrum branch uses a dilation factor of 2 in the linear layer to increase the feature dimension from N to 2N.
[0056] Furthermore, the fusion module is used to integrate the intermediate features of the two branches and perform weighted fusion through an attention mechanism.
[0057] Step 2: Fine-tune the amplitude spectrum and phase spectrum through the residual spectrum branch. The residual spectrum branch structure is symmetrical with the amplitude spectrum branch, and the same input signal is processed in parallel.
[0058] Step 3: Add the amplitude spectrum enhanced by the amplitude spectrum branch to the complex spectrum enhanced by the residual spectrum branch to obtain the final enhanced speech signal.
[0059] In specific implementation, this disclosure proposes a speech enhancement method based on a dual-branch dual-stream attention mechanism (Dual-Branch Dual-Stream Attention Model, DB-DSAM). By combining parallel amplitude masking and complex refinement branches with time-frequency dual-stream attention, it achieves efficient and high-fidelity speech enhancement. The model architecture is as follows: Figure 2 As shown. This method is based on a dual-branch neural network with an encoder-decoder structure. Step 1: The amplitude spectrum branch (upper branch) is used to estimate the amplitude spectrum of the clean speech signal. Step 2: The residual spectrum branch (lower branch) further refines the amplitude and phase spectra. Step 3: The enhanced amplitude spectrum information from the upper branch is added to the enhanced complex spectrum information from the lower branch to further enhance the speech signal, thereby alleviating the amplitude and phase compensation problem.
[0060] Specifically, the two branches run in parallel through their corresponding two-layer encoder, four-layer time-frequency dual-stream Transformer, fusion module, and two-layer decoder. The encoder consists of 2-D convolutional layers, followed by layer normalization (LN) and parameterized ReLU (PReLU) activation. The decoder is a mirror image of the encoder, in which all convolutional blocks are replaced by deconvolutional blocks.
[0061] The four-layer time-frequency dual-stream Transformer module in the middle is the key to this model. This structure consists of... Figure 2As shown. This module is based on a gated single-head Transformer architecture with dual-stream attention, employing local and global self-attention modes. It performs full self-attention operations on local segments and linearized, low-cost self-attention operations on the global segment. Furthermore, its powerful attention-gating mechanism allows for significantly simplified computational complexity of single-head self-attention, followed by the use of a GRU for further refined temporal modeling of the sequence. A time-frequency dual-stream attention module mainly consists of four convolutional modules, scaling and offset operation blocks, a dual-stream attention block, and a gated recurrent unit (GRU). The convolutional modules are as follows... Figure 4 As shown.
[0062] Signal processing in a time-frequency dual-stream Transformer module involves the following steps:
[0063] Step 1: The sequence is first divided into three parts: the middle signal is the complete sequence, and the two sides are the segmented sequences.
[0064] Step 2: Sequence processing is performed in the convolution module. The convolution module uses linear projection and depthwise convolution to process the sequence to extract local time-frequency features. The core of this module is depthwise two-dimensional separable convolution, which is used to extract local information while reducing computational complexity and the number of parameters.
[0065] Step 3: After the intermediate signal is processed by the convolution module, it enters the scaling and offset operation block for further processing. The function of the scaling and offset operation block is to generate independent scaling and offset parameters for multi-head attention through learnable parameters gamma and beta, so that different attention mechanisms can learn different feature interaction patterns and enhance expressive power.
[0066] Step 4: Use the intermediate signals from the scaling and offset operation blocks, along with the block sequences from both sides after passing through their respective convolutional modules, as input to the two-stream attention module, and then output the final joint attention result.
[0067] Step 5: The final joint attention result is input into the subsequent convolutional module through a gating mechanism to reduce the feature dimension from 2N back to N, and then added to the residual input.
[0068] Step Six: The process sequentially includes layer normalization, a GRU-based feedforward network, ReLU activation, a linear layer, residual connections, and layer normalization. The feedforward network uses a bidirectional GRU to replace the first fully connected layer in the traditional Transformer.
[0069] like Figure 3 As shown, the signal processing procedure of the dual-stream attention module includes the following steps:
[0070] Step 1: Split the input of the intermediate sequence into four heads (quadratic attention query, linear attention query, quadratic attention key, and linear attention key), which are used for quadratic and linear attention calculations respectively.
[0071] Step 2: The sequences U and V obtained after processing by the convolutional modules of the upper and lower branches are input into the dual-stream attention block. The linear layers of the convolutional modules of the upper and lower branches use an expansion factor of 2 to increase the feature dimension from N to 2N.
[0072] Step 3: Perform computational operations on the three inputs above, mainly in two parts: For global attention, we use the low-cost linearization form of Equation 1 to capture the long-distance global interaction between sequences V and U, where β=1 / S is the scaling factor. We reduce computation by skipping the Softmax operation. For local attention, we use the block-based strategy of Equation 2, dividing the sequence into H blocks of size P (H×P≥S), with each block independently calculating secondary attention. Here, γ=1 / P is the scaling factor. We use squared ReLU instead of softmax in Multi-Head Self-Attention (MHSA) to optimize computational performance, and the attention is shared and calculated only once.
[0073] Step 4: As shown in Formula 3, concatenate all local attention outputs along the time dimension to restore the complete sequence.
[0074] Step 5: As shown in Formula 4, add the local attention and global attention to obtain the final joint attention result.
[0075] , (1)
[0076] (2)
[0077] (3)
[0078] (4)
[0079] in, It is a scaling factor. It refers to the query and key of global attention. V represents the transpose of the key matrix; V and U are the sequence values obtained after processing by the convolutional modules of the upper and lower branches, respectively. These are the results of the upper and lower branches after global attention calculation, respectively. It is the square of the ReLU activation function. This is the query for the h-th block. It is the transpose of the key matrix of the h-th block. It is the value of the upper and lower branches of the h-th block; These are the results of the upper and lower branches after local attention calculations, respectively. These are the complete results after calculating the local attention for the upper and lower branches, respectively. These are the final joint attention results for the upper and lower branches, respectively.
[0080] This embodiment presents a speech enhancement method based on a dual-branch, dual-stream attention mechanism. By proposing a Transformer speech enhancement model with a dual-branch, dual-stream attention mechanism, it addresses the shortcomings of existing speech enhancement techniques, enhancing noisy speech in a highly efficient and low-complexity manner. In terms of speech enhancement effect, the model leverages the powerful sequence modeling capabilities and parallel computing advantages of the Transformer, combined with a dual-branch architecture, to effectively improve speech quality and readability. From the perspective of computational efficiency and complexity, the dual-stream attention technique, through a combination of local block-based secondary attention mechanisms and a global linear efficient attention mechanism, significantly reduces the model's computational complexity while maintaining performance. Compared to traditional multi-head attention mechanisms, when processing long sequences (with a large S), the computational complexity O(S²) of traditional attention mechanisms increases significantly, making them unsuitable for practical applications. The dual-stream attention mechanism captures long-range interactions globally using a linear approach, reducing computation by skipping the Softmax operation. Locally, it divides V, U, Q, and K into H non-overlapping blocks of size P (with zero-padding for boundary handling) and applies secondary attention independently. Finally, it replaces Softmax with squared ReLU to optimize performance. This reduces the computational complexity from O(S²) to near-linear O(S²). D+S P), D, P≪S. This design significantly improves the training and inference efficiency of long sequence speech enhancement tasks while maintaining model accuracy.
[0081] The method of this disclosure will be further described below with reference to a specific embodiment. The model training corresponding to the method of this disclosure uses the publicly available dataset VOICEBANK-DEMAND, which contains 28 speakers in the training set and 2 invisible speakers in the test set. The training set contains 11,572 speech samples, and the test set contains 824 noisy speech samples. For the training set, at four signal-to-noise ratios (SNRs) of 0dB, 5dB, 10dB, and 15dB, clean speech is noise-added using one of 10 types of noise, including 2 types of human noise and 8 types of noise from the DEMAND dataset. The test set is generated using 5 invisible noise samples selected from the DEMAND dataset, with SNRs of 2.5dB, 7.5dB, 12.5dB, and 17.5dB, respectively.
[0082] The hyperparameter settings of the model are shown in Table 1. The dimensions of the input and output of each layer of the encoder and decoder are given by (batch size × number of channels × number of features × time frame), and the hyperparameters are given in the form of (kernel size, stride, number of channels). The dual-stream attention module is given by (actual input and output dimensions and number of layers). The decoder is a mirror image of the encoder. The encoder consists of the following four layers: Temporal 2D convolutional block-1 has an input dimension of 1×1×161×T, hyperparameters of 1×1, (1, 1), 64, and an output dimension of 1×64×161×T; Temporal 2D convolutional block-2 has an input dimension of 1×64×161×T, hyperparameters of 1×3, (1, 2), 64, and an output dimension of 1×64×161×T; Complex 2D convolutional block-1 has an input dimension of 1×2×161×T, hyperparameters of 1×1, (1, 1), 64, and an output dimension of 1×64×161×T; Complex 2D convolutional block-2 has an input dimension of 1×64×161×T, hyperparameters of 1×3, (1, 2), 64, and an output dimension of 1×64×161×T. The two-stream attention module consists of the following two layers: a temporal Transformer with an input dimension of 128×161×T, hyperparameters of 128, 64, and 4, and an output dimension of 64×161×T; and a complex Transformer with an input dimension of 128×161×T, hyperparameters of 128, 64, and 4, and an output dimension of 64×161×T. The decoder consists of the following four layers: Temporal 2D deconvolution block-1 has an input dimension of 1×64×161×T, hyperparameters of 1×3, (1, 2), 64, and an output dimension of 1×64×161×T; Temporal 2D deconvolution block-2 has an input dimension of 1×64×161×T, hyperparameters of 1×1, (1, 1), 64, and an output dimension of 1×1×161×T; Complex 2D deconvolution block-1 has an input dimension of 1×64×161×T, hyperparameters of 1×3, (1, 2), 64, and an output dimension of 1×64×161×T; Complex 2D deconvolution block-2 has an input dimension of 1×64×161×T, hyperparameters of 1×1, (1, 1), 64, and an output dimension of 1×2×161×T.
[0083] This experiment was programmed in Python and implemented using the deep learning framework PyTorch. All speech data in the dataset was resampled at 16kHz, using a 320-point STFT to obtain 161-dimensional spectral features. A Hanning window of 25ms was selected, with 50% overlap between any two consecutive frames. Power compression was performed towards amplitude while maintaining phase invariance, and the optimal compression factor was set to 0.5. Optimization was performed using the Adam optimizer with a batch size of 1, training for 80 generations at a learning rate of 10. -4.
[0084] The model employs six evaluation metrics to reflect the performance of speech enhancement: Perceptual Speech Quality Assessment (PESQ), Short-Time Objective Speech Intelligibility (STOI), Naturalness of Enhanced Speech Signal (CSIG), Background Noise Suppression (CBAK), Overall Speech Quality (COVL), and Signal-to-Distortion Ratio (SSNR). With 2.93M parameters, this model achieves PESQ: 3.25, STOI: 95.3, CSIG: 4.51, CBAK: 3.80, COVL: 3.94, and SSNR: 10.31.
[0085] It should be understood that the various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof.
[0086] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A speech enhancement method based on a dual-branch, dual-stream attention mechanism, characterized in that, include: Step 1: Estimate the amplitude spectrum of the clean speech signal through the amplitude spectrum branch, wherein the amplitude spectrum branch includes two layers of encoders, four layers of time-frequency dual-stream Transformer modules, a fusion module and two layers of decoders connected in sequence. Step 2: Fine-tune the amplitude spectrum and phase spectrum through the residual spectrum branch. The residual spectrum branch structure is symmetrical with the amplitude spectrum branch, and the same input signal is processed in parallel. Step 3: Add the amplitude spectrum enhanced by the amplitude spectrum branch to the complex spectrum enhanced by the residual spectrum branch to obtain the final enhanced speech signal.
2. The method according to claim 1, characterized in that, The encoder includes a two-dimensional convolutional layer, a layer normalization function, and a parameterized activation function connected in sequence. The decoder is a mirror image of the encoder, with all convolutional blocks in the decoder replaced by deconvolutional blocks.
3. The method according to claim 1, characterized in that, The processing flow of the time-frequency dual-stream Transformer module includes: The input sequence is divided into a complete middle sequence and two side-part sequences. Local features are extracted through a convolution module, and the intermediate sequence is further processed by scaling and offset operation blocks to generate independent scaling and offset parameters. The processed intermediate sequence and the two side block sequences are input into the two-stream attention block to calculate local attention and global attention respectively; The local attention outputs are concatenated along the time dimension and then added to the global attention output to obtain the joint attention result; After dimensionality reduction via gating and convolutional modules, the result is added to the residual input, and then sequentially passed through layer normalization, feedforward network, activation function, linear layer, residual connection, and layer normalization output.
4. The method according to claim 3, characterized in that, The convolution module uses linear projection and depthwise separable convolution to process sequences, extracting time-frequency local features and reducing computational complexity.
5. The method according to claim 3, characterized in that, The scaling and offset operation blocks generate independent scaling and offset parameters for the attention mechanism through learnable parameters gamma and beta, thereby enhancing the expressive power of feature interactions.
6. The method according to claim 3, characterized in that, The processing flow of the dual-flow attention block includes: A linear attention approach is adopted, skipping the Softmax operation, and global attention is calculated accordingly. A block-based strategy is adopted to divide the sequence into multiple blocks. The quadratic attention is calculated independently for each block, and the squared ReLU is used instead of Softmax to calculate the local attention.
7. The method according to claim 3, characterized in that, The first fully connected layer of the feedforward network is a bidirectional gated cyclic unit.
8. The method according to claim 1, characterized in that, The convolutional module of the residual spectrum branch uses a dilation factor of 2 in the linear layer to increase the feature dimension from N to 2N.
9. The method according to claim 1, characterized in that, The fusion module is used to integrate the intermediate features of the two branches and perform weighted fusion through an attention mechanism.