Speech enhancement method and system based on selective state space model Mamba
By adopting the selective state space models Mamba and TF-mamba blocks in the speech enhancement system, the problem of computational complexity of the self-attention mechanism is solved, and a more efficient and accurate speech enhancement effect is achieved.
Patent Information
- Application Number
- CN202510369680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The computational complexity of self-attention mechanisms in existing voice enhancement systems makes it difficult to effectively deploy under limited computing resources.
The speech enhancement method based on the selective state space model Mamba is adopted to obtain the time-frequency domain signal through short-time Fourier transform, and the characteristics are extracted using the generator network and the TF-mamba block, and the forward and backward dependencies of speech signals at different resolutions are simulated, the local and global features of long-sequence speech are captured, and the enhancement signal is predicted through complex decoder and mask decoder.
It improves the efficiency and accuracy of the speech enhancement system, reduces the computational complexity, and makes the enhanced speech close to clean speech in subjective auditory and objective evaluation indicators.
Smart Images

Figure CN119993176A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech signal processing, in the direction of speech denoising, reverberation and enhancement, and mainly relates to a speech enhancement method and system based on a selective state space model Mamba. Background Art
[0002] In real-life speech applications, speech quality and intelligibility depend on the performance of the underlying speech enhancement (SE) systems, such as speech denoising, dereverberation, and acoustic echo cancellation. Therefore, the SE framework is an indispensable component in modern automatic speech recognition, telecommunication systems, and hearing aids. More and more research continues to try to push the performance limits of current SE systems, and most of these approaches have taken advantage of the latest advances in deep learning techniques and the increasing number of available speech datasets.
[0003] Currently, many SE models use transformers and conformers based on self-attention mechanisms to capture long-term dependencies in waveforms or spectrograms. In speech tasks, training usually includes the entire speech signal, which produces a long context, and the quadratic complexity of the self-attention mechanism relative to the sequence length poses a huge challenge to deploying these models with limited computing resources. Therefore, a more efficient and accurate method is urgently needed to improve the efficiency of SE tasks. Summary of the invention
[0004] The present invention aims at the problem of high computational complexity of the self-attention mechanism in the prior art, and proposes a speech enhancement method and system based on the selective state space model Mamba. First, the distorted speech is subjected to short-time Fourier transform to obtain a time-frequency domain signal, and the amplitude, real part and imaginary part of the time-frequency domain signal are spliced into a three-channel signal. The spliced signal is then expanded from three channels to multiple channels through the first convolution block of the encoder of the generator network. The encoder's expanded dense convolution layer is used to extract features of different resolutions and increase the receptive field. The last convolution block of the encoder is used to reduce the frequency dimension of the signal to 1 / 2 of the original to reduce the computational complexity. The extended signal is feature enhanced through N TF-mamba blocks, and the forward and backward dependencies of speech signals at different resolutions are simulated to capture the local and global features of long-sequence speech. The complex decoder and the mask decoder are used to predict the amplitude and phase of the enhanced signal to obtain the enhanced speech after denoising and dereverberation. A discriminator based on speech quality evaluation indicators is also used to distinguish enhanced speech from clean speech, and the gradients of the generator and discriminator networks are updated. Multiple rounds of network model training are repeated until the model converges, so that the enhanced speech is close to the clean speech in terms of both subjective hearing and objective evaluation indicators. The present invention improves the network training efficiency while ensuring the enhanced speech signal indicators.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is: a speech enhancement method based on a selective state space model Mamba, comprising at least the following steps:
[0006] S1. Data preprocessing: distorted speech signal input Preprocessing is performed to obtain the time-frequency domain signal through short-time Fourier transform, where L represents the time domain length of the speech; the amplitude, real part and imaginary part of the time-frequency domain signal are concatenated as the input signal Among them, B represents the batch size, T is the number of frames, F is the frequency, and 3 is the number of channels;
[0007] S2, signal expansion: The time-frequency domain signal obtained by the transformation in step S1 is dimensionally expanded through an encoder, from 3 channels to C channels; the encoder includes two convolution blocks and an expanded dense convolution layer, wherein the first convolution block is used to convert the input signal The channel dimension is expanded from 3 to C, forming a signal The expanded signal is obtained by expanding the dense convolution layer to extract features of different resolutions and increase the receptive field The second convolutional block is used to convert the signal Y″ after passing through the dilated dense convolutional layer stft The frequency F is downsampled to The expanded signal is
[0008] S3, feature enhancement: Through N TF-mamba blocks, the signal expanded in step S2 is feature enhanced to simulate the forward and backward dependencies of speech signals at different resolutions, and the feature-enhanced signal is obtained. The TF-mamba block is composed of a Time Mamba block and a Frequency Mamba block. Each Mamba block adopts a bidirectional SSM mode, and the input signal is processed in parallel by the forward Mamba and the backward Mamba.
[0009] S4, speech signal enhancement: the signal of step S3 is processed by a mask decoder and a complex decoder for amplitude and phase respectively to obtain predicted amplitude and phase, which are combined to obtain a predicted spectrum signal, and the predicted spectrum signal is subjected to inverse power compression and inverse short-time Fourier transform (ISTFT) to obtain an enhanced speech signal;
[0010] S5, enhanced speech discrimination: The enhanced speech signal obtained in step S4 and the clean speech signal are discriminated by the metric discriminator, and the gradients of the generator and the discriminator are updated after the discriminator is trained; the input signal Y of step S1 is inInput the generator after updating the gradient, repeat steps S2-S4 to obtain new enhanced speech, perform enhanced speech discrimination through the discriminator after gradient updating, repeat steps S2-S5 to update the generator network and discriminator to achieve speech enhancement.
[0011] As an improvement of the present invention, the specific method for obtaining the video domain signal in step S1 is: by STFT operation, the input distorted speech signal waveform is transformed into Convert to complex spectrum Where T and F represent the number of frames and frequency points respectively, and the compressed spectrum Y is obtained by power compression:
[0012]
[0013] Where c represents the power compression index, Y m , Y p , Y r and Y i is the amplitude, phase, real component and imaginary component of the spectrum graph Y; r , Y i and Y m Connect as input signal B represents the batch size, T represents the number of frames, and F represents the frequency.
[0014] As an improvement of the present invention, the encoder of step S2 is composed of two convolution blocks and a dilated dense convolution layer dilated DenseNet, each convolution block includes a two-dimensional convolution layer, an instance normalization layer and a PReLU activation function layer. The first convolution block takes the input Y stft Expanded to C channels, the dilated dense convolution layer contains four convolution blocks with dense residual connections. The dilation factors of the four convolution blocks are {1, 2, 4, 8} to effectively increase the receptive field; the second convolution block is responsible for reducing the frequency dimension to To reduce complexity, get the expanded signal
[0015] As another improvement of the present invention, in the encoder of step S2, the convolution kernel size of the first convolution block is (1, 1), and the step size is (1, 1); the convolution kernel size of the second convolution block is (1, 3), the step size is (1, 2), and the padding is (0, 1).
[0016] As another improvement of the present invention, in step S3, N TF-Mamba blocks are used to obtain global and local information of the speech signal, each TF-Mamba is composed of a Time Mamba block and a Frequency Mamba block, and the Time Mamba block has the same structure as the Frequency Mamba. The specific steps are:
[0017] S31: The input signal obtained in step S2 Reshape it into The local and global information is processed along the time domain dimension T by the Time Mamba block and then combined with Y tm Through skip connection
[0018] Y t =TimeMamba(Y tm )+Y tm
[0019] S32: The signal obtained from S31 Reshape it into The local and global information is processed along the frequency domain dimension F by the Frequency Mamba block and then combined with Y fm Through skip connection
[0020] Y f =FrequencyMamba(Y fm )+Y fm
[0021] S33: Reshape Y mamba Enter the next TF-Mamba block. A total of N TF-Mamba blocks need to be entered, that is, steps S31 and S32 need to be repeated N times to finally obtain the feature-enhanced signal.
[0022] As another improvement of the present invention, the step S4 specifically includes the following steps:
[0023] S41: In the mask decoder, a convolutional block is used The number of channels is compressed, and then the final amplitude is predicted by another convolutional layer with PReLU activation.
[0024] S42: In the complex decoder, the real and imaginary parts of the final enhanced signal are predicted by the convolution block
[0025] S43: The final amplitude predicted in step S41 Sum signal phase Y p The enhanced frequency domain signal is obtained by combining;
[0026] S44: The enhanced frequency domain signal obtained in step S43 is compared with the predicted output of the complex decoder in step S42. Add them in sequence to get the final enhanced frequency domain signal:
[0027]
[0028] S45: In the complex spectrum Perform inverse power compression and inverse short-time Fourier transform ISTFT to obtain the time domain signal
[0029] As another improvement of the present invention, the metric discriminator of step S5 is composed of four convolution blocks, each convolution block is composed of a two-dimensional convolution layer, an instance normalization and a PReLU activation layer, and the convolution block is followed by a global average pooling layer, two feedforward layers and a sigmoid activation layer.
[0030] As a further improvement of the present invention, the loss function of the generator in step S4 is is the amplitude loss function in the time-frequency domain And the complex loss function
[0031]
[0032] Among them, α represents the weight, X m represents the amplitude of clean speech, X r and X i Represents the real and imaginary parts of clean speech in the time-frequency domain;
[0033] The generator loss is the associated generator loss and time domain loss Combination of:
[0034]
[0035] in, represents the enhanced speech signal, X represents the clean speech signal, ‖·‖1 represents l 1 norm, ‖·‖2 represents l 2 norm, γ1, γ2, γ3 represent the corresponding loss weights.
[0036] In order to achieve the above object, the present invention also adopts a technical solution: a speech enhancement system based on the selective state space model Mamba, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0037] Compared with the prior art, the present invention has the following beneficial effects: the present invention provides a speech enhancement method and system based on the selective state space model Mamba, combines the amplitude and phase information of the time-frequency domain speech signal as input data for processing, and simultaneously estimates the amplitude and phase of the enhanced signal through the generator model. Combined with the TF-Mamba module, the efficiency of the generator model in processing complex long sequence data is effectively improved, and the multi-scale information capture capability is enhanced. The discriminator uses the speech perception quality evaluation score PESQ as the discrimination index, optimizes the performance of the speech enhancement model, and makes the enhanced speech perception closer to clean speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flow chart of the steps of the method of the present invention;
[0039] Figure 2 It is a schematic diagram of the data preprocessing process in the method of the present invention;
[0040] Figure 3 It is a structural schematic diagram of a speech enhancement network generator in the method of the present invention;
[0041] Figure 4 is a schematic diagram of the TF-Mamba block structure in the method of the present invention;
[0042] Figure 5 Schematic diagram of the structure of the speech enhancement network discriminator in the method of the present invention;
[0043] Figure 6 It is a clean speech spectrogram of the test set in the test example of the present invention;
[0044] Figure 7 It is a spectrogram of the noisy reverberation speech in the test example of the present invention;
[0045] Figure 8 This is the enhanced speech spectrogram in the test example of the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0047] Example 1
[0048] like Figure 1 As shown, the present invention provides a speech enhancement network method based on a selective state space model Mamba, comprising the following steps:
[0049] Step S1: Preprocess the input distorted speech signal Y, obtain the spectrum signal through short-time Fourier transform STFT, perform power compression on the signal, and stack its amplitude and phase information as the network input signal
[0050] The distorted speech signal is considered to be the speech signal obtained by superimposing the direct signal, the reverberation signal and the background noise. The background noise is unrelated to the direct signal, so the denoising task is considered to be a speech separation task; the reverberation signal is the attenuation and delay signal obtained by the reflection of the direct signal on the walls and ceiling of the room in a closed acoustic environment, which can be regarded as the convolution of the direct signal and the room impulse response. The reverberation signal is related to the direct signal. The distorted speech signal can be expressed as:
[0051] y(t)=x(t)*h(t)+n(t)
[0052] Where y(t) is the distorted speech, x(t) is the desired clean speech, h(t) is the background noise, and * represents the convolution operation.
[0053] For distorted speech waveforms The STFT operation first converts the waveform into a complex spectrum Where T and F represent the time dimension and frequency dimension respectively. Then, through power compression, the importance of smaller sounds is equal to that of larger sounds, which is closer to human perception of sound, and the compressed spectrum Y is obtained:
[0054]
[0055] c represents the power compression index, ranging from 0 to 1, and in this embodiment, c=0.3. m , Y p , Y r and Y i is the amplitude, phase, real component and imaginary component of the spectrum graph Y. r and Y i and Y m Connect as input signal B represents the batch size, T represents the number of frames, and F represents the frequency. The process is as follows Figure 2 The data preprocessing flow chart is shown in the figure.
[0056] Step S2: The signal in step S1 is dimensionally expanded by an encoder. The encoder consists of an expanded dense convolution layer and two convolution blocks. The input signal is expanded by the encoder to
[0057] The encoder structure is as follows Figure 3 As shown in the encoder section of Figure 1, the encoder consists of two convolutional blocks and a dilated DenseNet layer. Each convolutional block includes a 2D convolutional layer, an instance normalization layer, and a PReLU activation function layer. The first convolutional block takes the input Expanded to C channels, the dilated dense convolution layer contains four convolution blocks with dense residual connections. The dilation factors of the four convolution blocks are {1, 2, 4, 8} to effectively increase the receptive field. The second convolution block is responsible for reducing the frequency dimension to To reduce complexity.
[0058] In this embodiment, input First, a two-dimensional convolution block is used. The input channel of the convolution block is 3, the number of output channels is C, the convolution kernel size is designed to be (1, 1), the step size is (1, 1), and the instance normalization InstanceNorm and PReLU activation layer are used to obtain
[0059] Then Y′ stft Enter the dilated DenseNet with a depth of 4, which consists of four layers of two-dimensional convolutional blocks. The dilation factors of the four convolutional blocks are {1, 2, 4, 8}. The dilated convolution is combined with dense connections to gradually expand the receptive field and output Last Y″ stft Through the downsampling convolution layer, the number of input channels and output channels of the two-dimensional convolution block are both 64, the convolution kernel size is designed to be (1, 3), the step size is (1, 2), the padding is (0, 1), and the instance normalization InstanceNorm and PReLU activation layer are used. The function of the second convolution block is to convert Y″ stft Frequency F downsampled to To reduce the complexity, we get
[0060] Step S3: The signal in step S2 is passed through N TF-mamba blocks to enhance its time domain and frequency domain characteristics.
[0061] The TF-mamba block structure is as follows Figure 4 As shown, the TF-mamba block consists of the Time Mamba block and the Frequency Mamba block. The Time Mamba block and the Frequency Mamba block have the same structure. The specific steps are:
[0062] S31: The input signal obtained in step S2 Reshape it into The local and global information is processed along the time domain dimension T by the Time Mamba block and then combined with Y tm Through skip connection
[0063] Y t =TimeMamba(Y tm )+Y tm
[0064] S32: The signal obtained from S31 Reshape it into The local and global information is processed along the frequency domain dimension F by the Frequency Mamba block and then combined with Y fm Through skip connection
[0065] Y f =FrequencyMamba(Y fm )+Y fm
[0066] S33: Reshape Y mamba Enter the next TF-Mamba block. A total of N TF-Mamba blocks need to be entered, that is, steps S31 and S32 need to be repeated N times to finally obtain the feature-enhanced signal.
[0067] The TF-Mamba block structure consists of two mamba blocks: Time Mamba and Frequency Mamba. They perform forward and backward scanning respectively, integrating local and global information. The signal operation steps in the mamba block are as follows:
[0068] a. For the sample input data First, the forward scan is performed through the forward mamba block, and then the RMSNorm activation function is used, and then the Y eg Get Y through skip connection f :
[0069] Y f =RMS(FMamba(Y eg ))+Y eg
[0070] b. For example input data Flip it to get Y flip , and then scanned backwards through the backward mamba block, and then activated with the RMSNorm function, and then with Y flip Get Y through skip connection b :
[0071] Y b =RMS(FMamba(Y flip ))+Y flip
[0072] c. Y f and Y b After concatenation along the channel dimension C, transposed convolution is performed to obtain the output Y out :
[0073] Y out =TransConv(Concat(Y f ,Y b ))
[0074] Among them, RMS, FMamba, BMamba, Flip, and Concat represent forward output, backward output, RMS normalization, forward Mamba, backward Mamba, flip, and concatenation operations respectively;
[0075] Step S4: The signal of step S3 The amplitude and phase are processed by two independent decoders respectively to obtain the predicted amplitude and phase, and the predicted spectrum signal is obtained by combining them. The predicted signal is subjected to inverse power compression and inverse short-time Fourier transform (ISTFT) to obtain the enhanced speech signal.
[0076] like Figure 3 As shown in the decoder part of , both the mask decoder and the complex decoder consist of a dilated dense convolution layer, a dilated DenseNet, and a two-dimensional convolution block. The dilated DenseNet structure is the same as that in the encoder. After passing through the dilated DenseNet, the frequency F′ is upsampled to F using a sub-pixel convolutional layer. In the mask decoder, a convolutional block is first used to compress the number of channels to 1, and then another convolutional layer with PReLU activation is used to predict the final amplitude. In the complex decoder, the real and imaginary parts of the final enhanced signal are predicted by the convolution block.
[0077] First, the predicted enhanced signal amplitude Sum signal phase Y p Combined to obtain enhanced frequency domain signal, and then predicted by complex decoder output Add them in sequence to get the final enhanced frequency domain signal:
[0078]
[0079] Then in the complex spectrum Perform inverse power compression and inverse short-time Fourier transform ISTFT to obtain the time domain signal
[0080] Step S5: The enhanced signal obtained in step S4 is discriminated with the clean speech signal through a discriminator, and the discriminator is trained to estimate the enhanced signal PESQ score to update the gradients of the generator and the discriminator so that the model converges faster.
[0081] like Figure 5As shown, the clean speech signal X and the predicted enhanced signal Input to the metric discriminator, the discriminator consists of four convolution blocks, each of which consists of a two-dimensional convolution layer, instance normalization and PReLU activation layer. The convolution block is followed by a global average pooling layer, two feedforward layers and a sigmoid activation layer. The two-dimensional convolution layer has a kernel size of (4, 4), a stride of (2, 2), and padding of (1, 1). The sigmoid activation function is a nonlinear activation function that aims to map the discrimination result to between (0, 1).
[0082] like Figure 5 As shown, the perceptual evaluation score of speech quality PESQ is used to train the discriminator, and the amplitude of clean speech is used as reference and degraded input to estimate the maximum normalized PESQ score. The generator and discriminator are trained separately in each epoch, and the gradient is updated to generate enhanced speech similar to clean speech.
[0083] The loss function used by the generator is the amplitude loss function in the time-frequency domain And the complex loss function
[0084]
[0085] Among them, α represents the weight, X m represents the amplitude of clean speech, X r and X i Represents the real and imaginary parts of clean speech in the time-frequency domain.
[0086] The adversarial training loss of the entire network model is the discriminator loss and the associated generator loss
[0087]
[0088] Where Q PESQ is the normalized PESQ score.
[0089] Final generator loss for and time domain loss Combination of:
[0090]
[0091] in, represents the enhanced speech signal, X represents the clean speech signal, ‖·‖1 represents l 1 norm, ‖·‖2 represents l 2 norm, γ1, γ2, γ3 represent the corresponding loss weights.
[0092] Test Case
[0093] The simulation environment for the speech enhancement method experiment based on the selective state space model Mamba is: GPU NVIDIARTX4070TI SUPER, CPU I5 13600KF, Ubuntu 20.04LTS, CUDA 12.0, Pytorch 2.0.1.
[0094] This experiment uses the public dataset VCTK as a clean speech dataset. The dataset contains 108 44-hour recordings with various English accents and a sampling rate of 22.05KHz. First, the dataset speech is downsampled to 16KHz, then convolved with the RIR_NOISES dataset and 5-30dB random noise is added to obtain a noisy reverberation dataset. The sentences in the training set are cut into 2-second segments, and the sentences in the test set are not cut. The STFT step size is 400, the Hamming window is 25ms, and the frame shift is 6.25ms, that is, 75% overlap. The number of TF-Mamba blocks in the generator is set to 4, the batch size B is 2, and the number of channels C is 64. The number of channels in the metric discriminator is set to {16,32,64,128}. The number of training epochs is 100. During the training phase, both the generator and the discriminator are trained using the AdamW optimizer. The generator learning rate is set to 0.0004, the discriminator learning rate is set to 0.001, and the decay coefficient of the learning rate scheduler is 0.5 every 12 epochs. The generator loss weight is set to {γ1=1,γ2=0.01,γ3=1}.
[0095] Figure 6-8 They are the spectrograms of clean speech, noisy reverberation speech and enhanced speech in the test set of this test case, respectively. Figure 6 The clean speech spectrogram has a clear harmonic structure, obvious resonance peaks, and energy concentrated on the main frequency components of the speech. There is less background noise and the non-speech area is clean. Figure 7 Since noisy reverberation speech is obtained by superimposing clean speech with reverberation and background noise, its harmonic structure is fuzzy, the resonance peak is not obvious, the energy distribution is dispersed and extends to the non-main frequency area, the background noise increases, the energy distribution appears in the non-speech area, the time tailing phenomenon is obvious, and the speech component tails on the time axis. Figure 8 The enhanced speech spectrogram shows that the time tailing phenomenon has been significantly weakened, the background noise has also been reduced, and the overall spectrogram is closer to that of clean speech.
[0096] The control models of this experiment are MetricGAN+, CMGAN and TSTNN. MetricGAN is a speech enhancement network that connects the metric discriminator with the evaluation index for training, MetricGAN+ is a speech enhancement network that optimizes the sigmoid activation function based on MetricGAN, CMGAN is a speech enhancement network that combines MetricGAN with a conformer structure based on the self-attention mechanism, and TSTNN is a speech enhancement network that performs speech enhancement in the time domain through a transformer structure.
[0097] The experimental evaluation indicators are speech quality perception evaluation PESQ, speech intelligibility evaluation STOI, background noise intrusion CBAK, and speech signal distortion CSIG. PESQ simulates human auditory perception and calculates the difference between enhanced speech and original clean speech. The index range is [-0.5,4.5]. The higher the value, the better the speech quality; STOI is used to measure the intelligibility of speech. The index range is [0,1]. The higher the value, the better the speech intelligibility; CBAK is mainly used to measure the residual degree of background noise, that is, how much noise interference is in the enhanced speech. The range is [1,5]. The higher the score, the less background noise, that is, the better the enhancement effect; CSIG evaluates whether the enhanced speech signal maintains the original speech characteristics, that is, whether distortion is introduced during the enhancement process. The higher the score, the less distortion the enhanced speech has, and the better the speech quality. Table 1 below is the evaluation index table of speech enhancement in this test case:
[0098] Table 1
[0099] PESQ STOI CBAK CSIG Reverberation Speech 1.981 0.836 2.265 3.162 CMGAN 3.237 0.904 2.515 3.504 MetricGAN+ 3.216 0.892 2.477 3.36 TSTNN 3.158 0.887 2.51 3.375 TF-Mamba 3.34 0.925 2.615 3.471
[0100] As can be seen from Table 1, the PSEQ, STOI and CBAK indicators of the speech enhancement network proposed by this method are better than those of the comparison network model, and the CSIG indicator is better than MetricGAN+ and TSTNN, indicating that the speech enhancement effect of this method is better than the current mainstream speech enhancement network. Compared with CMGAN, when trained under the same batch size, the average training time per round of the model is reduced from 48 minutes to 31 minutes, and the training speed is increased by about 34.8%.
[0101] In summary, the TF-Mamba method proposed in the present invention can improve the model training speed. Compared with the transformer and conformer architectures, the training speed has been improved to a certain extent, and various evaluation indicators are better than the speech enhancement network of the transformer architecture. This method achieves more advanced SE performance with lower computational complexity.
[0102] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.
Claims
1. A speech enhancement method based on the selective state space model Mamba, characterized in that: At least the following steps are included: S1. Data preprocessing: distorted speech signal input Preprocessing is performed to obtain the time-frequency domain signal through short-time Fourier transform, where L represents the time domain length of the speech; the amplitude, real part and imaginary part of the time-frequency domain signal are concatenated as the input signal Among them, B represents the batch size, T is the number of frames, F is the frequency, and 3 is the number of channels; S2, signal expansion: The time-frequency domain signal obtained by the transformation in step S1 is dimensionally expanded through an encoder, from 3 channels to C channels; the encoder includes two convolution blocks and an expanded dense convolution layer, wherein the first convolution block is used to convert the input signal The channel dimension is expanded from 3 to C, forming a signal The expanded signal is obtained by expanding the dense convolution layer to extract features of different resolutions and increase the receptive field The second convolutional block is used to convert the signal Y″ after passing through the dilated dense convolutional layer stft The frequency F is downsampled to Get the extended signal S3, feature enhancement: Through N TF-mamba blocks, the signal expanded in step S2 is feature enhanced to simulate the forward and backward dependencies of speech signals at different resolutions, and the feature-enhanced signal is obtained. The TF-mamba block is composed of a Time Mamba block and a Frequency Mamba block. Each Mamba block adopts a bidirectional SSM mode, and the input signal is processed in parallel by the forward Mamba and the backward Mamba. S4, speech signal enhancement: the signal of step S3 is processed by a mask decoder and a complex decoder for amplitude and phase respectively to obtain predicted amplitude and phase, which are combined to obtain a predicted spectrum signal, and the predicted spectrum signal is subjected to inverse power compression and inverse short-time Fourier transform (ISTFT) to obtain an enhanced speech signal; S5, enhanced speech discrimination: The enhanced speech signal obtained in step S4 and the clean speech signal are discriminated by the metric discriminator, and the gradients of the generator and the discriminator are updated after the discriminator is trained; the input signal Y of step S2 is in Input the generator after updating the gradient, repeat steps S2-S4 to obtain new enhanced speech, perform enhanced speech discrimination through the discriminator after gradient updating, repeat steps S2-S5 to update the generator network and discriminator to achieve speech enhancement.
2. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 1, characterized in that: The method for obtaining the time-frequency domain signal in step S1 is specifically as follows: by STFT operation, the time-domain waveform of the input distorted speech signal is transformed into Convert to time-frequency domain signal Where L represents the length of speech time domain, T and F represent the number of frames and frequency points respectively, and the compressed spectrum Y is obtained by power compression: Where c represents the power compression index, Y m , Y p , Y r and Y i is the amplitude, phase, real component and imaginary component of the spectrum graph Y, ω(τ-t) is the window function, t is the time variable, f is the frequency variable, j is the imaginary unit of the complex number; Y r , Y i and Y m Connect as input signal B represents the batch size, T represents the number of frames, and F represents the frequency.
3. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 1, characterized in that: Each convolution block of the encoder in step S2 includes a two-dimensional convolution layer, an instance normalization layer and a PReLU activation function layer; the dilated dense convolution layer contains four convolution blocks with dense residual connections, and the dilation factors of the four convolution blocks are {1, 2, 4, 8}.
4. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 3, characterized in that: In the encoder of step S2, the convolution kernel size of the first convolution block is (1, 1), and the step size is (1, 1); the convolution kernel size of the second convolution block is (1, 3), the step size is (1, 2), and the padding is (0, 1).
5. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 1, characterized in that: The specific steps of the signal feature enhancement in step S3 are: S31: The input signal obtained in step S2 Reshape it into The local and global information is processed along the time domain dimension T by the Time Mamba block and then combined with Y tm Through skip connection AND t =TimeMamba(Y tm )+Y tm S32: The signal obtained from S31 Reshape it into The FrequencyMamba block processes local and global information along the frequency domain dimension F and then adds it to Y fm Through skip connection AND f =FrequencyMamba(Y fm )+Y fm S33: Reshape Y mamba Enter the next TF-Mamba block. A total of N TF-Mamba blocks need to be entered, that is, steps S31 and S32 need to be repeated N times to finally obtain the feature-enhanced signal.
6. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 5, characterized in that: The TF-Mamba block structure consists of two mamba blocks, Time Mamba and Frequency Mamba, which perform forward and backward scanning respectively and integrate local and global information. The operation steps of the signal in the mamba block are as follows: a. For input data First, the forward scan is performed through the forward mamba block, and then the RMSNorm activation function is used, and then the Y eg Get Y through skip connection f : AND f =RMS(FMamba(Y eg ))+Y eg b. For input data Flip it to get Y flip , and then scanned backwards through the backward mamba block, and then activated with the RMSNorm function, and then with Y flip Get Y through skip connection b : AND b =RMS(FMamba(Y flip ))+Y flip c. Y f and Y b After concatenation along the channel dimension C, transposed convolution is performed to obtain the output Y out : AND out =TransConv(Concat(And f ,AND b )) Among them, RMS, FMamba, BMamba, Flip, and Concat represent forward output, backward output, RMS normalization, forward Mamba, backward Mamba, flip, and concatenation operations, respectively.
7. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 1, characterized in that: The step S4 specifically includes the following steps: S41: In the mask decoder, a convolutional block is used The number of channels is compressed, and then the final amplitude is predicted by another convolutional layer with PReLU activation. S42: In the complex decoder, the real and imaginary parts of the final enhanced signal are predicted by the convolution block S43: The final amplitude predicted in step S41 Sum signal phase Y p The enhanced frequency domain signal is obtained by combining; S44: The enhanced frequency domain signal obtained in step S43 is compared with the predicted output of the complex decoder in step S42. Add them in sequence to get the final enhanced frequency domain signal: S45: In the complex spectrum Perform inverse power compression and inverse short-time Fourier transform ISTFT to obtain the time domain signal 8. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 7, characterized in that: The metric discriminator of step S5 is composed of four convolution blocks, each of which is composed of a two-dimensional convolution layer, an instance normalization layer, and a PReLU activation layer. The convolution block is followed by a global average pooling layer, two feed-forward layers, and a sigmoid activation layer.
9. The method for speech enhancement based on the selective state space model Mamba as claimed in claim 8, characterized in that: The loss function of the generator in step S5 is the amplitude loss function in the time-frequency domain And the complex loss function Among them, α represents the weight, X m represents the amplitude of clean speech, X r and X i Represents the real and imaginary parts of clean speech in the time-frequency domain; The generator loss is the associated generator loss and time domain loss Combination of: in, represents the enhanced speech signal, X represents the clean speech signal, ‖·‖1 represents l 1 norm, ‖·‖2 represents l 2 norm, γ1, γ2, γ3 represent the corresponding loss weights.
10. A speech enhancement system based on the selective state space model Mamba, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of any of the above methods are implemented.
Citation Information
Patent Citations
Noise-containing speech separation method based on selective state space model
CN118782065A
Scene-aware audiovisual speech enhancement method and device, medium and program product
CN118918913A
Monaural speech enhancement method and device based on double-branch network
CN119049489A
Channel information feedback method of large-scale MIMO system based on multi-scale feature fusion
CN119171953A
Denoising network model based on diffusion model and oracle rubbing image restoration method
CN119251088A
Cited By
Array speech enhancement method and system based on Hemma optimization and Mama
CN121506167A
MP-SENet speech enhancement method based on frequency domain channel attention and feature compression
CN122157686A
MP-SENet speech enhancement method based on frequency domain channel attention and feature compression
CN122157686B