Speech enhancement method for low-signal-to-noise-ratio audio signal

By using a complex encoder network and Conformer module to model the audio signal, combined with acoustic parameters and fine-grained spectrum feature extractor, multi-stage training and contrast learning methods are used to solve the noise residue and speech distortion of low signal-to-noise ratio audio signal in the existing technology, achieving better speech enhancement effect.

CN120164480AActive Publication Date: 2025-06-17DALIAN MARITIME UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510284302.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When existing speech enhancement methods face low signal-to-noise ratio audio signals, there are problems of noise residues and speech distortion.

Method used

The complex encoder network and Conformer module are used to model the real and imaginary parts of the audio signal, combined with the acoustic parameter extractor and the fine-grained spectrum feature extractor, and through multi-stage training and contrast learning methods, the loss function is constructed to optimize the network parameters.

Benefits of technology

Effectively reduce noise residue, improve the signal-to-noise ratio of speech signals, enhance the clarity and comprehensibility of speech, and significantly improve the voice enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164480A_ABST
    Figure CN120164480A_ABST
Patent Text Reader

Abstract

The invention provides a speech enhancement method for an audio signal with a low signal-to-noise ratio, and the method comprises the following steps: S1, inputting the frequency spectrum of the audio signal into a codec network, and obtaining a first-stage enhanced speech; s4, acoustic parameter loss is calculated; s5, calculating fine-grained spectrum feature loss; s6, combining the two losses and the frequency domain loss to construct a first-stage loss for training to obtain a first-stage intensifier; in the second stage of training, S7, features of the output in the S1, the pure voice and the residual noise are obtained through an encoder, and learning loss is calculated and compared; s9, constructing a second-stage loss number based on the first-stage loss and the comparative learning loss, and carrying out second-stage training to obtain a second-stage intensifier; and a reasoning stage: S10, inputting the noise-containing voice into the two-stage intensifier which is connected in sequence to obtain a target voice. According to the method, the influence of acoustic parameter features and fine-grained spectrum features on the enhanced voice signals is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maritime communication, and more particularly, to a method for enhancing the speech of audio signals with low signal-to-noise ratio. Background Art

[0002] Maritime communication audio signals are one of the most critical data recorded in ship voyage recorders and are also the main communication method during ship operations, being widely used for voice and data transmission between ships and shores. However, during the process of maritime ship navigation communication, maritime communication audio signals are often affected by the maritime environment and the complexity of ship navigation, resulting in various noises being mixed in the speech. The wind and wave sounds in the marine environment, the mechanical noises of ship engines, and the electromagnetic interference of other electronic devices will all interfere with VHF voice signals. Under the combined action of these factors, the signal-to-noise ratio of maritime communication audio signals is usually low, affecting the clarity and intelligibility of the speech. Therefore, enhancing maritime communication audio signals is directly related to the efficiency of ship-shore collaboration and maritime supervision, as well as the timely transmission of search and rescue information.

[0003] Speech enhancement has been developed for a quite long time. Traditional speech enhancement methods, such as minimum mean square error filtering and spectral subtraction filtering, improved signal subspace combined with Wiener filtering, and Kalman filtering algorithms, although all have certain effects, their performance is not satisfactory when facing invisible noise or low signal-to-noise ratio situations. The fundamental reason is that traditional methods rely too much on prior assumptions. If the noise assumption is underestimated, it will lead to noise residue, and spikes will appear in the spectrum, forming artificial noise. And overestimating the noise will damage the original speech information and cause speech distortion. Data-driven deep learning methods include time-domain waveform mapping, spectral mapping, and mask generation methods, etc. The time-domain waveform mapping method refers to the method of inputting the audio waveform into the model and the model directly outputting the pure speech waveform; the spectral mapping method transforms the time-domain input into the frequency domain and processes the enhancement task using the spectral mapping method; the output of the mask generation method is the mask value, and the enhanced speech is obtained based on the mask and the spectrum. Deep learning methods do not require prior assumptions and have high robustness to complex noises, and have gradually become the mainstream methods. However, existing deep learning methods have deficiencies in modeling the acoustic features and fine-grained spectra of speech, and there are problems of over-suppression and noise residue due to poor decoupling ability. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to propose a method for enhancing the speech of audio signals with low signal-to-noise ratio to solve the technical problem of noise residue existing in existing speech enhancement methods.

[0005] The technical means adopted by the present invention are as follows:

[0006] A method for speech enhancement of low signal-to-noise ratio audio signals, comprising the following steps:

[0007] First-stage training:

[0008] S1. Input the short-time Fourier transform spectrum of the audio signal into the complex encoder network, extract the real and imaginary part features, and obtain the output of the encoder network, i.e., the intermediate features;

[0009] S2. Input the intermediate features into the Conformer module, and the Conformer module models the temporal information using the attention network to obtain the output of the Conformer module;

[0010] S3. Concatenate the intermediate features and the output of the Conformer module and input them into the complex decoder network to obtain the frequency-domain output, and perform inverse Fourier transform on the frequency-domain output to obtain the first-stage time-domain enhanced speech;

[0011] S4. Input the clean speech and the first-stage time-domain enhanced speech in the audio signal into the acoustic parameter extractor to obtain the embeddings of the clean speech and the first-stage time-domain enhanced speech respectively, calculate the root mean square error of the embeddings of the clean speech and the first-stage time-domain enhanced speech, and construct an acoustic parameter loss function;

[0012] S5. Input the clean speech and the first-stage time-domain enhanced speech in the audio signal into the fine-grained spectral feature extractor, extract the output of the penultimate layer of the fine-grained spectral feature extractor network as the fine-grained spectral feature, and align the fine-grained spectral features using the root mean square error loss to construct a fine-grained feature loss function;

[0013] S6. Based on the acoustic parameter loss function, the fine-grained feature loss function and the spectral loss, construct the loss function of the first stage, use the loss function of the first stage for the first-stage training, and output the trained first-stage enhancer;

[0014] The training process is as follows: calculate the gradient of the first-stage loss function with respect to the network parameters, subtract the product of the gradient and the learning rate from the network parameters, repeat updating the network parameters to reduce the loss until the updated network parameters do not improve in performance for 20 consecutive iterations on the validation set, and then terminate the training;

[0015] Second-stage training:

[0016] S7. Respectively obtain features of the first-stage time-domain enhanced speech, clean speech, and residual noise through the encoder;

[0017] S8. Construct triplets based on the features obtained in S7, input the triplets into the contrastive learning method to obtain a contrastive learning loss function;

[0018] S9. Based on the fine-grained feature loss function, contrastive learning loss function, and L log-PCM Construct the loss function for the second stage, use the loss function of the second stage for the second stage training, and output the trained second stage enhancer;

[0019] Speech enhancement:

[0020] S10. Input the audio to be enhanced into the trained first stage enhancer, output the audio enhanced in the first stage, input the audio enhanced in the first stage into the trained second stage enhancer, and output the enhanced target audio.

[0021] Furthermore, S1 specifically includes the following steps:

[0022] Perform short-time Fourier transform on the audio signal to obtain the real part component and the imaginary part component;

[0023] Use the feature consistency module to extract the acoustic parameter features and fine-grained spectral features in the enhanced speech; input the obtained real part component and imaginary part component into the complex encoder network; the complex encoder network includes a complex convolutional layer, complex batch normalization, and PReLU activation function; process the real part component and the imaginary part component simultaneously through the complex convolutional layer that can perform convolution on both the real part component and the imaginary part component, and its specific representation of the calculation process is as follows:

[0024]

[0025] Among them, A k is the input of the complex two-dimensional convolution, W r and W i are respectively the weights of the real part and the imaginary part of the complex convolution kernel, and are respectively the real part and the imaginary part of the input signal;

[0026] A k After the batch normalization processing and PReLU activation function of the complex encoding layer, the corresponding encoder network output F k is obtained, and the formula is as follows:

[0027] F k = PReLU(Batchnorm2d(A k ))

[0028] Furthermore, in S2, the formula for the Conformer module to model the temporal information using the attention network is as follows:

[0029]

[0030] Among them, H is the output of the Conformer module, is the real part, is the imaginary part.

[0031] Furthermore, in S3, the frequency-domain output is expressed as:

[0032] O = Decoder(concatenate(F k , H)).

[0033] Furthermore, the acoustic parameter loss function in S4 and the fine-grained feature loss function in S5 are as follows:

[0034]

[0035] where T is the acoustic parameter embedding corresponding to the audio signal in the training set, is the acoustic parameter embedding after passing through the acoustic parameter extractor, F is the fine-grained spectral embedding corresponding to the audio signal in the training set, is the fine-grained spectral embedding after passing through the post-acoustic parameter extractor.

[0036] Furthermore, the formula for obtaining features through the encoder in S7 is as follows:

[0037]

[0038] where S is the output of the first stage, X is the clean signal, w, w + and w - respectively represent the first-stage output features, clean audio features, and residual noise features extracted after passing through the encoder.

[0039] Furthermore, the loss function of the first stage is:

[0040] L PFL-stage1 = L log-PCM + αL TAP + βL FM

[0041] The loss function of the second stage is:

[0042] L PFL-stage2 = L log-PCM + αL TAP + βL FM + λL CL

[0043] where α, β, and λ are hyperparameters;

[0044] L log-PCM is:

[0045]

[0046] where, S(t,f) and are the coefficients of the STFT of s and respectively, T is the number of time frames, F is the number of frequency points; r and i represent the real and imaginary parts of a complex number respectively; L SM is the mean absolute error loss between the pure short-time Fourier transform coefficients and the estimated short-time Fourier coefficients;

[0047] The contrast loss function is as follows:

[0048]

[0049] where, N is the number of samples, τ is the temperature function, and the triplet is [w, w + , w - , where [w, w + is used as the positive sample pair, and [w, w - is used as the negative sample pair.

[0050] The present invention also provides a storage medium, the storage medium includes a stored program, wherein when the program runs, it executes any one of the above-mentioned speech enhancement methods for low signal-to-noise ratio audio signals.

[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes any one of the above-mentioned speech enhancement methods for low signal-to-noise ratio audio signals through the running of the computer program.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] The present invention takes into account the influence of the acoustic parameter characteristics and fine-grained spectral characteristics of the signal on the construction of the enhanced speech signal. The real and imaginary parts of the speech signal are simultaneously modeled through a complex codec network, the temporal characteristics of the speech signal are modeled using a Conformer module, and the speech is reconstructed through a complex decoder network.

[0054] The proposed feature consistency module in this method can model 25 acoustic features such as fundamental frequency (F0), jitter, center frequencies and bandwidths of the first three formants through the acoustic parameter modeling part, and can capture the minute spectral details in the audio signal through the fine-grained spectral feature module, so as to achieve a better speech enhancement effect.

[0055] The contrastive learning module proposed by this method improves the decoupling ability of the model, thus realizing the removal of residual noise and the correction of over-suppression. In addition, the loss function proposed by this method combines the frequency-domain loss constrained by phase, the acoustic parameter loss, the fine-grained spectral feature loss, and the contrastive learning loss to guide the network to update parameters, obtaining better network parameters and enhancement effects. Experimental results show that the enhancement effect of the present invention on speech signals is significantly higher than other methods, fully demonstrating the effectiveness of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a flowchart of the method of the present invention.

[0058] Figure 2 It is a time-domain waveform diagram of the audio before enhancement of the present invention.

[0059] Figure 3 It is a spectrogram of the audio before enhancement of the present invention.

[0060] Figure 4 It is a time-domain waveform diagram of the audio after enhancement of the present invention.

[0061] Figure 5 It is a spectrogram of the audio after enhancement of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] As Figure 1 shown, the present invention provides a speech enhancement method for low signal-to-noise ratio audio signals, including the following steps:

[0065] First, perform a short-time Fourier transform on the audio signal to obtain a real component and an imaginary component.

[0066] The present invention adopts segmented training. In the first stage, a feature consistency module is used to align the acoustic parameter features and fine-grained spectral features in the enhanced speech. The obtained real component and imaginary component are input into a complex encoder network so that the model can better reconstruct the real and imaginary parts of the speech. The complex encoder network consists of a complex convolutional layer, complex batch normalization, and a PReLU activation function. The real and imaginary components of the speech are processed simultaneously by the complex convolutional layer that can perform convolution on the real and imaginary parts at the same time, more effectively capturing and processing the complex features of the speech signal. If A k is the output of the complex two-dimensional convolution, its specific representation of the calculation process is as follows:

[0067]

[0068] where, W r and W i respectively represent the weights of the real part and the imaginary part of the complex convolution kernel, and are respectively the real part and the imaginary part of the input signal. Then, A k undergoes batch normalization processing and PReLU activation function of the complex encoding layer to obtain the corresponding encoder network output F k . The formula is as follows:

[0069] F k = PReLU(Batchnorm2d(A k )) (2)

[0070] The intermediate features extracted by the complex encoder are fed into the Conformer module, which uses its attention mechanism to model the temporal information. Its output can be expressed as:

[0071]

[0072] Then, the output F k of the encoder is concatenated with the output H of the Conformer and input into the decoder. Then, the frequency-domain output O of the decoder can be expressed as:

[0073] O = Decoder(concatenate(F k , H)) (4)

[0074] The inverse Fourier transform is performed on the frequency-domain output O to obtain the first-stage time-domain enhanced speech S.

[0075] Then, the clean speech and the enhanced speech are input into the acoustic parameter extractor to obtain the embeddings of the clean speech and the enhanced speech respectively. The root mean square error is calculated to reduce the distance between these embeddings to ensure the consistency of the acoustic parameters between the enhanced speech and the clean speech. At the same time, the clean speech and the enhanced speech are input into the fine-grained spectral feature extractor, and the output of the penultimate layer of its network is extracted as the fine-grained spectral feature. The root mean square error loss is used to align the fine-grained spectral features, and the acoustic parameter loss function and the fine-grained spectral feature loss function can be expressed as:

[0076]

[0077] where T is the acoustic parameter embedding corresponding to the audio signal in the training set, is the acoustic parameter embedding obtained through the acoustic parameter extractor, F is the fine-grained spectral embedding corresponding to the audio signal in the training set, is the fine-grained spectral embedding obtained through the fine-grained spectral feature extractor.

[0078] The fine-grained spectral feature extractor network consists of 16 convolutional layers and 3 fully connected layers connected in sequence.

[0079] In the second stage, contrastive learning is introduced. The output of the first stage, the clean audio, and the residual noise also pass through the Encoder with shared weights to obtain features, which can be expressed as:

[0080]

[0081] where R is the residual noise, S is the output of the first stage, X is the clean signal, w, w + and w - represent the first-stage output features, clean audio features, and residual noise features extracted after passing through the encoder respectively.

[0082] For each sample, a triple of [w, w + , w - is constructed, where [w, w + is used as a positive sample pair, and [w, w - is used as a negative sample pair. In contrastive learning, a contrastive loss function is usually used to maximize the consistency between positive samples and minimize the consistency between negative samples in the learned representation space, thereby enhancing the decoupling ability of the model, as shown in the following formula:

[0083]

[0084] When training the model, the loss function used is the one proposed in the present invention, L PFL , which uses L that can limit the phase on the main body log-PCM , and at the same time adds a feature consistency module, which can model the acoustic parameters and fine-grained spectral features of the audio. Finally, a contrastive learning loss function is added to improve the decoupling ability of the model.

[0085] Therefore, the loss function of the first stage is defined as:

[0086] L PFL-stage1 = L log-PCM + αL TAP + βL FM (8)

[0087] The loss function of the second stage is defined as:

[0088] L PFL-stage2 = L log-PCM + αL TAP + βL FM + λL CL (9)

[0089] Among them, α, β, and λ are hyperparameters, which are set to 0.1, 0.6, and 1.8 respectively here.

[0090] L log-PCM can be defined as:

[0091]

[0092] Among them, S(t, f) and are the coefficients of the STFT of s and respectively, T is the number of time frames, and F is the frequency range. The subscripts r and i represent the real and imaginary parts of the complex number respectively. L SM is the mean absolute error loss between the pure short-time Fourier transform coefficients and the estimated short-time Fourier coefficients.

[0093] The network structure adopted in the two-stage training is a complex encoder-Conformer-complex decoder structure. Both the encoder and the decoder contain six complex convolutional layers, and the corresponding number of channels for each layer is set to (32, 64, 128, 256, 256, 256). There are two layers of Conformer modules between the complex encoder and the complex decoder, and an eight-head attention network is adopted. The training processes of the first stage and the second stage are the same, and the training process is as follows: calculate the gradient of the loss function with respect to the network parameters, subtract the product of the gradient and the learning rate from the network parameters, and repeat updating the network parameters to reduce the loss until the updated network parameters do not improve in performance for 20 consecutive iterations on the validation set, then terminate the training;

[0094] All audio data is 16 kHz. The window function uses a Hamming window, and the number of points for the Fourier transform is 512. The learning rate during the experiment is set to 0.001, and the Adam optimizer is used. The epoch is set to 30, and the batch size used is 1. The server version used for training the model is Ubuntu 18.04.6, the method is implemented using Python 3.8, and the PyTorch framework is used. If the weights of the training result do not improve in performance for 20 consecutive iterations on the validation set, then terminate the training.

[0095] Figure 2 This is the time-domain waveform diagram of the audio before enhancement of the present invention. It can be observed that there is a lot of noise in the speech before enhancement. This noise consists of wind and wave sounds, electromagnetic interference noise, and mechanical noise, resulting in a very low signal-to-noise ratio of the speech. Figure 3 This is the spectrogram of the audio before enhancement of the present invention. In the spectrogram of the audio before enhancement, the speech spectrum is covered by the noise spectrum, and the harmonic structure cannot be captured. Figure 4 This is the time-domain waveform diagram of the audio after enhancement of the present invention. The noise in the audio after enhancement is suppressed, the speech content of the speaker is retained in the speech waveform diagram, and the signal-to-noise ratio is improved. Figure 5 This is the spectrogram of the audio after enhancement of the present invention. The harmonic structure in the spectrogram of the speech after enhancement is clear.

[0096] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for speech enhancement of low signal-to-noise ratio audio signals, characterized in that: The steps include: First stage training: S1. Input the short-time Fourier transform spectrum of the audio signal into the complex encoder network, extract the real and imaginary part features, and obtain the output of the encoder network, i.e., the intermediate features; S2, input the intermediate features into the Conformer module, the Conformer module uses the attention network to model the temporal information, and obtain the output of the Conformer module; S3, concatenating the intermediate features and the output of the Conformer module and inputting them into the complex decoder network to obtain the frequency domain output, and performing an inverse Fourier transform on the frequency domain output to obtain the first stage time domain enhanced speech; S4, inputting the clean speech and the first-stage time-domain enhanced speech in the audio signal into the acoustic parameter extractor to obtain the embedding of the clean speech and the first-stage time-domain enhanced speech respectively, calculating the root mean square error of the embedding of the clean speech and the first-stage time-domain enhanced speech, and calculating the acoustic parameter loss function; S5, inputting the clean speech and the first-stage time-domain enhanced speech in the audio signal into the fine-grained spectral feature extractor, extracting the penultimate layer output of the fine-grained spectral feature extractor network as the fine-grained spectral feature, aligning the fine-grained spectral features using the root mean square error loss, and calculating the fine-grained feature loss function; S6. Construct a first-stage loss function based on the acoustic parameter loss function, the fine-grained feature loss function and the spectrum loss, use the first-stage loss function to perform the first-stage training, and output the trained first-stage enhancer; The training process is as follows: calculate the gradient of the first-stage loss function to the network parameters, subtract the product of the gradient and the learning rate from the network parameters, and repeatedly update the network parameters to reduce the loss until the updated network parameters do not improve for 20 consecutive iterations on the validation set, then terminate the training; Second stage training: S7, obtaining features of the first-stage time-domain enhanced speech, clean speech and residual noise through encoders respectively; S8, constructing a triplet based on the features obtained in S7, inputting the triplet into a contrastive learning method, and obtaining a contrastive learning loss function; S9, based on acoustic parameter loss function, fine-grained feature loss function, contrastive learning loss function and L log-PCM Construct the loss function of the second stage, use the loss function of the second stage to perform the second stage training, and output the trained second stage enhancer; Voice Enhancement: S10, input the audio to be enhanced into the trained first-stage enhancer, output the audio enhanced by the first stage, input the audio enhanced by the first stage into the trained second-stage enhancer, and output the enhanced target audio.

2. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 1, characterized in that: S1 specifically includes the following steps: Perform short-time Fourier transform on the audio signal to obtain real and imaginary components; The feature consistency module is used to extract acoustic parameter features and fine-grained spectral features in enhanced speech; The obtained real and imaginary components are input into the complex encoder network; the complex encoder network includes a complex convolution layer, a complex batch normalization, and a PReLU activation function; the real and imaginary components are processed simultaneously by a complex convolution layer that can simultaneously convolve the real and imaginary components, and the specific calculation process is as follows: Among them, A k is the output of the complex 2D convolution, W r and W i are the weights of the real part and the imaginary part of the complex convolution kernel, respectively. and are the real and imaginary parts of the input signal respectively; A k After batch normalization and PReLU activation function of the complex encoding layer, the corresponding encoder network output F is obtained. k , the formula is as follows: F k =PReLU(Batchnorm2d(A k ))。 3. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 1, characterized in that: In S2, the formula for the Conformer module to model temporal information using the attention network is as follows: Among them, H is the output of the Conformer module, is the real part, is the imaginary part.

4. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 1, characterized in that: In S3, the frequency domain output is expressed as: O=Decoder(concatenate(F k ,H))。 5. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 1, characterized in that: The acoustic parameter loss function in S4 and the fine-grained feature loss function in S5 are as follows: Where T is the acoustic parameter embedding corresponding to the audio signal in the training set, is the acoustic parameter embedding after passing through the post-acoustic parameter extractor, F is the fine-grained spectral embedding corresponding to the audio signal in the training set, It is a fine-grained spectral embedding after passing through the post-acoustic parameter extractor.

6. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 1, characterized in that: The formula for obtaining features through the encoder in S7 is as follows: Among them, S is the output of the first stage, X is the pure signal, w, w + and w - They represent the first-stage output features, clean audio features, and residual noise features extracted after the encoder, respectively.

7. The method for speech enhancement of low signal-to-noise ratio audio signals according to claim 5, characterized in that: The loss function of the first stage is: L PFL-stage1 =L log-PCM +αL TAP +βL FM The loss function of the second stage is: L PFL-stage2 =L log-PCM +αL TAP +βL FM +λL CL Among them, α, β, and λ are hyperparameters; L log-PCM for: Among them, S(t,f) and They are s and The coefficients of the STFT of , T is the number of time frames, F is the number of frequency points; r and i represent the real and imaginary parts of the complex number respectively; L SM is the mean absolute error loss between the pure StF coefficients and the estimated StF coefficients; The contrastive learning loss function is as follows: Where N is the number of samples, τ is the temperature function, and the triple is [w,w + ,w - ], where [w,w + ] as a positive sample pair, [w,w - ] as negative sample pairs.

8. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the speech enhancement method for a low signal-to-noise ratio audio signal according to any one of claims 1 to 7 is executed.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the speech enhancement method for a low signal-to-noise ratio audio signal according to any one of claims 1 to 7 by running the computer program.

Citation Information

Patent Citations

  • Speech enhancement method for ship VHF communication audio signal

    CN117409793A

  • Scene-aware audiovisual speech enhancement method and device, medium and program product

    CN118918913A

  • Speech recognition model training method and apparatus, storage medium, and electronic device

    WO2024011902A1