Noise reduction method and device, mobile terminal and medium
By using pre-stored noise reduction model and estimator gain function on mobile terminals, the existing speech enhancement algorithms have solved the problem of noise residue and high computational complexity in complex noise environments, and efficient suppression of steady-state noise and clear restoration of target sounds are achieved.
Patent Information
- Application Number
- CN202410032286.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
Existing speech enhancement algorithms are difficult to accurately distinguish between steady-state noise and target sound in complex noise environments, resulting in noise residues, speech distortion and high computational complexity, and cannot be widely used in mobile terminals.
Using the pre-stored noise reduction model and the first estimator gain function, the noise intensity of the steady-state noise signal is estimated by performing frame-based windowing and short-time Fourier transform on the noise signal, and noise suppression is performed using the optimal logarithmic amplitude spectrum estimator to retain the target sound signal.
It improves the speed and accuracy of noise estimation, reduces the calculation amount, and realizes efficient steady-state noise suppression on mobile terminals, avoids noise residues and voice distortion, and improves user experience.
Smart Images

Figure CN120299469A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of mobile terminals, and in particular, to a noise reduction method, apparatus, mobile terminal, and medium. Background Art
[0002] Voice is the most direct and fast way in the process of information transmission. People can use mobile communication technology to achieve long-distance real-time voice communication and improve the efficiency of long-distance communication. In an ideal environment, that is, there is only the voice of the target speaker and no noise interference, voice communication can achieve very good results. However, in daily life, real-time voice communication is affected by noise interference, which generates information masking for the target voice signal, reduces the call quality, causes auditory fatigue, and makes communication between both sides difficult. And by increasing the playback volume to improve the recognition of the target voice signal in the noisy voice signal, it is easy to cause hearing damage. Therefore, it is necessary to perform noise reduction processing on the noisy voice signal to improve the voice quality. Summary of the Invention
[0003] To overcome the problems in the related art, the present disclosure provides a noise reduction method, apparatus, mobile terminal, and medium.
[0004] According to the first aspect of the embodiments of the present disclosure, a noise reduction method is provided, which is applied to a mobile terminal. The noise reduction method includes:
[0005] Obtain a noisy voice signal, where the noisy voice signal includes a steady-state noise signal and a target voice signal;
[0006] Based on the noisy voice signal and a pre-stored noise reduction model, obtain the noise intensity of the steady-state noise signal;
[0007] Based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function, remove the steady-state noise signal from the noisy voice signal to obtain the noise intensity of the target voice signal;
[0008] Based on the noise intensity of the target voice signal, obtain the target voice signal.
[0009] In some exemplary embodiments of the present disclosure, the obtaining the noise intensity of the steady-state noise signal based on the noisy voice signal and a pre-stored noise reduction model includes:
[0010] Based on the noisy voice signal, obtain the noise intensity of the noisy voice signal;
[0011] Input the noise intensity of the noisy voice signal into a pre-stored noise reduction model to obtain the gain function of the steady-state noise signal;
[0012] Obtain the noise intensity of the steady-state noise signal based on the gain function of the steady-state noise signal and the frequency-domain characteristics of the noisy voice signal.
[0013] In some exemplary embodiments of the present disclosure, obtaining the noise intensity of the noisy voice signal based on the noisy voice signal includes:
[0014] Perform frame windowing processing on the noisy voice signal and perform short-time Fourier transform to obtain the frequency-domain characteristics of each frame of the noisy voice signal;
[0015] Based on the frequency-domain characteristics of each frame of the noisy voice signal, obtain the noise intensity of each frame of the noisy voice signal.
[0016] In some exemplary embodiments of the present disclosure, obtaining the noise intensity of the target voice signal based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function includes:
[0017] Based on the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal in each frame of the noisy voice signal, determine the variable parameters in the first estimator gain function;
[0018] Based on the determined variable parameters and the first estimator gain function, obtain a second estimator gain function;
[0019] Based on the second estimator gain function and the frequency-domain characteristics of the noisy voice signal, obtain the noise intensity of the target voice signal.
[0020] In some exemplary embodiments of the present disclosure, determining the variable parameters in the first estimator gain function based on the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal in each frame of the noisy voice signal includes:
[0021] Based on the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal in each frame of the noisy voice signal, obtain the posterior signal-to-noise ratio of each frame of the noisy voice signal;
[0022] Based on the posterior signal-to-noise ratio of each frame of the noisy voice signal, determine the prior signal-to-noise ratio of this frame of the noisy voice signal;
[0023] Based on the posterior signal-to-noise ratio and the prior signal-to-noise ratio of each frame of the noisy voice signal, determine the variable parameters of the first estimator gain function.
[0024] In some exemplary embodiments of the present disclosure, determining the prior signal-to-noise ratio of a frame of the noisy voice signal based on the posterior signal-to-noise ratio of each frame of the noisy voice signal includes:
[0025] The a priori signal-to-noise ratio of each frame of the noisy voice signal is determined based on the a posteriori signal-to-noise ratio of this frame and the a priori signal-to-noise ratio of the previous frame of the noisy voice signal.
[0026] In some exemplary embodiments of the present disclosure, the method for obtaining the noise reduction model includes:
[0027] Obtain a training noisy voice signal, where the training noisy voice signal includes a target signal and a non-target signal, the target signal is steady-state noise, and the non-target signal includes one or more of non-steady-state noise, music, speech, and animal calls;
[0028] Train a neural network model based on the training noisy voice signal to obtain the noise reduction model.
[0029] In some exemplary embodiments of the present disclosure, the method for obtaining the noise reduction model further includes:
[0030] Determine a reference gain function based on the frequency domain characteristics of the training noisy voice signal and the frequency domain characteristics of the target signal;
[0031] Modify the noise reduction model based on the reference gain function.
[0032] According to a second aspect of the embodiments of the present disclosure, there is provided a noise reduction device applied to a mobile terminal. The noise reduction device includes:
[0033] An acquisition module for acquiring a noisy voice signal, where the noisy voice signal includes a steady-state noise signal and a target voice signal;
[0034] A processing module for obtaining the noise intensity of the steady-state noise signal based on the noisy voice signal and a pre-stored noise reduction model;
[0035] The processing module is further configured to remove the steady-state noise signal from the noisy voice signal based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function to obtain the noise intensity of the target voice signal;
[0036] The processing module is further configured to obtain the target voice signal based on the noise intensity of the target voice signal.
[0037] According to a third aspect of the embodiments of the present disclosure, there is provided a mobile terminal, where the mobile terminal includes:
[0038] A processor;
[0039] A memory for storing instructions executable by the processor;
[0040] Wherein, the processor is configured to execute executable instructions in the memory to implement the noise reduction method provided in the first aspect of the present disclosure.
[0041] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program capable of being run is stored. When the computer program is executed by a processor of a mobile terminal, the noise reduction method provided in the first aspect of the present disclosure is implemented.
[0042] Adopting the above method of the present disclosure has the following beneficial effects: The present disclosure uses a pre-stored noise reduction model to estimate the noise of a steady-state noise signal, which can effectively improve the speed and accuracy of noise estimation, obtain a better noise reduction effect, avoid problems such as noise residue, voice distortion, and sudden changes in volume due to inaccurate noise estimation, and the learning difficulty of the noise reduction model is low, reducing the computational amount in the noise reduction process and improving the applicability of the noise reduction method.
[0043] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0044] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0045] Figure 1 is a flowchart of a noise reduction method shown according to an exemplary embodiment.
[0046] Figure 2 is a flowchart of a noise reduction method shown according to an exemplary embodiment.
[0047] Figure 3 is a flowchart of a noise reduction method shown according to an exemplary embodiment.
[0048] Figure 4 is a block diagram of a noise reduction method shown according to an exemplary embodiment.
[0049] Figure 5 is a block diagram of a noise reduction device shown according to an exemplary embodiment.
[0050] Figure 6 is a block diagram of a mobile terminal shown according to an exemplary embodiment. Detailed Embodiments
[0051] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0052] Speech signal processing has always been a research hotspot in academia and industry, and speech enhancement technology for improving speech quality has been developing rapidly and steadily. Speech enhancement technology aims to extract the target speech from noisy speech and improve the perceptual quality and intelligibility of the enhanced speech. Speech enhancement technology is widely applied to the front-end modules of various speech systems to enhance the noise robustness of the speech system.
[0053] There are two technical solutions for existing speech enhancement technologies. One is the Optimally Modified Log Spectral Amplitude (OMLSA) speech enhancement algorithm based on Minima Controlled Recursive Averaging (MCRA), and the other is the speech enhancement algorithm based on deep neural networks.
[0054] Among them, the Optimally Modified Log Spectral Amplitude based on minimum control recursive averaging is the most mature traditional speech enhancement algorithm at present and has been widely applied in the industry. The gain formula of the Optimally Modified Log Spectral Amplitude in this algorithm is:
[0055] G(t,f) = G H (t,f) p(t,f) ×G min (1-p(t,f))
[0056] where t represents the frame index, f represents the frequency index, p(t,f) represents the presence probability of the target speech, and G min is the gain threshold value, which is determined according to the subjective standard of noise naturalness. G H (t,f) is expressed as:
[0057]
[0058] where ξ f and γ f represent the a priori signal-to-noise ratio and the a posteriori signal-to-noise ratio respectively. The a priori signal-to-noise ratio is solved using the decision-directed method, and the a posteriori signal-to-noise ratio is obtained from the power spectrum of the noisy sound signal and the power spectrum of the noise.
[0059] In this method, the presence probability p(t,f) of the target speech and the power spectrum of the noise are variables to be solved, and the minimum control recursive averaging (MCRA) needs to be used for estimation. Among them, the minimum control is used to calculate the presence probability of the target speech, and the recursive averaging is used to calculate the noise estimation, that is, the noise estimation is made based on the presence probability of the target speech. In the implementation process, MCRA needs to determine the length of the search window according to the actual situation, and usually 0.5s - 1.5s is selected. This parameter determines the tracking speed of the algorithm for the noise.
[0060] Since in the actual life scenario, the noise environment is very complex, the background noise in the speech will change, and the algorithm adopted in the above noise reduction method takes time to track the noise. During the process of tracking the noise, due to the incomplete convergence of the algorithm, noise residue will occur, resulting in an obvious noise signal. Although accelerating the convergence time of the algorithm can alleviate the problem of noise residue in the scenario with background noise switching. However, if the convergence time of the algorithm is too short, some audio that is expected to be retained will be misjudged as steady-state noise and removed, resulting in serious speech distortion problems.
[0061] In addition, when the microphone pickup scenario is a music scenario, since the music contains a very rich variety of sound types and there are many types of quasi-steady sounds, the algorithm cannot distinguish between steady-state noise and music, and often treats the music as steady-state noise, suppressing some music signals, resulting in music distortion and the problem of the sound being sometimes loud and sometimes soft.
[0062] The above problems can only be alleviated by reducing the degree of noise reduction, which makes the traditional speech enhancement algorithm unable to accurately estimate the noise, resulting in limited speech enhancement performance. Therefore, in applications with strict requirements for speech distortion, this algorithm cannot obtain a satisfactory enhancement effect.
[0063] Another commonly used noise reduction method is the speech enhancement algorithm based on deep neural networks. This algorithm relies on a large amount of training data and does not make any assumptions about the statistical characteristics of speech signals. Therefore, in a complex noise environment, this algorithm exhibits excellent enhancement performance. The speech enhancement algorithm based on deep neural networks mainly consists of two stages: training and noise reduction. In the training stage, noisy speech is synthesized using the target speech and noise in the ideal state prepared in advance. The features of the noisy speech are used as the input of the deep neural network. The training objective is calculated using the noisy speech and noise, and the network parameters are optimized through the BackPropagation (BP) algorithm. After being trained with a large amount of training data, the deep neural network has the ability to estimate the target speech from the features of the noisy speech. In the noise reduction stage, the features of the noisy speech are used as the input of the deep neural network. The deep neural network estimates the features of the target speech, and the phase of the noisy speech is added to restore the time-domain waveform of the target speech. The disadvantage of this algorithm is that the computational complexity of the deep neural network model is too large, requiring a large amount of computing power support and generating a large amount of power consumption. This makes the use of this algorithm subject to many restrictions and difficult to apply to mobile terminals such as mobile phones, and it cannot be widely promoted.
[0064] In practical applications, the noise environment is complex, and the boundary between noise and the target sound is very blurred. Therefore, in the process of reducing the noise of the sound signal, it is not necessary to suppress all noises, but only to suppress the steady-state noise. Due to factors such as the real-time nature of noise reduction and the computational complexity of noise reduction, traditional speech enhancement algorithms and speech enhancement algorithms based on deep neural networks cannot solve the problem of steady-state noise in all scenarios.
[0065] To solve the above problems, the present disclosure provides a noise reduction method, which is applied to a mobile terminal. The steady-state noise signal in the noisy sound signal is estimated using a pre-stored noise reduction model to obtain the noise intensity of the steady-state noise signal; the steady-state noise signal is suppressed using a pre-stored first estimator gain function to remove the steady-state noise signal from the noisy sound signal and obtain the noise intensity of the target sound signal in the noisy sound signal; based on the noise intensity of the target sound signal, the target sound signal is obtained. The present disclosure uses the noise reduction model to estimate the noise of the steady-state noise signal, which can effectively improve the speed and accuracy of noise estimation, obtain a better noise reduction effect, avoid problems such as noise residue, speech distortion, and sudden changes in volume due to inaccurate noise estimation, and the learning difficulty of the noise reduction model is low, effectively reducing the computational amount in the noise reduction process and improving the applicability of the noise reduction method.
[0066] Exemplary embodiments of the present disclosure provide a noise reduction method, which is applied to a mobile terminal. The mobile terminal may be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a smart watch, etc. Since the noise reduction method in the present disclosure uses a simple algorithm and requires less computing power during the noise reduction process, it can be applied to mobile terminals, which can not only ensure the noise reduction effect, but also improve the noise reduction speed, and has a wide range of application scenarios.
[0067] As Figure 1 shown, the noise reduction method shown in the present disclosure includes:
[0068] S101. Obtain a noisy sound signal, which includes a steady-state noise signal and a target sound signal;
[0069] S102. Based on the noisy sound signal and a pre-stored noise reduction model, obtain the noise intensity of the steady-state noise signal;
[0070] S103. Based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function, remove the steady-state noise signal from the noisy sound signal to obtain the noise intensity of the target sound signal;
[0071] S104. Based on the noise intensity of the target sound signal, obtain the target sound signal.
[0072] In step S101, the noisy sound signal may be obtained based on the microphone device built in the mobile terminal. For example, when the microphone of the mobile terminal is turned on, the microphone collects the user's voice input (including the sounds in the surrounding environment when the user is speaking), and saves it as the noisy sound signal; the noisy sound signal may also be obtained based on network communication. For example, when the user uses an instant messaging application such as WeChat to make a voice call with relatives and friends, the mobile terminal obtains the noisy sound signal sent by the relatives and friends through network interaction.
[0073] For the sound signal collected by the mobile terminal, the user's voice is the sound signal that wants to be retained, and other sound signals (such as the sound signals in the environment) can be considered as noise signals. Noise signals can be divided into steady-state noise and non-steady-state noise. Steady-state noise refers to noise whose amplitude and frequency do not change significantly over time, such as the sound of an air conditioner running, the sound of a car driving, etc. Non-steady-state noise refers to sounds whose amplitude and frequency change over time, such as the sound of thunder, the sound of birds chirping, etc. Since in actual life scenarios, the sound environment is complex and the existence of steady-state noise cannot be avoided, the noisy sound signal obtained by the mobile terminal includes a steady-state noise signal and a target sound signal including speech content.
[0074] In some scenarios, the noisy voice signal includes not only a steady-state noise signal and a target voice signal, but also a non-steady-state noise signal. Since the non-steady-state noise signal can reflect the user's scenario information and exists for a short time, it will not have a great impact on the user's interpretation of the target voice signal. For example, during a phone call in a thunderstorm, the thunder sound, although it belongs to the non-steady-state noise signal, can reflect the environment where the user is located and can enrich the call experience of the person on the other end of the call during the user's call; another example is that when a user makes a voice call through an application in the forest, the chirping of insects and birds in the forest can give the person on the other end of the call an immersive feeling. Therefore, when performing noise reduction processing on the voice signal, the non-steady-state noise signal can be left unprocessed, and only the steady-state noise signal is processed for noise reduction. The main purpose of the present disclosure is to remove the steady-state noise signal in the noisy voice signal through noise reduction processing, retain and restore the target voice signal in the noisy voice signal, so that the user can clearly know the voice information carried in the target voice signal.
[0075] In step S102, the noise reduction model is pre-stored in the mobile terminal and is used to estimate the steady-state noise signal in the noisy voice signal. When signal processing of the noisy voice signal is required, the noise reduction model can be called. In order to enable the noise reduction model to better estimate the steady-state noise signal, the mobile terminal needs to extract the features of the noisy voice signal, complete operations such as frame division and windowing, Fourier transform, etc., convert the noisy voice signal from a time-domain signal to a frequency-domain signal, and obtain the feature information of the noisy voice signal, and input it into the pre-stored noise reduction model.
[0076] The noise reduction model is obtained by training a multi-layer neural network using a large number of training noisy voice signals. During the training process, the gain function of the steady-state noise signal in the training noisy voice signal is used as the learning target, so that the trained noise reduction model can output the gain function of the steady-state noise signal based on the noisy voice signal. Since the gain function reflects the ratio of the energy levels of the steady-state noise signal and the noisy voice signal at each frequency point, the mobile terminal can calculate and obtain the noise intensity of the steady-state noise signal by using the gain function of the steady-state noise signal output by the noise reduction model and the noisy voice signal.
[0077] The basis of the noise reduction model can be a stack of one or more of CNN (Convolutional Neural Networks), DNN (Deep Neural Networks), RNN (Recurrent Neural Networks), LSTM (Long Short-Term Memory), activation functions, etc. Since the noise reduction model needs to process the steady-state noise signal in the noisy voice signal, and the steady-state noise signal has relatively small changes in amplitude and frequency over time and has a small information entropy. Compared with estimating the intensities of non-steady-state noise signals and target voice signals, the difficulty of noise estimation for the steady-state noise signal by the noise reduction model is much reduced. Therefore, in order to reduce the computing power consumption of the mobile terminal, on the premise of ensuring the accuracy of noise estimation, the parameter scale of the noise reduction model can be greatly reduced, and the number of nodes in each layer of neural network in the noise reduction model can be reduced to reduce the algorithm complexity of the noise reduction model.
[0078] Since the noise reduction method in this disclosure is expected to be used on mobile terminals, and the computing power of mobile terminals is relatively weak compared to that of large computing devices such as servers, when establishing a noise reduction model, a relatively simple noise reduction model can be selected, so as to ensure the noise reduction effect and reduce the noise tracking time.
[0079] In step S103, the mobile terminal uses the optimal log amplitude spectrum estimator to estimate the target voice signal in the noisy voice signal, and the first estimator gain function is the optimal log amplitude spectrum estimator gain function. Since the noise intensity of the steady-state noise signal will affect the parameter values of the first estimator gain function, therefore, based on the first estimator gain function, the energy proportion of the steady-state noise signal in the noisy voice signal can be calculated. The mobile terminal suppresses the steady-state noise signal in the noisy voice signal based on the result of the first estimator gain function. The suppression process is to adjust the frequency values corresponding to the steady-state noise signal on the amplitude spectrum of the noisy voice signal according to the result of the first estimator gain function, so as to achieve the filtering effect of removing the steady-state noise signal from the noisy voice signal without having too much impact on the voice signals of other frequencies in the target voice signal except for the steady-state noise signal. The amplitude spectrum of the noisy voice signal after filtering is the amplitude spectrum of the target voice signal, and the amplitude spectrum can reflect the distribution of the amplitude of the target voice signal at each frequency point, so as to obtain the noise intensity of the target voice signal.
[0080] In step S104, after obtaining the noise intensity of the target voice signal, the mobile terminal can use the inverse Fourier transform to reconstruct the time-domain waveform of the target voice signal, and then the target voice signal in the frequency domain can be obtained, and the voice signal without the steady-state noise signal can be restored, enabling the user to clearly hear the target voice signal.
[0081] In the present disclosure, using a pre-stored noise reduction model to estimate the noise of a steady-state noise signal can effectively improve the speed and accuracy of noise estimation, obtain a better noise reduction effect, avoid problems such as noise residue, speech distortion, and sudden changes in volume due to inaccurate noise estimation, and the learning difficulty of the noise reduction model is low, effectively reducing the computational amount in the noise reduction process and improving the applicability of the noise reduction method.
[0082] According to an exemplary embodiment, as Figure 2 shown, the noise reduction method in this embodiment includes:
[0083] S201. Obtain a noisy voice signal, where the noisy voice signal includes a steady-state noise signal and a target voice signal;
[0084] S202. Based on the noisy voice signal, obtain the noise intensity of the noisy voice signal;
[0085] S203. Input the noise intensity of the noisy voice signal into the pre-stored noise reduction model to obtain the gain function of the steady-state noise signal;
[0086] S204. Based on the gain function of the steady-state noise signal and the frequency domain characteristics of the noisy voice signal, obtain the noise intensity of the steady-state noise signal;
[0087] S205. Based on the noise intensity of the steady-state noise signal and the pre-stored first estimator gain function, remove the steady-state noise signal from the noisy voice signal to obtain the noise intensity of the target voice signal;
[0088] S206. Based on the noise intensity of the target voice signal, obtain the target voice signal.
[0089] Among them, the implementation manners of steps S201, S205, and S206 are the same as those of steps S101, S103, and S104 in the above embodiment, and will not be elaborated here.
[0090] In step S202, since the noisy voice signal includes a steady-state noise signal and a target voice signal, in order to enable the noise reduction model to accurately estimate the steady-state noise signal, the mobile terminal needs to perform signal processing on the noisy voice signal to obtain the noise intensity of the noisy voice signal, which reflects the energy distribution of each frequency point in the noisy voice signal. The mobile terminal uses the noise intensity of the noisy voice signal as the feature information of the noisy voice signal and inputs it into the noise reduction model.
[0091] In some embodiments, step S202 based on the noisy voice signal to obtain the noise intensity of the noisy voice signal includes:
[0092] Frame and window the noisy voice signal and perform short-time Fourier transform to obtain the frequency-domain features of each frame of the noisy voice signal;
[0093] Based on the frequency-domain features of each frame of the noisy voice signal, obtain the noise intensity of each frame of the noisy voice signal.
[0094] Since the noisy voice signal has short-term stationarity, in order to ensure real-time processing of the noisy voice signal, frame and window the noisy voice signal. Framing means segmenting a segment of the noisy voice signal according to a specified length and dividing several sampling points into one frame. To avoid large changes between adjacent frames, overlapping sampling point data is retained between frames. Windowing means introducing a window function to make the framed noisy voice signal continuous, and each frame will exhibit the characteristics of a periodic function. Perform short-time Fourier transform on the framed and windowed noisy voice signal to convert the time-domain signal into a frequency-domain signal and obtain the frequency-domain features of each frame of the noisy voice signal.
[0095] The noisy voice signal is represented by the symbol y(n), the target voice signal is represented by the symbol x(n), and the steady-state noise signal is represented by the symbol d(n), that is, y(n) = x(n) + d(n), where n represents the number of sampling points, and each sampling point corresponds to a digital signal. Perform short-time Fourier transform on the noisy voice signal y(n) after framing and windowing to obtain the amplitude spectrum Y(t,f) of the noisy voice signal. The frequency-domain expression of the amplitude spectrum of the noisy voice signal: Y(t,f) = X(t,f) + D(t,f). Among them, t represents the frame index, and f represents the frequency index, that is, the corresponding number of frames or frequencies of the noisy voice signal can be located through the frame index or frequency index.
[0096] In order to enable the noise reduction model to accurately estimate the steady-state noise signal, feature extraction needs to be performed on the noisy voice signal. Since the Log Power Spectra (LPS) can intuitively reflect the frequency components of the signal and highlight the intensity differences of different frequency components, the logarithmic power spectrum of the noisy voice signal can be calculated based on the frequency-domain features of each frame of the noisy voice signal to obtain the noise intensity of each frame of the noisy voice signal.
[0097] Among them, the logarithmic power spectrum of the noisy voice signal y(n) is expressed as: Y LPS (t,f) = log|Y(t,f)| 2 . The logarithmic power spectrum arranges the logarithm values of the spectrum in frequency order, thereby converting the non-linear relationship into a linear relationship, intuitively reflecting the frequency components and intensity distribution of the noisy voice signal, and facilitating the noise reduction model to estimate the steady-state noise signal of the noisy voice signal.
[0098] Here, it should be noted that in the noise reduction methods of each embodiment of the present disclosure, each frame of the noisy voice signal is processed, and after each frame is processed, it is finally combined into a complete voice signal for playback.
[0099] In step S203, the noise reduction model processes the noise intensity of the input noisy voice signal and outputs the gain function of the steady-state noise signal. Since the gain function reflects the ratio of the steady-state noise signal to the noisy voice signal, step S204 can calculate the amplitude spectrum of the steady-state noise signal based on the gain function of the steady-state noise signal and the frequency domain characteristics of the noisy voice signal. Obtaining the noise intensity of the steady-state noise signal reflects the energy level of the steady-state noise signal in the noisy voice signal. Specifically, the amplitude spectrum of the steady-state noise signal The calculation process is as follows:
[0100]
[0101] In some embodiments, step S205 obtains the noise intensity of the target voice signal based on the noise intensity of the steady-state noise signal and the pre-stored first estimator gain function, including:
[0102] Determine the variable parameter in the first estimator gain function based on the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal in each frame of the noisy voice signal;
[0103] Obtain the second estimator gain function based on the determined variable parameter and the first estimator gain function;
[0104] Obtain the noise intensity of the target voice signal based on the second estimator gain function and the frequency domain characteristics of the noisy voice signal.
[0105] The first estimator gain function is G LSA (ξ f , v f ), where v f is the variable parameter in the first estimator gain function, and the variable parameter v f is determined based on the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal in each frame of the noisy voice signal. In the above process, through signal processing of the noisy voice signal, the noise intensity of the noisy voice signal can be obtained, and based on the noise reduction model, the noise intensity of the steady-state noise signal can be obtained, that is, the noise intensity of the steady-state noise signal and the noise intensity of the noisy voice signal are known quantities, so the variable parameter in the first estimator gain function can be determined. Substitute the determined variable parameter into the first estimator gain function to obtain the second estimator gain function.
[0106] The second estimator gain function can calculate the energy ratio of the steady-state noise signal in the noisy voice signal, and obtain the respective proportions of the target voice signal and the steady-state noise signal in a segment of the noisy voice signal. For example, using the second estimator gain function, it can be known that in a certain segment of the noisy voice signal, the target voice signal accounts for 80% of the noisy voice signal, and the steady-state noise signal accounts for 20% of the noisy voice signal. In the frequency domain, multiplying the second estimator gain function by the frequency domain characteristics of the noisy voice signal can filter out the steady-state noise signal and only retain the target voice signal to obtain the noise intensity of the target voice signal.
[0107] In some embodiments, based on the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal, determine the variable parameters in the first estimator gain function, including:
[0108] Based on the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal, obtain the posterior signal-to-noise ratio of each frame of the noisy voice signal;
[0109] Based on the posterior signal-to-noise ratio of each frame of the noisy voice signal, determine the prior signal-to-noise ratio of this frame of the noisy voice signal;
[0110] Based on the posterior signal-to-noise ratio and the prior signal-to-noise ratio of each frame of the noisy voice signal, determine the variable parameters of the first estimator gain function.
[0111] The first estimator gain function is where v f represents the variable parameter, and ξ f represents the prior signal-to-noise ratio. Since the posterior signal-to-noise ratio γ f characterizes the power ratio of the noisy voice signal and the steady-state noise signal, the posterior signal-to-noise ratio γ can be determined by the following formula using the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal f :
[0112]
[0113] where is the amplitude spectrum of the steady-state noise signal in each frame of the noisy voice signal, Y(t,f) is the amplitude spectrum of each frame of the noisy voice signal, t represents the frame index, and f represents the frequency index.
[0114] After determining the posterior signal-to-noise ratio γ f , the posterior signal-to-noise ratio γ f can be further used to determine the prior signal-to-noise ratio ξ f , and then the variable parameter v f is determined by the following formula:
[0115]
[0116] In some embodiments, determining the prior signal-to-noise ratio of each frame of noisy voice signal based on the posterior signal-to-noise ratio of each frame of noisy voice signal includes:
[0117] The prior signal-to-noise ratio of each frame of noisy voice signal is determined based on the posterior signal-to-noise ratio of this frame and the prior signal-to-noise ratio of the previous frame of noisy voice signal.
[0118] Since the prior signal-to-noise ratio ξ f represents the power ratio of the target voice signal and the steady-state noise signal, and the target voice signal is unknown, the prior signal-to-noise ratio ξ of each frame of noisy voice signal can be determined by the posterior signal-to-noise ratio of this frame and the prior signal-to-noise ratio of the previous frame of noisy voice signal f , that is, ξ is determined by the following formula f :
[0119] ξ t,f = aξ t-1,f +(1 - a)max(γ t,f - 1, 0)
[0120] where a is a smoothing parameter, and its value range is [0, 1], which can be determined according to the actual usage scenario. By adjusting the smoothing parameter a, ξ t,f can have a smoothing effect and effectively eliminate abnormal voice signals caused by the algorithm. The smoothing parameter is usually set and determined before the mobile terminal leaves the factory and does not change during the signal processing process. ξ t-1,f represents the prior signal-to-noise ratio of the previous frame of noisy voice signal, γ t,f represents the posterior signal-to-noise ratio of this frame, and ξ t,f represents the prior signal-to-noise ratio of this frame.
[0121] Through the above process, step S205 can obtain all the parameters of the first estimator gain function G LSA (ξ f , v f ). After obtaining G LSA (ξ f , v f ), the amplitude spectrum of the target voice signal can be solved, that is
[0122]
[0123] The amplitude spectrum of the target voice signal can reflect the distribution of the amplitude of the target voice signal at each frequency point to obtain the noise intensity of the target voice signal.
[0124] In the present disclosure, the noise reduction model can accurately estimate the steady-state noise signal in the noisy sound signal, and then multiply the amplitude spectrum of the noisy sound signal by the first estimator gain function to achieve the effect of filtering the steady-state noise signal, obtain the target sound signal without noise residue, improve the noise reduction effect of the speech signal, and enhance the user experience.
[0125] According to an exemplary embodiment, as Figure 3 shown, the noise reduction method in this embodiment includes:
[0126] S301. Obtain the training noisy sound signal;
[0127] S302. Train the neural network model based on the training noisy sound signal to obtain the noise reduction model;
[0128] S303. Obtain the noisy sound signal, where the noisy sound signal includes a steady-state noise signal and a target sound signal;
[0129] S304. Based on the noisy sound signal and the pre-stored noise reduction model, obtain the noise intensity of the steady-state noise signal;
[0130] S305. Based on the noise intensity of the steady-state noise signal and the pre-stored first estimator gain function, remove the steady-state noise signal from the noisy sound signal to obtain the noise intensity of the target sound signal;
[0131] S306. Based on the noise intensity of the target sound signal, obtain the target sound signal.
[0132] Among them, the implementation manners of steps S303 - S306 are the same as those of steps S101 - S104 in the above embodiment, and will not be elaborated here.
[0133] In step S301, in order to ensure that the noise reduction model can accurately distinguish the steady-state noise signal from other noises, the obtained training noisy sound signal includes a target signal and a non-target signal. Among them, the target signal is the steady-state noise, represented by the symbol d(n), and the non-target signal includes one or more of non-steady-state noise, music, speech, and animal calls, represented by the symbol x(n). The target signal d(n) and the non-target signal x(n) are mixed through a certain signal-to-noise ratio to obtain the training noisy sound signal y(n):
[0134] y(n) = x(n) + d(n)
[0135] In order to ensure that the training noisy audio signals can be diverse and the neural network model can be well trained, there is no restriction on the signal-to-noise ratio. Different signal-to-noise ratios can be used to mix the same target signal and non-target signal. Different signal-to-noise ratios can also be used to mix different target signals and non-target signals to obtain multiple training noisy audio signals.
[0136] For example, a steady-state noise can be mixed with a non-steady-state noise at a signal-to-noise ratio A to form a training noisy sound signal a; a steady-state noise can be mixed with a non-steady-state noise at a signal-to-noise ratio B to form a training noisy sound signal b; a steady-state noise can be mixed with a piece of music at a signal-to-noise ratio A to form a training noisy sound signal c; a steady-state noise can be mixed with a voice at a signal-to-noise ratio B to form a training noisy sound signal d; a steady-state noise can be mixed with an animal's call at a signal-to-noise ratio B to form a training noisy sound signal e; a steady-state noise can be mixed with a non-steady-state noise and an animal's call at a signal-to-noise ratio C to form a training noisy sound signal f; a steady-state noise can be mixed with a non-steady-state noise, a voice, and an animal's call at a signal-to-noise ratio D to form a training noisy sound signal g; a steady-state noise can be mixed with a non-steady-state noise, a voice, a piece of music, and an animal's call at a signal-to-noise ratio E to form a training noisy sound signal h. The mixed training noisy sound signal ah is used to train the neural network model.
[0137] In step S302, all training noisy sound signals are processed by frame segmentation, windowing, short-time Fourier transform, etc. to obtain the logarithmic power spectrum of the training noisy sound signal. The logarithmic power spectrum of the training noisy sound signal is input into the neural network model for repeated training, so that the neural network model finally trained can output the gain function of the target signal, and the trained neural network model is the noise reduction model.
[0138] Among them, the neural network model can be a single or multiple stacks of CNN, DNN, RNN, LSTM, activation function, etc. Since the target signal to be processed is steady-state noise, the amplitude and frequency will not change with time, so the structure of the neural network model does not need to be very complicated, and the number of nodes in each layer of the neural network can be appropriately reduced, as long as the neural network model can output the gain function of the target signal That's it.
[0139] In some embodiments, the method for obtaining a noise reduction model further includes:
[0140] Determining a reference gain function based on the frequency domain characteristics of the training noisy sound signal and the frequency domain characteristics of the target signal;
[0141] Modify the noise reduction model based on the reference gain function.
[0142] Since the target signal in the training noisy voice signal is known, after performing short-time Fourier transform on the training noisy voice signal, the frequency domain characteristics of the training noisy voice signal and the frequency domain characteristics of the target signal can be obtained, that is, the power spectrum of the training noisy voice signal and the power spectrum of the target signal can be obtained, and then the reference gain function G(t,f) of the training noisy voice signal can be calculated. The calculation process is as follows:
[0143]
[0144] where D(t,f) is the amplitude spectrum of the target signal, and Y(t,f) is the amplitude spectrum of the training noisy voice signal.
[0145] To improve the accuracy of the noise reduction model, calculate the loss function of the gain function output by the noise reduction model and the reference gain function G(t,f) to obtain error information, and optimize the noise reduction model through the backpropagation algorithm. In this way, after training with a large number of training noisy voice signals and optimizing the noise reduction model, it can accurately output the gain function of the steady-state noise signal in the noisy voice signal.
[0146] In the present disclosure, use the trained neural network model to estimate the steady-state noise signal. Since the steady-state noise signal changes less with time, the estimation difficulty of the neural network model is reduced. Therefore, a smaller neural network model can be used to obtain better training results, avoiding the problem that it cannot be applied to mobile terminals due to high algorithm complexity. Moreover, using the reference gain function to modify and optimize the noise reduction model can effectively improve the accuracy of the noise reduction model in estimating the steady-state noise signal.
[0147] In one example, as Figure 4 shown, the noise reduction method is divided into two parts: the training stage and noise estimation. In the training stage, the target signal and the non-target signal are mixed according to a certain signal-to-noise ratio to form a training noisy voice signal, extract the characteristics of the training noisy voice signal, input them into the neural network model, train the neural network model, and use the trained neural network model as the noise reduction model. To improve the output performance of the neural network model, calculate the loss function of the gain function output by the neural network model and the reference gain function of the training noisy voice signal, and use the backpropagation algorithm to optimize the neural network model, so that the finally determined noise reduction model can accurately estimate the steady-state noise signal.
[0148] In the noise estimation process, first, feature extraction is performed on the noisy speech signal as the input of the noise reduction model. The gain function of the stationary noise signal is estimated using the noise reduction model, and thus the amplitude spectrum of the stationary noise signal is calculated. Using the amplitude spectrum of the stationary noise signal and the amplitude spectrum of the noisy speech signal, the a priori signal-to-noise ratio of the noisy speech signal can be calculated. Using the a priori signal-to-noise ratio, the first estimator gain function can be calculated. Using the first estimator gain function, the amplitude spectrum of the target sound signal is calculated, and then the time-domain waveform of the target sound signal is reconstructed to obtain the target sound signal.
[0149] An exemplary embodiment of the present disclosure provides a noise reduction device applied to a mobile terminal. As Figure 5 shown, a block diagram of a noise reduction device shown in the present disclosure.
[0150] The block diagram includes: an acquisition module 51 and a processing module 52. The acquisition module 51 is used to acquire a noisy sound signal, which includes a stationary noise signal and a target sound signal; the processing module 52 is used to obtain the noise intensity of the stationary noise signal based on the noisy sound signal and a pre-stored noise reduction model; the processing module 52 is further used to remove the stationary noise signal from the noisy sound signal based on the noise intensity of the stationary noise and a pre-stored first estimator gain function to obtain the noise intensity of the target sound signal; the processing module 52 is further used to obtain the target sound signal based on the noise intensity of the target sound signal.
[0151] In an exemplary embodiment of the present disclosure, the processing module 52 is further used to: obtain the noise intensity of the noisy sound signal based on the noisy sound signal; input the noise intensity of the noisy sound signal into a pre-stored noise reduction model to obtain the gain function of the stationary noise signal; obtain the noise intensity of the stationary noise signal based on the gain function of the stationary noise signal and the frequency-domain characteristics of the noisy sound signal.
[0152] In an exemplary embodiment of the present disclosure, the processing module 52 is further used to: perform frame windowing processing and short-time Fourier transform on the noisy sound signal to obtain the frequency-domain characteristics of each frame of the noisy sound signal; obtain the noise intensity of each frame of the noisy sound signal based on the frequency-domain characteristics of each frame of the noisy sound signal.
[0153] In an exemplary embodiment of the present disclosure, the processing module 52 is further used to: determine the variable parameters in the first estimator gain function based on the noise intensity of the stationary noise signal in each frame of the noisy sound signal and the noise intensity of the noisy sound signal; obtain the second estimator gain function based on the determined variable parameters and the first estimator gain function; obtain the noise intensity of the target sound signal based on the second estimator gain function and the frequency-domain characteristics of the noisy sound signal.
[0154] In an exemplary embodiment of the present disclosure, the processing module 52 is further configured to: obtain the posterior signal-to-noise ratio of each frame of noisy voice signal based on the noise intensity of the steady-state noise signal in each frame of noisy voice signal and the noise intensity of the noisy voice signal; determine the prior signal-to-noise ratio of the frame of noisy voice signal based on the posterior signal-to-noise ratio of each frame of noisy voice signal; and determine the variable parameters of the first estimator gain function based on the posterior signal-to-noise ratio and the prior signal-to-noise ratio of each frame of noisy voice signal.
[0155] In an exemplary embodiment of the present disclosure, the processing module 52 is further configured to: determine the prior signal-to-noise ratio of each frame of noisy voice signal based on the posterior signal-to-noise ratio of the frame and the prior signal-to-noise ratio of the previous frame of noisy voice signal.
[0156] In an exemplary embodiment of the present disclosure, the acquisition module 51 is further configured to: acquire a training noisy voice signal, where the training noisy voice signal includes a target signal and a non-target signal, the target signal is steady-state noise, and the non-target signal includes one or more of non-steady-state noise, music, speech, and animal calls;
[0157] Train a neural network model based on the training noisy voice signal to obtain a noise reduction model.
[0158] In an exemplary embodiment of the present disclosure, the processing module 52 is further configured to: determine a reference gain function based on the frequency domain features of the training noisy voice signal and the frequency domain features of the target signal; and correct the noise reduction model based on the reference gain function.
[0159] Regarding the noise reduction device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0160] Figure 6 It is a block diagram of a mobile terminal 600 shown according to an exemplary embodiment. For example, the mobile terminal 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0161] Refer to Figure 6 , the mobile terminal 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0162] The processing component 602 generally controls the overall operations of the mobile terminal 600, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above - mentioned methods. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0163] The memory 604 is configured to store various types of data to support the operations of the mobile terminal 600. Examples of such data include instructions for any application or method operating on the mobile terminal 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read - only memory (EEPROM), erasable programmable read - only memory (EPROM), programmable read - only memory (PROM), read - only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0164] The power component 606 provides power to various components of the mobile terminal 600. The power component 606 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the mobile terminal 600.
[0165] The multimedia component 608 includes a screen that provides an output interface between the mobile terminal 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 608 includes a front - facing camera and / or a rear - facing camera. When the mobile terminal 600 is in an operating mode, such as a shooting mode or a video mode, the front - facing camera and / or the rear - facing camera can receive external multimedia data. Each front - facing camera and rear - facing camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0166] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the mobile terminal 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.
[0167] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a start button, and a lock button.
[0168] The sensor component 614 includes one or more sensors for providing an assessment of various aspects of the state of the mobile terminal 600. For example, the sensor component 614 can detect the on / off state of the mobile terminal 600, the relative positioning of components, such as the display and keypad of the mobile terminal 600. The sensor component 614 can also detect a change in the position of the mobile terminal 600 or a component of the mobile terminal 600, the presence or absence of user contact with the mobile terminal 600, the orientation or acceleration / deceleration of the mobile terminal 600, and a change in the temperature of the mobile terminal 600. The sensor component 614 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 614 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0169] The communication component 616 is configured to facilitate communication between the mobile terminal 600 and other devices in a wired or wireless manner. The mobile terminal 600 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0170] In an exemplary embodiment, the mobile terminal 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0171] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the mobile terminal 600 to complete the above noise reduction method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0172] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of the mobile terminal, enables the processing device of the mobile terminal to execute the noise reduction method provided by the exemplary embodiments of the present disclosure.
[0173] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the content disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0174] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A noise reduction method, characterized in that, Applied to a mobile terminal, the noise reduction method includes: Obtain a noisy voice signal, where the noisy voice signal includes a steady-state noise signal and a target voice signal; Based on the noisy voice signal and a pre-stored noise reduction model, obtain the noise intensity of the steady-state noise signal; Based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function, remove the steady-state noise signal from the noisy voice signal to obtain the noise intensity of the target voice signal; Based on the noise intensity of the target voice signal, obtain the target voice signal.
2. The noise reduction method according to claim 1, wherein The step of obtaining the noise intensity of the steady-state noise signal based on the noisy voice signal and a pre-stored noise reduction model includes: Based on the noisy voice signal, obtain the noise intensity of the noisy voice signal; Input the noise intensity of the noisy voice signal into a pre-stored noise reduction model to obtain the gain function of the steady-state noise signal; Based on the gain function of the steady-state noise signal and the frequency domain characteristics of the noisy voice signal, obtain the noise intensity of the steady-state noise signal.
3. The noise reduction method according to claim 2, wherein The step of obtaining the noise intensity of the noisy voice signal based on the noisy voice signal includes: Perform frame addition and windowing processing on the noisy voice signal and perform short-time Fourier transform to obtain the frequency domain characteristics of each frame of the noisy voice signal; Based on the frequency domain characteristics of each frame of the noisy voice signal, obtain the noise intensity of each frame of the noisy voice signal.
4. The noise reduction method according to claim 3, wherein The step of obtaining the noise intensity of the target voice signal based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function includes: Based on the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal, determine the variable parameters in the first estimator gain function; Based on the determined variable parameters and the first estimator gain function, obtain a second estimator gain function; Based on the second estimator gain function and the frequency domain characteristics of the noisy voice signal, obtain the noise intensity of the target voice signal.
5. The noise reduction method according to claim 4, wherein The step of determining the variable parameters in the first estimator gain function based on the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal includes: Based on the noise intensity of the steady-state noise signal in each frame of the noisy voice signal and the noise intensity of the noisy voice signal, obtain the posterior signal-to-noise ratio of each frame of the noisy voice signal; Based on the posterior signal-to-noise ratio of each frame of the noisy voice signal, determine the prior signal-to-noise ratio of this frame of the noisy voice signal; Based on the posterior signal-to-noise ratio and the prior signal-to-noise ratio of each frame of the noisy voice signal, determine the variable parameters of the first estimator gain function.
6. The noise reduction method according to claim 5, wherein The step of determining the prior signal-to-noise ratio of a frame of the noisy voice signal based on the posterior signal-to-noise ratio of each frame of the noisy voice signal includes: The prior signal-to-noise ratio of each frame of the noisy voice signal is determined based on the posterior signal-to-noise ratio of this frame and the prior signal-to-noise ratio of the previous frame of the noisy voice signal.
7. The noise reduction method according to any one of claims 1-6, characterized in that, The method for obtaining the noise reduction model includes: Obtain a training noisy voice signal, where the training noisy voice signal includes a target signal and a non-target signal, the target signal is steady-state noise, and the non-target signal includes one or more of non-steady-state noise, music, speech, and animal calls; Train a neural network model based on the training noisy voice signal to obtain the noise reduction model.
8. The noise reduction method according to claim 7, wherein The method for obtaining the noise reduction model further includes: Determine a reference gain function based on the frequency domain characteristics of the training noisy voice signal and the frequency domain characteristics of the target signal; Modify the noise reduction model based on the reference gain function.
9. A noise reduction device, characterized in that, Applied to a mobile terminal, the noise reduction device includes: An acquisition module, configured to acquire a noisy voice signal, where the noisy voice signal includes a steady-state noise signal and a target voice signal; A processing module, configured to obtain the noise intensity of the steady-state noise signal based on the noisy voice signal and a pre-stored noise reduction model; The processing module is further configured to remove the steady-state noise signal from the noisy voice signal based on the noise intensity of the steady-state noise signal and a pre-stored first estimator gain function to obtain the noise intensity of the target voice signal; The processing module is further configured to obtain the target voice signal based on the noise intensity of the target voice signal.
10. A mobile terminal, characterized in that, The mobile terminal includes: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the executable instructions in the memory to implement the noise reduction method according to any one of claims 1 to 8.
11. A non - transitory computer - readable storage medium, characterized in that, A computer program capable of being run is stored thereon, and when the computer program is executed by the processor of the mobile terminal, it is used to implement the noise reduction method according to any one of claims 1 to 8.