Voice noise reduction method and device, equipment and storage medium
By combining the first and second noise reduction algorithms, gain calculation and noise reduction processing are performed for different types of voice frames, the problem of poor voice noise reduction effect in the prior art is solved, and the clarity and user experience of the voice signal are improved.
Patent Information
- Application Number
- CN202510138156.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art lacks tracking of non-stationary noise during speech noise reduction. Traditional noise reduction algorithms may change the spectrum characteristics of the audio, resulting in the loss of high frequency of speech and poor user experience.
By determining the spectrum of the speech signal to be processed, the gain is calculated based on the first noise reduction algorithm, and combined with the second noise reduction algorithm, the corresponding gain is determined for different types of speech frames, and the noise reduction process is performed.
It improves the voice noise reduction effect, reduces the distortion of the clear sound components, and improves the user experience.
Smart Images

Figure CN120148534A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of signal processing, and particularly to a voice noise reduction method, device, equipment and storage medium. Background Art
[0002] In audio data processing, voice noise reduction is a very important technology, and its main goal is to recover the clean voice we want from the "contaminated" voice with different noises collected by the microphone.
[0003] Voice noise reduction involves a very wide range of application fields, including voice calls, teleconferences, scene recordings, military eavesdropping, hearing aid devices and speech recognition devices, etc., and has become a preprocessing module for many speech coding and recognition systems. The algorithms used for voice noise reduction include traditional noise reduction and artificial intelligence (AI) noise reduction, etc. In the process of voice noise reduction, traditional noise reduction algorithms have problems of insufficient tracking ability for non-stationary noise and poor estimation ability, while basic AI noise reduction algorithms may change the spectral characteristics of the audio during the process of enhancing noise reduction, resulting in the loss or distortion of some details of the original speech signal. This loss is mainly manifested as the loss of high-frequency speech, that is, the loss of voiceless sounds. If the above methods are used for voice noise reduction in related devices, the user experience will be poor. Summary of the Invention
[0004] In view of this, this application provides a voice noise reduction method, device, equipment and storage medium, which helps to solve the problems of poor voice noise reduction effect and low user experience in the prior art.
[0005] In a first aspect, an embodiment of this application provides a voice noise reduction method, including: determining a first spectrum of a first voice signal to be processed; determining a first gain of the first spectrum based on a first noise reduction algorithm, and determining a second spectrum according to the first gain and the first spectrum; determining types of each voice frame included in the second spectrum according to the first gain and the second spectrum; determining a second gain of the first spectrum based on a second noise reduction algorithm; obtaining a third gain of each voice frame according to the types of each voice frame, the first gain and the second gain; determining a second voice signal based on the third gain of each voice frame and the second spectrum; the second voice signal is the voice signal after noise reduction of the first voice signal.
[0006] In a possible implementation manner of the first aspect, the determining a first spectrum of a first voice signal to be processed includes: performing frame addition and windowing processing on the first voice signal to obtain a plurality of voice frames; performing Fourier transform on the voice frames to obtain a first spectrum corresponding to each voice frame.
[0007] In a possible implementation of the first aspect, determining the first gain of the first spectrum based on the first noise reduction algorithm includes:
[0008] Performing noise estimation on the first spectrum; calculating a priori signal-to-noise ratio according to the noise estimation result; and determining the first gain according to the a priori signal-to-noise ratio by using a preset gain algorithm.
[0009] In a possible implementation of the first aspect, determining the types of each speech frame included in the second spectrum according to the first gain and the second spectrum includes: for each speech frame included in the second spectrum, calculating the difference between the low-frequency energy value and the high-frequency energy value in the speech frame; if the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than a first preset threshold, determining that the speech frame is a voiced audio signal; if the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than the first preset threshold, determining a target value corresponding to the speech frame based on the first gain; the target value is the sum of high-frequency gains or the sum of low-frequency gains corresponding to the speech frame, and when the sum of high-frequency gains corresponding to the speech frame is greater than the sum of low-frequency gains, the target gain value is the sum of high-frequency gains corresponding to the speech frame, or when the sum of low-frequency gains corresponding to the speech frame is greater than the sum of high-frequency gains, the target gain value is the sum of low-frequency gains corresponding to the speech frame; if the target value corresponding to the speech frame is greater than a second preset threshold, determining that the speech frame is an unvoiced audio signal; if the target value corresponding to the speech frame is not greater than the second preset threshold, determining that the speech frame is a voiced audio signal.
[0010] In a possible implementation of the first aspect, obtaining the third gain of each speech frame according to the types of each speech frame, the first gain, and the second gain includes: for each speech frame included in the second spectrum, determining a fusion coefficient when the speech frame is an unvoiced audio signal; performing a merging process on the first gain and the second gain based on the fusion coefficient to obtain the third gain of the speech frame; when the speech frame is a voiced audio signal, using the second gain as the third gain of the speech frame.
[0011] In a possible implementation of the first aspect, determining the second gain of the first spectrum based on the second noise reduction algorithm includes: inputting the first spectrum into a prediction network model to obtain a mask vector; and determining the mask vector as the second gain.
[0012] In a possible implementation of the first aspect, determining the second speech signal based on the third gain of each speech frame and the second spectrum includes: multiplying each speech frame in the second spectrum by the corresponding third gain to obtain a third spectrum; and obtaining the second speech signal after overlapping and adding processing and inverse Fourier transform of the third spectrum.
[0013] In a second aspect, an embodiment of the present application provides a speech noise reduction device, including: a determination unit configured to determine a first spectrum of a first speech signal to be processed; the determination unit is further configured to determine a first gain of the first spectrum based on a first noise reduction algorithm, and determine a second spectrum according to the first gain and the first spectrum; the determination unit is further configured to determine a frequency band type of each speech frame included in the second spectrum according to the first gain and the second spectrum; the determination unit is further configured to determine a second gain of the first spectrum based on a second noise reduction algorithm; a processing unit configured to perform a merging process on the first gain and the second gain according to the frequency band type of each speech frame to obtain a third gain of each speech frame; the processing unit is further configured to determine a second speech signal based on the third gain of each speech frame and the second spectrum; the second speech signal is the speech signal after noise reduction of the first speech signal.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute the method according to any one of the first aspects described above.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of the first aspects described above.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes executable instructions, and when the executable instructions are executed on a computer, the computer is caused to execute the method according to any one of the first aspects described above.
[0017] Adopt the solution provided by the embodiment of the present application to determine the first spectrum of the first voice signal to be processed; determine the first gain of the first spectrum based on the first noise reduction algorithm, and determine the second spectrum according to the first gain and the first spectrum; determine the types of each voice frame included in the second spectrum according to the first gain and the second spectrum; determine the second gain of the first spectrum based on the second noise reduction algorithm; obtain the third gain of each voice frame according to the types of each voice frame, the first gain and the second gain; determine the second voice signal based on the third gain of each voice frame and the second spectrum. In this way, in the embodiment of the present application, different third gains can be determined for different types of voice frames for noise reduction processing with different effects, so that different types of noise such as stationary noise and non-stationary noise can be subjected to different noise reduction processing, thereby reducing the distortion degree of the voiceless component in the voice signal, improving the voice noise reduction effect, and further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart of a voice noise reduction method provided by an embodiment of the present application;
[0020] Figure 2 It is a schematic flowchart of another voice noise reduction method provided by an embodiment of the present application;
[0021] Figure 3 It is a schematic flowchart of another voice noise reduction method provided by an embodiment of the present application;
[0022] Figure 4 It is a schematic flowchart of another voice noise reduction method provided by an embodiment of the present application;
[0023] Figure 5 It is a schematic flowchart of another voice noise reduction method provided by an embodiment of the present application;
[0024] Figure 6 It is a schematic structural diagram of a voice noise reduction device provided by an embodiment of the present application;
[0025] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the drawings.
[0027] It should be clear that the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0028] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the" and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0029] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0030] The application fields involved in voice noise reduction are very extensive, including voice calls, teleconferences, scene recordings, military eavesdropping, hearing aid devices, and speech recognition devices, etc., and it has become a preprocessing module for many speech coding and recognition systems. The algorithms used for voice noise reduction include traditional noise reduction and artificial intelligence (AI) noise reduction, etc. The traditional speech enhancement algorithm is based on statistical characteristics and assumes that the noise is more stationary than the speech during noise estimation, resulting in the traditional speech enhancement algorithm being only suitable for processing relatively stationary noise, but it is difficult to process non-stationary noise. In recent years, deep learning has been widely used in the noise reduction task. Compared with the disadvantages of traditional signal processing methods that are difficult to process diverse and sudden non-stationary noises, the AI noise reduction method based on deep learning can achieve better noise reduction performance in non-stationary environments on the premise of a large amount of training data and a good model design. However, the AI noise reduction method may change the spectral characteristics of the audio during the voice noise reduction process, resulting in the loss or distortion of some details of the original speech signal. This loss is mainly manifested as the loss of high-frequency speech, that is, the loss of unvoiced sounds.
[0031] In view of the above problems, the embodiments of the present application provide a voice noise reduction method, apparatus, device and storage medium. The method determines the first spectrum of the first voice signal to be processed; determines the first gain of the first spectrum based on the first noise reduction algorithm, and determines the second spectrum according to the first gain and the first spectrum; determines the types of each voice frame included in the second spectrum according to the first gain and the second spectrum; determines the second gain of the first spectrum based on the second noise reduction algorithm; obtains the third gain of each voice frame according to the types of each voice frame, the first gain and the second gain; determines the second voice signal based on the third gain of each voice frame and the second spectrum. In this way, in the embodiments of the present application, different third gains can be determined for different types of voice frames of different voice frames, and noise reduction processing with different effects can be performed. In this way, different noise reduction processing can be performed for different types of noise such as stationary noise and non-stationary noise, so as to reduce the distortion degree of the voiceless components in the voice signal, improve the voice noise reduction effect, and further improve the user experience. The following will be described in detail.
[0032] See Figure 1 , which is a schematic flowchart of a voice noise reduction method provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0033] Step S101, determine the first spectrum of the first voice signal to be processed.
[0034] In the embodiments of the present application, when the voice noise reduction apparatus obtains the first voice signal to be processed, the first voice signal is usually a time-domain signal. In order to more accurately identify and filter the noise in the first voice signal, the first voice signal can be converted into a frequency-domain signal. At this time, the voice noise reduction apparatus can use the Fourier transform to analyze the first voice signal and convert it into the first spectrum.
[0035] In some embodiments, in order to more accurately convert the first voice signal into a frequency-domain signal, the above-mentioned determining the first spectrum of the first voice signal to be processed includes: performing frame addition and windowing processing on the first voice signal to obtain a plurality of voice frames; performing Fourier transform on the voice frames to obtain the first spectrum corresponding to each voice frame.
[0036] Under normal circumstances, a complete speech signal is usually non-stationary, that is, its characteristics basically change with time. However, due to the inherent characteristics of human oral and laryngeal vocalization, within a short time range, generally between 10 and 30 ms (milliseconds), its characteristics basically remain unchanged, that is, the speech signal is short-term stationary. Therefore, when performing frequency-domain analysis on the speech signal, it needs to be processed frame by frame, and the frame length is generally taken as 10 to 30 ms. Therefore, the first speech signal can be framed. To prevent spectral leakage when converting the first speech signal into a frequency-domain signal, the framed first speech signal can be windowed to obtain multiple speech frames. The Fourier transform is performed on each of the multiple speech frames respectively, that is, each speech frame is converted into a frequency-domain signal, so that the first spectrum corresponding to each speech frame can be obtained respectively. For the convenience of description, the first speech signal can be represented by S 1 (t), and the first spectrum signal can be represented by S 1 (k, λ), where k represents the frequency index and λ represents the frame index.
[0037] In some embodiments, when performing the Fourier transform on the speech frame, the short-time Fourier transform can be performed, or the fast Fourier transform can be performed. Of course, it can also be other types of Fourier transforms, and the embodiments of the present application do not limit this.
[0038] Step S102: Determine the first gain of the first spectrum based on the first noise reduction algorithm, and determine the second spectrum according to the first gain and the first spectrum.
[0039] In the implementation of the present application, after the speech noise reduction device converts the first speech signal S 1 (t) into the first spectrum S 1 (k, λ), the first noise reduction algorithm can be used to perform corresponding calculations on the first spectrum S 1 (k, λ) to obtain the first gain corresponding to each spectrum signal. Among them, the first noise reduction algorithm can be preset. For example, it can be a traditional noise reduction algorithm. After obtaining the first gain, the first spectrum can be noise-reduced according to the first gain to obtain the second spectrum.
[0040] In some embodiments, as Figure 2 shown, the determination of the first gain of the first spectrum based on the first noise reduction algorithm in the above step S102 includes:
[0041] Step S1021: Estimate the noise of the first spectrum.
[0042] In the embodiments of the present application, the speech noise reduction device can perform noise estimation on the first spectrum S 1Noise estimation is performed for (k, λ), where the voice noise reduction device can use algorithms such as the Minimum Statistics (MS) algorithm, the Minima Controlled Recursive Averaging (MCRA) algorithm, the Minima Controlled Recursive Averaging - 2 (MCRA - 2) algorithm, or the Improved Minima Controlled Recursive Averaging (IMCRA) algorithm to perform noise estimation.
[0043] Step S1022: Calculate the a priori signal - to - noise ratio based on the noise estimation.
[0044] The voice noise reduction device can calculate the a priori signal - to - noise ratio according to the noise estimation result. Among them, the a priori signal - to - noise ratio can be obtained by calculating the ratio between the clean speech signal and the noisy speech signal. That is, the a priori signal - to - noise ratio can be calculated from the result of the noise estimation and the non - noise speech signal in the first spectrum. The voice noise reduction device can calculate the a priori signal - to - noise ratio according to the noise estimation result by means of the decision - directed method, the improved decision - directed method, etc.
[0045] Step S1023: Determine the first gain according to the a priori signal - to - noise ratio using a preset gain algorithm.
[0046] The voice noise reduction device calculates the first gain using the a priori signal - to - noise ratio. The voice noise reduction device can calculate the first gain by methods such as the spectral subtraction method, the Wiener filter method, the MMSE (Minimum Mean Squared Error) method, etc. After the voice noise reduction device calculates the first gain, it can use the product of the first gain and the first spectrum as the second spectrum. That is, the second spectrum can be calculated by using the formula S 2 (k, λ) = S 1 (k, λ) * G1(k, λ), where S 2 (k, λ) represents the second spectrum, and G1(k, λ) represents the first gain.
[0047] Step S103: Determine the types of each speech frame included in the second spectrum according to the first gain and the second spectrum.
[0048] In the embodiments of the present application, voice frames can be divided into voiceless audio signals and voiced audio signals. The voice noise reduction device can perform different noise reduction processes on the voiceless audio signals and voiced audio signals to reduce the distortion degree of the voiceless components in the voice signal. Based on this, the voice noise reduction device can confirm the voiceless or voiced type of each voice frame in the second spectrum according to the first gain and the second spectrum.
[0049] In some embodiments, as Figure 3 shown, the above step S103 determines the types of each voice frame included in the second spectrum according to the first gain and the second spectrum, including:
[0050] Step S1031: For each voice frame included in the second spectrum, calculate the difference between the low-frequency energy value and the high-frequency energy value in the voice frame.
[0051] In the embodiments of the present application, it is necessary to determine the voiceless type and the voiced type of each voice frame in the second spectrum. Therefore, the following process needs to be executed for each voice frame included in the second spectrum. For the convenience of description, the following takes the process of determining the type of a voice frame as an example. Since the vocal cords do not vibrate when making a voiceless sound, the corresponding fundamental frequency of the vibration is zero, and the voiceless signal is a high-frequency signal. Therefore, the voice noise reduction device can calculate the difference between the low-frequency energy value and the high-frequency energy value in this voice frame.
[0052] In some embodiments, the voice noise reduction device can use the formula to calculate the difference between the low-frequency energy value and the high-frequency energy value in the voice frame. Wherein, a represents the initial index value of the low-frequency energy, b represents the initial index value of the high-frequency energy, and n represents the number of frequency domain points mapped by the high-low frequency energy coverage bandwidth.
[0053] Step S1032: Whether the difference between the low-frequency energy value and the high-frequency energy value in the voice frame is less than the first preset threshold.
[0054] That is, the voice noise reduction device can compare the difference between the low-frequency energy value and the high-frequency energy value in the voice frame with the first preset threshold to detect whether there is more high-frequency energy in the voice frame. If there is more, it means that the probability that this voice frame is a voiceless audio signal is greater. If not, it means that this voice frame is not a voiceless audio signal.
[0055] It should be understood that the first preset threshold can be preset according to actual requirements and is used to determine whether a speech frame is a voiced audio signal based on the difference between the low-frequency energy value and the high-frequency energy value in the speech frame. When the speech noise reduction device determines that the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than the first preset threshold, the following step S1033a is executed. When the speech noise reduction device determines that the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than the first preset threshold, the following step S1033b is executed.
[0056] Step S1033a: If the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than the first preset threshold, it is determined that the speech frame is a voiced audio signal.
[0057] That is, when the speech noise reduction device determines that the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than the first preset threshold, it indicates that there is more low-frequency energy in the speech frame, and the speech frame is not an unvoiced audio signal, so it is a voiced audio signal.
[0058] In some embodiments, to facilitate marking the type of speech frame, when it is determined that the speech frame is a voiced audio signal, the unvoiced flag (label) of the speech frame can be set to 0. At this time, in subsequent steps, the speech noise reduction device can determine the type of the speech frame by detecting the unvoiced flag value of the speech frame.
[0059] Step S1033b: If the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than the first preset threshold, the target value corresponding to the speech frame is determined based on the first gain.
[0060] Wherein, the target value is the sum value of the high-frequency gain or the sum value of the low-frequency gain corresponding to the speech frame. When the sum value of the high-frequency gain corresponding to the speech frame is greater than the sum value of the low-frequency gain, the target value is the sum value of the high-frequency gain corresponding to the speech frame. Or, when the sum value of the low-frequency gain corresponding to the speech frame is greater than the sum value of the high-frequency gain, the target value is the sum value of the low-frequency gain corresponding to the speech frame.
[0061] That is, when the speech noise reduction device determines that the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than the first preset threshold, it indicates that there is more high-frequency energy in the speech frame, and it may be an unvoiced audio signal, which needs further judgment. The speech noise reduction device calculates the maximum value of the sum of the high-frequency and low-frequency gains of the first gain. The speech noise reduction device can calculate the target value corresponding to the speech frame according to the formula where a represents the initial index value of the low-frequency energy, b represents the initial index value of the high-frequency energy, and n represents the number of frequency domain points mapped by the high-low frequency energy coverage bandwidth. represents the sum value of the first gain of the low frequency, Represents the sum value of the first gain at high frequencies. In this step, the reason for MAX is to prevent the first noise reduction algorithm from overly suppressing audio signals of the voiceless type and ensure that the target value is the larger value among the first gains in the high and low frequency bands.
[0062] Step S1034: Whether the target value corresponding to the speech frame is greater than the second preset threshold.
[0063] After calculating the target value corresponding to the speech frame, compare the target value corresponding to the speech frame with the second preset threshold to determine whether the speech frame is a voiceless audio signal.
[0064] It should be understood that the second preset threshold can be set in advance according to actual needs and is used as a threshold for determining the type of speech frame. If the target value corresponding to the speech frame is not greater than the second preset threshold, then execute the following step S1035b; if the target value corresponding to the speech frame is greater than the second preset threshold, then execute step S1035a.
[0065] Step S1035a: When the target value corresponding to the speech frame is greater than the second preset threshold, determine that the speech frame is a voiceless audio signal.
[0066] That is, when the target value corresponding to the speech frame is greater than the second preset value, it indicates that the first gain corresponding to this speech frame is larger, and this speech frame can be judged as a voiceless audio signal. At this time, the speech noise reduction device can set the voiceless flag value of this speech frame to 1.
[0067] Step S1035b: When the target value corresponding to the speech frame is not greater than the second preset threshold, determine that the speech frame is a voiced audio signal.
[0068] That is, when the target value corresponding to the speech frame is not greater than the second preset value, it indicates that the possibility of this speech noise exists greater, and there is no voiceless type audio signal, and this speech frame can be judged as a voiced audio signal. At this time, the speech noise reduction device can set the voiceless flag value of this speech frame to 0.
[0069] In this way, through the above method, the speech noise reduction device can determine the type of each speech frame in the second spectrum.
[0070] Step S104: Determine the second gain of the first spectrum based on the second noise reduction algorithm.
[0071] In the embodiments of the present application, the second noise reduction algorithm and the first noise reduction algorithm are two noise reduction algorithms with different emphases on noise reduction functions. For example, the first noise reduction algorithm can be a traditional noise reduction algorithm, and the second noise reduction algorithm can be an AI noise reduction algorithm. In order to more accurately perform noise reduction processing on the first speech signal, the speech noise reduction device can calculate the second gain of the first spectrum based on the second noise reduction algorithm.
[0072] In some embodiments, referring to as Figure 2 shown, the above step S104 of determining the second gain of the first spectrum based on the second noise reduction algorithm includes:
[0073] Step S1041: Input the first spectrum into the prediction network model to obtain a mask vector.
[0074] In the embodiments of the present application, the network model can be pre-trained to output a mask vector according to the input spectrum. When the prediction network model is pre-trained, the voice noise reduction device can use the first spectrum as the input of the prediction network model. The prediction network model performs feature analysis on the first spectrum to obtain a mask vector and outputs the mask vector.
[0075] It should be noted that the above prediction network model can include deep learning neural network models such as Deep Xi neural network, Convolutional Neural Network (CNN), and Long Short Term Memory (LSTM). The embodiments of the present application do not limit this.
[0076] Step S1042: Determine the mask vector as the second gain.
[0077] That is, after the voice noise reduction device obtains the mask vector output by the prediction network model, it can use the mask vector as the second gain. For convenience of description, the second gain can be represented by G2(k, λ).
[0078] It should be noted that the execution order between step S104 and steps S102 - S103 is not limited. Steps S102 - S103 can be executed first, and then step S104 can be executed. Or step S104 can be executed first, and then steps S102 - S103 can be executed. Or steps S102 - S103 and step S104 can be executed simultaneously, as Figure 4 shown. The illustration is only a schematic and not a specific limitation in the embodiments of the present application.
[0079] Step S105: Obtain the third gain of each speech frame according to the type, first gain, and second gain of each speech frame.
[0080] In the embodiments of the present application, for the speech frames of the voiceless type, in order to protect the voiceless audio signal and reduce the distortion of the voiceless audio signal, when denoising the speech frames of the voiceless type, the first gain and the second gain need to be combined. For the speech frames of the voiced type, in order to perform more accurate denoising, the second gain can be used for denoising. Based on this, the voice noise reduction device can determine the third gain of each speech frame according to the type, first gain, and second gain of each speech frame.
[0081] In some embodiments, as Figure 5 shown, the above step S105 obtains the third gain of each speech frame according to the type, the first gain, and the second gain of each speech frame, including: for each speech frame included in the second spectrum, when the speech frame is a voiceless audio signal, determining a fusion coefficient. Merging the first gain and the second gain based on the fusion coefficient to obtain the third gain of the speech frame. When the speech frame is a voiced audio signal, using the second gain as the third gain of the speech frame.
[0082] In the embodiments of the present application, for each speech frame included in the second spectrum, it is necessary to determine the third gain corresponding to each speech frame. For the convenience of description, the following takes determining the third gain of any speech frame as an example for illustration. The speech noise reduction device needs to determine the third gain according to the type of the speech frame. When the speech frame is a voiceless audio signal, a fusion coefficient can be determined. The fusion coefficient can be a value preset according to actual needs. The fusion coefficient is usually a value greater than 0 and less than 1. In some embodiments, the fusion coefficient can be 0.8. After determining the fusion coefficient, the speech noise reduction device can perform a fusion process on the first gain and the second gain according to the fusion coefficient to obtain the third gain of the speech frame. In some embodiments, the speech noise reduction device can calculate the third gain of the speech frame according to the formula G3(k, λ) = G1(k, λ) * β + G2(k, λ) * (1 - β). Where β represents the fusion coefficient, G1(k, λ) is the first gain of the speech frame, and G2(k, λ) is the second gain of the speech frame.
[0083] When the speech frame is a voiced audio signal, the second gain can be directly used to perform noise reduction on it to improve the noise reduction effect. At this time, the speech noise reduction device can determine the second gain of the speech frame as the third gain, that is, G3(k, λ) = G2(k, λ).
[0084] Step S106: Determine a second speech signal based on the third gain of each speech frame and the second spectrum.
[0085] The second speech signal is the speech signal after the first speech signal is noise-reduced.
[0086] That is, after the speech noise reduction device determines the third gain of each speech frame, it can perform noise reduction processing on each speech frame in the second spectrum to obtain the noise-reduced speech signal, which is the second speech signal.
[0087] In some embodiments, determining the second speech signal based on the third gain of each speech frame and the second spectrum includes:
[0088] Multiply each speech frame in the second spectrum by the corresponding third gain to obtain a third spectrum; perform overlap-and-add processing and inverse Fourier transform on the third spectrum to obtain a second speech signal.
[0089] That is, after the voice noise reduction device determines the third gain of each speech frame, it can multiply each speech frame in the second spectrum by the corresponding third gain to perform noise reduction processing on each speech frame in the second spectrum, and obtain a noise-reduced third spectrum. For example, it can be obtained through the formula S 3 (k, λ) = S 2 (k, λ) * G3(k, λ) to obtain the third spectrum. Since the third spectrum is a frequency-domain signal and needs to be converted into a time-domain signal, at this time, in order to accurately convert the third spectrum signal into a time-domain signal, the voice noise reduction device can perform overlap-and-add processing on the third spectrum and then perform inverse Fourier transform to obtain the second speech signal.
[0090] In this way, in the embodiment of the present application, for a speech signal of the voiceless type, the first gain and the second gain can be fused to obtain a third gain, and the second spectrum can be noise-reduced through the third gain. The protection strength of the voiceless type speech signal can be controlled by adjusting the smoothing parameter, greatly reducing the distortion degree of the voiceless type speech signal and improving the protection of the voiceless type speech signal. And for the voiced type speech signal, the second gain is used to perform noise reduction processing on the second spectrum, which can improve the noise reduction effect and the accuracy of noise reduction processing for non-stationary noise. In this way, through the combined noise reduction processing of the first noise reduction algorithm and the second noise reduction algorithm, the removal effect of the voice noise component can be made more significant, and the distortion degree of the voiceless component can be reduced. The quality of the noise-reduced speech signal is higher, improving the user experience.
[0091] See Figure 6 for a schematic structural diagram of a voice noise reduction device provided by an embodiment of the present application. As Figure 6 shown, the voice noise reduction device includes:
[0092] A determination unit 601, configured to determine the first spectrum of the first speech signal to be processed;
[0093] The determination unit 601 is further configured to determine the first gain of the first spectrum based on the first noise reduction algorithm, and determine the second spectrum according to the first gain and the first spectrum.
[0094] The determination unit 601 is further configured to determine the frequency band type of each speech frame included in the second spectrum according to the first gain and the second spectrum.
[0095] The determination unit 601 is further configured to determine the second gain of the first spectrum based on the second noise reduction algorithm.
[0096] A processing unit 602, configured to merge the first gain and the second gain according to the frequency band types of the respective speech frames to obtain a third gain for each speech frame.
[0097] The processing unit 602 is further configured to determine a second speech signal based on the third gain and the second spectrum of each speech frame.
[0098] Wherein, the second speech signal is the speech signal after noise reduction of the first speech signal.
[0099] As a possible implementation manner, the determining unit 601 is specifically configured to perform frame segmentation and windowing processing on the first speech signal to obtain a plurality of speech frames; perform Fourier transform on the speech frames to obtain a first spectrum corresponding to each speech frame.
[0100] As a possible implementation manner, the determining unit 601 is specifically configured to perform noise estimation on the first spectrum; calculate a priori signal-to-noise ratio according to the noise estimation result; determine the first gain by using a preset gain algorithm according to the a priori signal-to-noise ratio.
[0101] As a possible implementation manner, the determining unit 601 is specifically configured to calculate the difference between the low-frequency energy value and the high-frequency energy value in each speech frame included in the second spectrum; if the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than a first preset threshold, determine that the speech frame is a voiced audio signal; if the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than the first preset threshold, determine a target value corresponding to the speech frame based on the first gain; the target value is the sum value of the high-frequency gain or the sum value of the low-frequency gain corresponding to the speech frame, and when the sum value of the high-frequency gain corresponding to the speech frame is greater than the sum value of the low-frequency gain, the target gain value is the sum value of the high-frequency gain corresponding to the speech frame, or, when the sum value of the low-frequency gain corresponding to the speech frame is greater than the sum value of the high-frequency gain, the target gain value is the sum value of the low-frequency gain corresponding to the speech frame; if the target value corresponding to the speech frame is greater than a second preset threshold, determine that the speech frame is an unvoiced audio signal; if the target value corresponding to the speech frame is not greater than the second preset threshold, determine that the speech frame is a voiced audio signal.
[0102] As a possible implementation manner, the processing unit 602 is specifically configured to, for each speech frame included in the second spectrum, determine a fusion coefficient when the speech frame is an unvoiced audio signal; merge the first gain and the second gain based on the fusion coefficient to obtain a third gain of the speech frame; when the speech frame is a voiced audio signal, use the second gain as the third gain of the speech frame.
[0103] As a possible implementation manner, the determining unit 601 is specifically configured to input the first spectrum into a prediction network model to obtain a mask vector; determine the mask vector as the second gain.
[0104] As a possible implementation, the processing unit 602 is specifically configured to multiply each speech frame in the second spectrum by the corresponding third gain to obtain a third spectrum; and perform overlap-and-add processing and inverse Fourier transform on the third spectrum to obtain a second speech signal.
[0105] Corresponding to the above embodiments, the present application further provides an electronic device. Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 700 may include: a processor 701, a memory 702, and a communication unit 703. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation to the embodiments of the present invention. It may be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0106] Among them, the communication unit 703 is used to establish a communication channel so that the electronic device can communicate with other devices. Receive user data sent by other devices or send user data to other devices.
[0107] The processor 701 is the control center of the electronic device. It connects various parts of the entire electronic device through various interfaces and lines. By running or executing software programs, instructions, and / or modules stored in the memory 702, and by calling data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC). For example, it may be composed of a single packaged IC, or may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 701 may only include a central processing unit (CPU). In the embodiment of the present invention, the CPU may be a single operation core or may include multiple operation cores.
[0108] The memory 702 is used to store the execution instructions of the processor 701. The memory 702 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0109] When the execution instructions in the memory 702 are executed by the processor 701, the electronic device 700 is enabled to execute some or all of the steps in the above-described embodiments.
[0110] In a specific implementation, the present invention further provides a computer storage medium. The computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in the embodiments of the voice noise reduction method provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.
[0111] In a specific implementation, the present invention further provides a computer program product. The computer program product includes executable instructions. When the executable instructions are executed on a computer, the computer is caused to execute some or all of the steps in the embodiments of the simulation scenario generation method provided by the present invention.
[0112] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0113] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments and the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.
Claims
1. A speech noise reduction method, characterized in that: include: Determining a first frequency spectrum of a first speech signal to be processed; Determine a first gain of the first spectrum based on a first noise reduction algorithm, and determine a second spectrum according to the first gain and the first spectrum; Determining the type of each speech frame included in the second spectrum according to the first gain and the second spectrum; determining a second gain of the first spectrum based on a second noise reduction algorithm; Obtaining a third gain for each speech frame according to the type of each speech frame, the first gain, and the second gain; Based on the third gain of each speech frame and the second spectrum, a second speech signal is determined; the second speech signal is a speech signal after noise reduction of the first speech signal.
2. The method according to claim 1, characterized in that Determining a first frequency spectrum of a first speech signal to be processed comprises: Performing frame division and windowing processing on the first speech signal to obtain a plurality of speech frames; Perform Fourier transform on the speech frames to obtain a first spectrum corresponding to each speech frame.
3. The method according to claim 1, characterized in that Determining the first gain of the first spectrum based on the first noise reduction algorithm includes: performing noise estimation on the first spectrum; According to the noise estimation result, the priori signal-to-noise ratio is calculated; The first gain is determined by using a preset gain algorithm according to the priori signal-to-noise ratio.
4. The method according to claim 1, characterized in that: The determining, according to the first gain and the second spectrum, the type of each speech frame included in the second spectrum comprises: For each speech frame included in the second spectrum, calculating the difference between the low-frequency energy value and the high-frequency energy value in the speech frame; If the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is less than a first preset threshold, then determining that the speech frame is a voiced audio signal; If the difference between the low-frequency energy value and the high-frequency energy value in the speech frame is not less than a first preset threshold, a target value corresponding to the speech frame is determined based on the first gain; the target value is a high-frequency gain sum value or a low-frequency gain sum value corresponding to the speech frame, and when the high-frequency gain sum value corresponding to the speech frame is greater than the low-frequency gain sum value, the target gain value is the high-frequency gain sum value corresponding to the speech frame, or when the low-frequency gain sum value corresponding to the speech frame is greater than the high-frequency gain sum value, the target gain value is the low-frequency gain sum value corresponding to the speech frame; When the target value corresponding to the speech frame is greater than a second preset threshold, determining that the speech frame is an unvoiced audio signal; When the target value corresponding to the speech frame is not greater than the second preset threshold, it is determined that the speech frame is a voiced audio signal.
5. The method according to claim 4, characterized in that The step of obtaining a third gain for each speech frame according to the type of each speech frame, the first gain, and the second gain includes: For each speech frame included in the second spectrum, when the speech frame is an unvoiced audio signal, determining a fusion coefficient; Combining the first gain and the second gain based on the fusion coefficient to obtain a third gain of the speech frame; When the speech frame is a voiced audio signal, the second gain is used as the third gain of the speech frame.
6. The method according to claim 1, characterized in that Determining the second gain of the first spectrum based on the second noise reduction algorithm includes: Inputting the first spectrum into a prediction network model to obtain a mask vector; The mask vector is determined as the second gain.
7. The method according to claim 1, characterized in that The determining the second speech signal based on the third gain of each speech frame and the second spectrum comprises: multiplying each speech frame in the second spectrum by the corresponding third gain to obtain a third spectrum; The second speech signal is obtained by performing a concatenated addition process and an inverse Fourier transform on the third spectrum.
8. A speech noise reduction device, characterized in that: include: A determination unit, configured to determine a first frequency spectrum of a first speech signal to be processed; The determining unit is further configured to determine a first gain of the first spectrum based on a first noise reduction algorithm, and determine a second spectrum according to the first gain and the first spectrum; The determining unit is further configured to determine a frequency band type of each speech frame included in the second spectrum according to the first gain and the second spectrum; The determining unit is further configured to determine a second gain of the first spectrum based on a second noise reduction algorithm; a processing unit, configured to combine the first gain and the second gain according to the frequency band type of each voice frame to obtain a third gain for each voice frame; The processing unit is further used to determine a second speech signal based on the third gain of each speech frame and the second spectrum; the second speech signal is a speech signal after noise reduction of the first speech signal.
9. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.