Audio noise reduction model training method, system, device and medium
By performing feature normalization and root mean square normalization of the input of the audio noise reduction model, the problems of instability in training and uneven noise reduction effects in the existing technology are solved, and a more efficient and balanced noise reduction effect is achieved.
Patent Information
- Application Number
- CN202510128877.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing audio noise reduction model training methods are used to process the amplitude spectrum, there are problems such as training instability, slow convergence speed and uneven noise reduction effect. Especially when the noise is high or the speech is weak, the effect is not ideal.
By preprocessing and time-frequency analysis of the initial speech information, the amplitude spectrum information is obtained, and the feature normalization is performed to make different features have similar scales and ranges. At the same time, before calculating the loss function, the output and label of the model are normalized to ensure that the contribution of speech of different energies to the loss is close.
It improves the convergence speed and training stability of the model, enhances the balance of the noise reduction effect, can better cope with various complex environmental noises, and improves the overall noise reduction performance.
Smart Images

Figure CN120071955A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio noise reduction, and in particular, to a method, system, device, and medium for training an audio noise reduction model. Background Art
[0002] With the popularization of intelligent devices such as smart phones, wireless earphones, and voice assistants, clear and clean voice signals have become one of the core requirements for improving the user experience. When conducting voice communication in a noisy environment, the interference of background noise seriously affects the quality and recognition rate of the voice. Therefore, it is necessary to separate clear voice signals from noise. By training a deep neural network to learn to remove noise from noisy speech, the AI noise reduction model can effectively improve the voice quality, so it is widely used in scenarios such as speech recognition, audio processing, and voice communication.
[0003] The current noise reduction model adopts a training method based on the magnitude spectrum. The input of the model is usually the magnitude spectrum of noisy speech. After forward inference of the neural network, a predicted magnitude spectrum is obtained. The model uses the mean square error (MSE) as the loss function, and the magnitude spectrum of clean speech as the label. By calculating the loss and performing backpropagation, the network parameters are updated. This training method improves the noise reduction effect to a certain extent and can handle different types of noise.
[0004] However, the numerical values of the magnitude spectrum features in the current training method usually have a large range, resulting in instability during the training process, and the convergence speed of the model is slow. During the training process, due to the different magnitudes of the magnitude spectra in different frequency bands, it is difficult for the model to achieve balanced learning across the entire frequency spectrum, thus affecting the overall training efficiency and effect. Moreover, when the mean square error loss function processes the magnitude spectrum, the parts with larger energy contribute more to the loss, while the parts with smaller energy are more sensitive, making the model not learn the details of the low-energy parts sufficiently during the training process. The dominant role of the high-energy frequency bands in the loss causes the model to overly focus on the high-energy parts when updating the parameters and ignore the details of the low-energy parts, resulting in unbalanced noise reduction effects. Therefore, under the influence of the existing training method, in the case of large noise or weak speech, the noise reduction effect often fails to reach the expected level, and the above problems need to be solved. Summary of the Invention
[0005] In order to enable the noise reduction system to have higher adaptability and performance, cope with various complex environmental noises, and improve the noise reduction effect, this application provides a method, system, device, and medium for training an audio noise reduction model, and adopts the following technical solutions:
[0006] In a first aspect, this application provides a method for training an audio noise reduction model, including:
[0007] Obtain initial voice information, and perform preprocessing on the initial voice information to obtain preprocessed voice information;
[0008] Perform time-frequency analysis on the preprocessed speech information to obtain amplitude spectrum information;
[0009] Perform feature normalization on the amplitude spectrum information to obtain feature normalization scaling value information;
[0010] Input the feature normalization scaling value information into the noise reduction model to be trained for forward inference to obtain amplitude spectrum mask information, and perform noise reduction analysis based on the amplitude spectrum mask information and the amplitude spectrum information to obtain noise reduction amplitude spectrum information;
[0011] Perform root mean square normalization on the initial speech information to obtain root mean square normalization scaling parameters;
[0012] Obtain label information, and scale the noise reduction amplitude spectrum information and the label information respectively according to the root mean square normalization scaling value information to obtain root mean square normalization result information.
[0013] Preferably, it further includes:
[0014] Backpropagate the root mean square normalization result information to the noise reduction model to update the model weights.
[0015] Preferably, the specific steps of performing time-frequency analysis on the preprocessed speech information to obtain amplitude spectrum information are:
[0016] Perform short-time Fourier transform on the preprocessed speech information to obtain amplitude information;
[0017] Perform amplitude spectrum conversion on the amplitude information to obtain amplitude spectrum information.
[0018] Preferably, the specific steps of performing feature normalization on the amplitude spectrum information to obtain feature normalization scaling value information are:
[0019] Obtain sliding window length information and selected time information, and perform forgetting ratio analysis according to the sliding window length information and the selected time information to obtain forgetting ratio coefficient information;
[0020] Perform scaling coefficient analysis on the forgetting ratio coefficient information to obtain feature normalization scaling coefficients;
[0021] Perform feature normalization scaling on the amplitude spectrum information according to the feature normalization scaling coefficients to obtain feature normalization scaling value information.
[0022] Preferably, the feature normalization scaling coefficient is:
[0023]
[0024] where t is the selected time information, L is the sliding window length information, and α tis the forgetting ratio coefficient information, μ t-1 is the normalized scaling coefficient at the previous moment.
[0025] Preferably, the root mean square normalized scaling parameter is:
[0026]
[0027] where x is the initial voice information and N is the length.
[0028] In a second aspect, the present application provides an audio noise reduction model training system, including:
[0029] A preprocessing module for obtaining initial voice information, preprocessing the initial voice information, and obtaining preprocessed voice information;
[0030] A time-frequency analysis module for performing time-frequency analysis on the preprocessed voice information to obtain amplitude spectrum information;
[0031] A feature normalization processing module for performing feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information;
[0032] A noise reduction analysis module for inputting the feature normalization scaling value information into the noise reduction model to be trained for forward inference, obtaining amplitude spectrum mask information, and performing noise reduction analysis based on the amplitude spectrum mask information and the amplitude spectrum information to obtain noise reduction amplitude spectrum information;
[0033] A root mean square normalization processing module for performing root mean square normalization processing on the initial voice information to obtain a root mean square normalization scaling parameter;
[0034] A scaling processing module for obtaining label information, and scaling the noise reduction amplitude spectrum information and the label information respectively according to the root mean square normalization scaling value information to obtain root mean square normalization result information.
[0035] Preferably, it further includes:
[0036] A model update module for backpropagating the root mean square normalization result information to the noise reduction model to update the model weights.
[0037] In a third aspect, the present application provides an audio noise reduction model training device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the audio noise reduction model training method as described above.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the audio noise reduction model training method as described above when running.
[0039] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by this application at least include:
[0040] In this application, the initial voice information is first preprocessed, the obtained preprocessed voice information is subjected to time-frequency analysis to obtain amplitude spectrum information, and a feature normalization operation that can meet the requirements of real-time inference is performed on the amplitude spectrum features input to the model, so that different features have similar scales and ranges, making it easier for the model to be optimized and converge. Through the forward inference of the noise reduction model, amplitude spectrum mask information is obtained, and then the noise reduction amplitude spectrum information is obtained through noise reduction analysis. And before calculating the loss function, root mean square normalization is also performed on the output and labels of the model at the same time to obtain the root mean square normalization scaling parameter, so that when calculating the MSE loss, the system is not overly sensitive to voices with different energies, and the contributions of voices with different energies to the loss are relatively close, reducing the occurrence of parameter bias when updating model parameters, which helps the model to converge, so that the noise reduction system can have higher adaptability and performance, cope with various complex environmental noises, and improve the noise reduction effect. Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of an active noise control method described in an embodiment of this application.
[0042] Figure 2 It is a schematic structural diagram of an active noise control system described in an embodiment of this application.
[0043] Description of the Reference Numerals:
[0044] 1. Preprocessing module; 2. Time-frequency analysis module; 3. Feature normalization processing module; 4. Noise reduction analysis module; 5. Root mean square normalization processing module; 6. Scaling processing module. Detailed Embodiments
[0045] The following will be combined with Figure 1 - Figure 2 This application is further described in detail. The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to be limiting.
[0046] Referring to Figure 1 , an audio noise reduction model training method involved in this application specifically includes:
[0047] Step S1: Obtain the initial voice information, and preprocess the initial voice information to obtain the preprocessed voice information;
[0048] Step S2: Perform time-frequency analysis on the preprocessed voice information to obtain amplitude spectrum information;
[0049] Step S3: Perform feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information;
[0050] Step S4: Input the feature normalization scaling value information into the denoising model to be trained for forward inference to obtain the magnitude spectrum mask information, and perform denoising analysis based on the magnitude spectrum mask information and the magnitude spectrum information to obtain the denoised magnitude spectrum information;
[0051] Step S5: Perform root mean square normalization processing on the initial speech information to obtain the root mean square normalization scaling parameter;
[0052] Step S6: Obtain the label information, and scale the denoised magnitude spectrum information and the label information respectively according to the root mean square normalization scaling value information to obtain the root mean square normalization result information.
[0053] Specifically, in the process of training the AI model in the current technology, the range of magnitude spectrum features is not considered to be very large, resulting in unstable model training and slow convergence. At the same time, when calculating the mean square error loss, the part with large energy contributes more to the loss, resulting in insensitivity to the part with small energy when updating the parameters by backpropagation. As a result, the existing AI denoising model based on the magnitude spectrum cannot provide a satisfactory denoising effect.
[0054] In this application, the initial speech information is preprocessed first, the obtained preprocessed speech information is subjected to time-frequency analysis to obtain the magnitude spectrum information, and the magnitude spectrum features input to the model are subjected to feature normalization operations that can meet the requirements of real-time inference, so that different features have similar scales and ranges, making it easier for the model to be optimized and converge. Through forward inference of the denoising model, the magnitude spectrum mask information is obtained, and thus the denoised magnitude spectrum information is obtained through denoising analysis. And before calculating the loss function, the output and label of the model are also simultaneously subjected to root mean square normalization to obtain the root mean square normalization scaling parameter, so that when calculating the MSE loss, the speech with different energies will not be overly sensitive, and the contributions of the speech with different energies to the loss are relatively close, reducing the parameter bias when updating the model parameters, which helps the model to converge, so that the denoising system can have higher adaptability and performance, cope with various complex environmental noises, and improve the denoising effect.
[0055] As one of the implementation manners, obtaining the initial speech information and preprocessing the initial speech information to obtain the preprocessed speech information specifically means first collecting the original speech signal from a speech source such as a microphone or a recording device, and then performing some preprocessing operations on the original speech signal, such as removing the DC offset, adjusting the sampling rate, etc., so that the signal is within an appropriate range, and dividing the speech signal into short-time frames for subsequent processing.
[0056] As one of the implementation manners, the specific steps for performing time-frequency analysis on the preprocessed speech information to obtain the magnitude spectrum information are:
[0057] Perform short-time Fourier transform on the preprocessed speech information to obtain the magnitude information;
[0058] Perform amplitude spectrum conversion on the amplitude information to obtain amplitude spectrum information.
[0059] Specifically, the embodiment of the present application uses STFT, that is, the short-time Fourier transform method to process the preprocessed speech information to obtain amplitude information. STFT can transform a signal from the time domain to the frequency domain, and the processed result is a complex number. Since a complex number can also be represented by amplitude and phase, the amplitude and phase information of the signal at a certain time period and frequency can be represented. The signal is decomposed into the amplitude and phase information of different frequency components, and further forms an amplitude spectrum and a phase spectrum. The specific amplitude spectrum is:
[0060]
[0061] where Re represents the real part of the complex number, and Im represents the imaginary part of the complex number.
[0062] As one of the implementation manners, the specific steps for performing feature normalization processing on the amplitude spectrum information to obtain the feature normalization scaling value information are as follows:
[0063] Obtain the sliding window length information and the selected time information, perform forgetting ratio analysis according to the sliding window length information and the selected time information to obtain the forgetting ratio coefficient information;
[0064] Perform scaling coefficient analysis on the forgetting ratio coefficient information to obtain the feature normalization scaling coefficient;
[0065] Perform feature normalization scaling processing on the amplitude spectrum information according to the feature normalization scaling coefficient to obtain the feature normalization scaling value information.
[0066] Specifically, before the forward inference of the model, the embodiment of the present application performs forgetting normalization operations on the features input to the model to meet the requirements of real-time inference, so that different features have similar scales and ranges, making it easier for the model to be optimized and converge. During the training process, by using the forgetting normalization operation that meets the real-time processing requirements of the noise reduction system, the input amplitude spectrum features are normalized to balance the range of the value domain, ensuring that the size of the feature values input to the model is within a relatively small range, and the difference between values is not too large.
[0067] After the STFT processing in the embodiments of the present application, the frequency-domain amplitude spectrum of the noisy speech is extracted, and the size of the non-DC part is FFT_n / 2, where FFT_n represents the number of points of the Fourier transform of the STFT. In the embodiments of the present application, we preferably use 512 as this value, and then use the non-DC part as the input of the model. Before inputting the amplitude spectrum into the Encoder, it is normalized. Since the speech noise reduction process needs to be carried out in real time, forgetting normalization is used instead of using the common global features that rely on future information for normalization processing.
[0068] The forgetting normalization based on the sliding window uses the neighboring frames within a certain time range of the current frame to calculate its average value for normalization, where the time range is denoted as L. Since future frame information is not used, it meets the real-time requirement of AI noise reduction, and the proportional coefficient α is used to control the proportion of the current frame and the average value of the neighboring frames. The amplitude features after normalization consider the influence of other frames within a certain time range in the past, making the eigenvalue ranges of the front and back frames related, which not only makes the normalization effect better but also enables real-time processing. In the embodiments of the present application, it is preferred that L is 500 frames, and the proportional coefficient α is calculated through L. This method is used to process the input features during both the training and inference processes.
[0069] As one of the implementation manners, the feature normalization scaling coefficient is:
[0070]
[0071] where t is the selected time information, L is the sliding window length information, α t is the forgetting proportional coefficient information, and μ t-1 is the normalization scaling coefficient of the previous moment.
[0072] Specifically, in the embodiments of the present application, forgetting normalization is performed based on the sliding window, mainly for calculating the normalization scaling coefficient μ t and the forgetting proportional coefficient α t . The length of the sliding window is set as L. When the time t is less than L, the forgetting proportional coefficient is calculated When the time t is not less than L, α t is fixed as
[0073] where μ t-1 represents the normalization scaling coefficient of the previous moment, and mean(X t ′ ) represents taking the mean value of the amplitude spectrum X at time t ′ t . represents scaling the amplitude spectrum X at time t ′ t .
[0074] The amplitude spectrum AI noise reduction model uses the amplitude spectrum as the input, and performs a normalization operation on the features using forgetting normalization based on a sliding window before the amplitude spectrum features enter the model. The forgetting normalization adopts the operation of a sliding window and does not use future frame information, meeting the real-time processing requirements of the AI noise reduction system, and can perform noise reduction processing on speech online. By adjusting the weight ratio of the current frame and other frames within the sliding window through a scale factor, considering the changing factors of speech over time, the normalization of the current frame uses the results of the feature values of past frames within a certain range, making the entire normalization operation smoother and more effective, avoiding the weight bias of the input features, enabling the contribution of each feature to the model to be relatively balanced, and reducing the deviation between features.
[0075] As one of the implementation manners, the feature normalization scaling value information is input into the noise reduction model to be trained for forward inference to obtain amplitude spectrum mask information, and noise reduction analysis is performed based on the amplitude spectrum mask information and the amplitude spectrum information to obtain noise reduction amplitude spectrum information. Specifically, the model performs forward inference to obtain an amplitude spectrum mask, and applies the amplitude spectrum mask to the original amplitude spectrum to obtain the noise-reduced amplitude spectrum.
[0076] Specifically, the original amplitude spectrum in the embodiment of the present application is the result obtained through the above complex number conversion. The amplitude spectrum mask predicted by the model is a matrix with the same shape as the original amplitude spectrum, and its value ranges from 0 to 1. The mask is close to 0 where the model determines it to be noise, and close to 1 where it determines it to be human voice. Applying the amplitude spectrum mask to the original amplitude spectrum is to multiply the two. Through such multiplication, the human voice in the original amplitude can be retained, the noise can be suppressed, and the noise-reduced amplitude spectrum can be obtained.
[0077] The process of noise reduction by the amplitude spectrum model is as follows: first, the complex representation of the noisy speech is obtained through STFT, and then it is converted into an amplitude spectrum and a phase spectrum; the model uses the amplitude spectrum as the input to obtain an amplitude spectrum mask, and then multiplies them to obtain the noise-reduced amplitude spectrum; the noise-reduced amplitude spectrum and the original phase spectrum are converted back into a complex form representation, and then through iSTFT transformation into a time-domain signal to obtain the noise-reduced speech, that is, the noise reduction amplitude spectrum information.
[0078] As one of the implementation manners, the root mean square normalization scaling parameter is:
[0079]
[0080] where x is the initial speech information and N is the length.
[0081] Specifically, in the embodiment of the present application, before calculating the MSE loss using the noise-reduced amplitude spectrum and the clean amplitude spectrum, RMS normalization is performed on both of them simultaneously to reduce the influence of the energy magnitude on the result of the loss function.
[0082] When using backpropagation to update and optimize model parameters, it is necessary to first calculate the MSE loss function. In the embodiments of the present application, the RMS is calculated using the original noisy speech, and the prediction results and labels are normalized to balance the contributions of results with different energy levels to the loss function. When calculating the loss function, real-time performance does not need to be considered, so the global speech at the sentence level is used to calculate the RMS value. The preferred speech length is 15 seconds. First, the RMS is calculated using the complete noisy speech, and then the prediction results and labels are divided by this value to obtain volume-balanced prediction results and labels. Finally, the mean squared error loss is calculated for backpropagation to update the model parameters, specifically as follows:
[0083]
[0084] In the embodiments of the present application, first calculate to obtain the normalized scaling parameter RMS noisy , where x represents the noisy speech with a length of N. Then, through and scale the magnitude spectrum Mag 预测 after noise reduction and the magnitude spectrum Mag 标签 of the clean speech.
[0085] After obtaining the prediction results through forward inference of the magnitude spectrum model, the embodiments of the present application need to calculate the loss function for backpropagation to update the model parameters. First, the prediction results and labels are normalized using the root mean square, and then the loss function is calculated. The normalization operation can ensure that data with different energy levels have a comparable impact on the results of the loss function. It can avoid the parameter optimization direction tending to data with greater contributions when updating the model parameters through backpropagation. It can improve the model robustness and the effect of small-volume data.
[0086] The embodiments of the present application provide an AI speech noise reduction system based on the magnitude spectrum, including an acquisition unit, a calculation unit, a data transmission unit, and a terminal unit.
[0087] Among them, the acquisition unit consists of 1 microphone and 1 ADC hardware chip, and is used to convert the analog speech signal in the environment into a digital signal; the calculation unit consists of 1 single-chip microcomputer or a computing chip with an operating system, and is used for the calculation of the noise reduction model; the data transmission unit consists of 1 network system capable of transmitting data, and is used to transmit the calculated data; the terminal unit consists of any real-time conference communication device with network access, and is used to play the processed audio data.
[0088] Referring to Figure 2 , the embodiments of the present application provide an audio noise reduction model training system, including:
[0089] A preprocessing module, configured to obtain initial voice information, preprocess the initial voice information, and obtain preprocessed voice information;
[0090] A time-frequency analysis module, configured to perform time-frequency analysis on the preprocessed voice information to obtain amplitude spectrum information;
[0091] A feature normalization processing module, configured to perform feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information;
[0092] A noise reduction analysis module, configured to input the feature normalization scaling value information into a noise reduction model to be trained for forward inference, obtain amplitude spectrum mask information, and perform noise reduction analysis based on the amplitude spectrum mask information and the amplitude spectrum information to obtain noise reduction amplitude spectrum information;
[0093] A root mean square normalization processing module, configured to perform root mean square normalization processing on the initial voice information to obtain a root mean square normalization scaling parameter;
[0094] A scaling processing module, configured to obtain label information, and scale the noise reduction amplitude spectrum information and the label information respectively according to the root mean square normalization scaling value information to obtain root mean square normalization result information.
[0095] As one of the implementation manners, it further includes:
[0096] A model update module, configured to backpropagate the root mean square normalization result information to the noise reduction model to update the model weights.
[0097] An audio noise reduction model training device provided by an embodiment of the present application includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the audio noise reduction model training method as described above.
[0098] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, wherein the computer program is configured to execute the audio noise reduction model training method as described above when running.
[0099] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and products described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0100] In several embodiments provided by the present application, it should be understood that the disclosed methods, systems, devices, and program products can be implemented in other ways.
[0101] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0102] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for training an audio noise reduction model, characterized in that: include: Acquire initial voice information, pre-process the initial voice information, and obtain pre-processed voice information; Perform time-frequency analysis on the preprocessed speech information to obtain amplitude spectrum information; Performing feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information; Inputting the feature normalization scaling value information into the denoising model to be trained for forward reasoning to obtain the amplitude spectrum mask information, performing denoising analysis based on the amplitude spectrum mask information and the amplitude spectrum information to obtain the denoised amplitude spectrum information; Performing root mean square normalization processing on the initial speech information to obtain a root mean square normalization scaling parameter; The label information is obtained, and the denoised amplitude spectrum information and the label information are scaled according to the root mean square normalization scaling value information to obtain the root mean square normalization result information.
2. The audio noise reduction model training method according to claim 1, characterized in that: Also includes: The RMS normalized result information is back-propagated to the denoising model to update the model weights.
3. The audio noise reduction model training method according to claim 1, characterized in that: The specific steps of performing time-frequency analysis on the pre-processed speech information to obtain the amplitude spectrum information are: Perform short-time Fourier transform on the preprocessed speech information to obtain amplitude information; The amplitude information is converted into an amplitude spectrum to obtain the amplitude spectrum information.
4. The audio noise reduction model training method according to claim 3, characterized in that: The specific steps of performing feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information are: Obtaining sliding window length information and selected time information, performing forgetting ratio analysis according to the sliding window length information and the selected time information, and obtaining forgetting ratio coefficient information; Perform scaling factor analysis on the forgotten ratio coefficient information to obtain feature normalization scaling factor; The amplitude spectrum information is subjected to feature normalization scaling processing according to the feature normalization scaling coefficient to obtain feature normalization scaling value information.
5. The audio noise reduction model training method according to claim 4, characterized in that: The feature normalization scaling factor is: Among them, t is the selected time information, L is the sliding window length information, α t is the forgetting ratio coefficient information, μ t-1 is the normalized scaling factor of the previous moment.
6. The audio noise reduction model training method according to claim 1, characterized in that: The RMS normalization scaling parameter is: Among them, x is the initial voice information, and N is the length.
7. An audio noise reduction model training system, characterized in that: include: A preprocessing module, used to obtain initial voice information, preprocess the initial voice information, and obtain preprocessed voice information; A time-frequency analysis module is used to perform time-frequency analysis on the pre-processed speech information to obtain amplitude spectrum information; A feature normalization processing module is used to perform feature normalization processing on the amplitude spectrum information to obtain feature normalization scaling value information; A denoising analysis module is used to input the feature normalization scaling value information into the denoising model to be trained for forward reasoning to obtain the amplitude spectrum mask information, and to perform denoising analysis based on the amplitude spectrum mask information and the amplitude spectrum information to obtain denoised amplitude spectrum information; A root mean square normalization processing module, used for performing root mean square normalization processing on the initial voice information to obtain a root mean square normalization scaling parameter; The scaling processing module is used to obtain the label information, and scale the noise reduction amplitude spectrum information and the label information respectively according to the root mean square normalization scaling value information to obtain the root mean square normalization result information.
8. The audio noise reduction model training system according to claim 7, characterized in that: Also includes: The model updating module is used to back-propagate the RMS normalization result information to the denoising model to update the model weights.
9. An audio noise reduction model training device, characterized in that: It comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the audio noise reduction model training method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the audio noise reduction model training method according to any one of claims 1 to 6 when running.
Citation Information
Cited By
Segmented energy weighted noisy speech generation method and system, medium and equipment
CN120895018A