Speech enhancement model training method and device, electronic equipment and storage medium
By performing feature extraction and noise reduction processing on speech data, and adjusting model parameters in combination with loss function, the problem of poor processing of weak noise background in traditional speech noise reduction models is solved, and efficient speech enhancement effect is achieved.
Patent Information
- Application Number
- CN202411907760.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The traditional voice noise reduction model has huge computing requirements, and it cannot efficiently eliminate noise for weak noise backgrounds, which affects the model accuracy and leads to poor voice enhancement effect.
By collecting audio data containing ambient noise and audio data without noise, extracting features and using a preset model for noise reduction processing, calculating the loss function between the noise reduction signals, and adjusting the model parameters to generate a speech enhancement model.
The noise reduction process is simplified, the noise reduction efficiency is improved, and the edge effect of the signal in the time domain or frequency domain is reduced or eliminated, making the speech signal smoother, thereby generating an efficient speech enhancement model.
Smart Images

Figure CN119920263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech enhancement technology, and in particular to a training method, device, electronic equipment and storage medium for a speech enhancement model. Background Art
[0002] Environmental noise mainly refers to the sound generated in industrial production, construction, transportation and social life that interferes with the surrounding living environment. It is a sound phenomenon that is processed through adaptive noise suppression technology. This technology is an algorithm for signal processing of a fully digital machine. It analyzes the sound signal, distinguishes the noise and voice components, suppresses the noise as needed, and enhances the voice to improve the speech clarity in a noisy environment.
[0003] Adaptive filter technology is widely used in the existing speech noise reduction models on the market. This technology shows good results in strong noise backgrounds, especially adaptive filters based on adaptive noise cancellation. Adaptive filters enhance speech signals by adaptively adjusting filter parameters to adapt to the ever-changing noise environment. However, the computational requirements of traditional speech noise reduction models are usually very large, and they cannot achieve efficient noise elimination for weaker noise backgrounds, which seriously affects the accuracy of the model and thus cannot achieve good speech enhancement effects. Summary of the invention
[0004] In view of this, the present invention aims to propose a training method, device, electronic device and storage medium for a speech enhancement model to solve the problem that the traditional speech denoising model has large computational requirements and cannot achieve efficient noise elimination for weaker noise backgrounds, which seriously affects the accuracy of the model and thus cannot achieve a good speech enhancement effect.
[0005] According to a first aspect of the present invention, a method for training a speech enhancement model is provided, the method comprising:
[0006] Collecting audio data, where the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise;
[0007] Extracting a first audio feature of the first audio and a second audio feature of the second audio;
[0008] Using a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio;
[0009] Converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio;
[0010] Calculating a loss function between the first denoised signal and the second denoised signal;
[0011] The loss function is used to adjust the enhancement parameters of the first large model to generate a speech enhancement model, where the speech enhancement model is used to eliminate environmental noise in the audio data.
[0012] Optionally, after collecting the audio data, the method includes:
[0013] Separating a first audio containing environmental noise and a second audio not containing environmental noise from the audio data;
[0014] A negative sample data set is constructed using the first audio containing environmental noise, and a positive sample data set is constructed using the second audio not containing environmental noise, wherein the second audio in the positive sample data set corresponds one-to-one to the first audio in the negative sample data set.
[0015] Optionally, after collecting the audio data, the method includes:
[0016] Inputting the first audio in the negative sample data set into a preset second largest model for noise reduction, and obtaining a first noise-reduced audio output by the second largest model;
[0017] Determine a second audio corresponding to the first audio from the positive sample data set;
[0018] Comparing the first noise reduction audio and the second audio to obtain a comparison result, and generating a first loss function according to the comparison result;
[0019] The denoising parameters of the second large model are adjusted according to the first loss function to generate the first large model.
[0020] Optionally, the calculating a loss function between the first denoised signal and the second denoised signal includes:
[0021] Comparing the first noise reduction signal with the second noise reduction signal to obtain a second comparison result, and generating a second loss function according to the comparison result;
[0022] Using the loss function to adjust the enhancement parameters of the first large model includes:
[0023] The second loss function is used to adjust the enhancement parameters of the first large model.
[0024] Optionally, extracting a first audio feature of the first audio and a second audio feature of the second audio includes:
[0025] Divide the first audio and the second audio into frames to obtain a plurality of first audio segments of the first audio and a plurality of second audio segments of the second audio;
[0026] Converting the first audio segment into a first audio sub-signal, and converting the second audio segment into a second audio sub-signal;
[0027] multiplying the first audio sub-signal and the second audio sub-signal by a preset window function respectively to obtain a third audio sub-signal corresponding to the first audio sub-signal and a fourth audio sub-signal corresponding to the second audio sub-signal;
[0028] Performing a forward time-frequency transform on the third audio sub-signal and the fourth audio sub-signal to obtain a first frequency spectrum corresponding to the third audio sub-signal and a second frequency spectrum corresponding to the fourth audio sub-signal;
[0029] A first audio feature of the first audio is extracted from the first frequency spectrum, and a second audio feature of the second audio is extracted from the second frequency spectrum.
[0030] Optionally, the third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and the performing forward time-frequency transform on the third audio sub-signal and the fourth audio sub-signal includes:
[0031] Convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain to obtain a first real part and a first imaginary part corresponding to the first audio point, and a second real part and a second imaginary part corresponding to the second audio point;
[0032] A first frequency spectrum corresponding to the third audio sub-signal is determined according to the first real part and the first imaginary part, and a second frequency spectrum corresponding to the fourth audio sub-signal is determined according to the second real part and the second imaginary part.
[0033] Optionally, converting the frequency domains of the first noise reduction feature and the second noise reduction feature includes:
[0034] Acquire first frequency domain information of the first noise reduction feature according to the first frequency spectrum, and acquire second frequency domain information of the second noise reduction feature according to the second frequency spectrum;
[0035] Performing inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain a first time domain frame of the first noise reduction feature and a second time domain frame of the second noise reduction feature;
[0036] Acquire a frame shift when framing the first audio and the second audio, where the frame shift is used to represent an overlapping distance between two consecutive frames;
[0037] Superimposing the first time domain frame according to the frame shift to obtain a first superimposed signal, and adjusting a first amplitude and a first phase of the first superimposed signal to obtain a first noise reduction signal;
[0038] The second time domain frame is superimposed according to the frame shift to obtain a second superimposed signal, and a second amplitude and a second phase of the second superimposed signal are adjusted to obtain a second noise reduction signal.
[0039] According to another aspect of the present invention, there is provided a training device for a speech enhancement model, the device comprising:
[0040] An audio collection module, used for collecting audio data, wherein the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise;
[0041] a feature extraction module, configured to extract a first audio feature of the first audio and a second audio feature of the second audio;
[0042] A first noise reduction module, configured to use a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio;
[0043] A time-frequency conversion module, used for converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio;
[0044] A loss calculation module, used to calculate a loss function between the first denoised signal and the second denoised signal;
[0045] A model generation module is used to use the loss function to adjust the enhancement parameters of the first large model to generate a speech enhancement model, and the speech enhancement model is used to eliminate the environmental noise in the audio data.
[0046] Optionally, the audio acquisition module includes:
[0047] An audio separation module, used to separate a first audio containing environmental noise and a second audio not containing environmental noise from the audio data;
[0048] The sample data set construction module is used to construct a negative sample data set using the first audio containing environmental noise, and to construct a positive sample data set using the second audio not containing environmental noise, wherein the second audio in the positive sample data set corresponds one-to-one to the first audio in the negative sample data set.
[0049] Optionally, after the audio acquisition module, the device includes:
[0050] A first noise reduction audio generation module is used to input the first audio in the negative sample data set into a preset second large model for noise reduction, and obtain a first noise reduction audio output by the second large model;
[0051] A second audio determination module, configured to determine a second audio corresponding to the first audio from the positive sample data set;
[0052] A first loss function generating module, configured to compare the first noise reduction audio and the second audio to obtain a comparison result, and generate a first loss function according to the comparison result;
[0053] The first large model generation module is used to adjust the noise reduction parameters of the second large model according to the first loss function to generate the first large model.
[0054] Optionally, the loss calculation module is specifically used to compare the first noise reduction signal with the second noise reduction signal to obtain a second comparison result, and generate a second loss function according to the comparison result;
[0055] The model generation module is specifically used to adjust the enhancement parameters of the first large model using the second loss function.
[0056] Optionally, the feature extraction module includes:
[0057] an audio framing module, configured to frame the first audio and the second audio to obtain a plurality of first audio segments of the first audio and a plurality of second audio segments of the second audio;
[0058] A first signal conversion module, configured to convert the first audio segment into a first audio sub-signal, and convert the second audio segment into a second audio sub-signal;
[0059] a first signal conversion module, configured to multiply the first audio sub-signal and the second audio sub-signal by a preset window function respectively, to obtain a third audio sub-signal corresponding to the first audio sub-signal and a fourth audio sub-signal corresponding to the second audio sub-signal;
[0060] a forward time-frequency transformation module, configured to perform forward time-frequency transformation on the third audio sub-signal and the fourth audio sub-signal to obtain a first frequency spectrum corresponding to the third audio sub-signal and a second frequency spectrum corresponding to the fourth audio sub-signal;
[0061] The feature extraction submodule is used to extract a first audio feature of the first audio from the first spectrum, and to extract a second audio feature of the second audio from the second spectrum.
[0062] Optionally, the third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and the forward time-frequency transform module includes:
[0063] a time domain conversion module, configured to convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain to obtain a first real part and a first imaginary part corresponding to the first audio point, and a second real part and a second imaginary part corresponding to the second audio point;
[0064] The spectrum determination module is configured to determine a first spectrum corresponding to the third audio sub-signal according to the first real part and the first imaginary part, and to determine a second spectrum corresponding to the fourth audio sub-signal according to the second real part and the second imaginary part.
[0065] Optionally, the time-frequency conversion module includes:
[0066] A frequency domain information acquisition module, configured to acquire first frequency domain information of the first noise reduction feature according to the first frequency spectrum, and acquire second frequency domain information of the second noise reduction feature according to the second frequency spectrum;
[0067] an inverse time-frequency transformation module, configured to perform an inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain a first time domain frame of the first noise reduction feature and a second time domain frame of the second noise reduction feature;
[0068] A frame shift acquisition module, used to acquire a frame shift when framing the first audio and the second audio, wherein the frame shift is used to represent an overlapping distance between two consecutive frames;
[0069] A first noise reduction signal generating module, configured to obtain a first superimposed signal by superimposing the first time domain frame according to the frame shift, and adjust a first amplitude and a first phase of the first superimposed signal to obtain a first noise reduction signal;
[0070] The second noise reduction signal generating module is used to obtain a second superimposed signal by superimposing the second time domain frame according to the frame shift, and adjust the second amplitude and the second phase of the second superimposed signal to obtain a second noise reduction signal.
[0071] According to another aspect of the present invention, there is also provided an electronic device, comprising:
[0072] processor;
[0073] a memory for storing instructions executable by the processor;
[0074] Wherein, the processor is configured to execute the instructions to implement the training method of the speech enhancement model as described above.
[0075] According to another aspect of the present invention, a readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the training method of the speech enhancement model as described above are implemented.
[0076] The training method of the speech enhancement model provided by the present invention collects audio data including at least a first audio containing environmental noise and a second audio not containing environmental noise, extracts a first audio feature of the first audio and a second audio feature of the second audio; uses a preset first large model to reduce noise on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio; converts the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio; calculates a loss function between the first noise reduction signal and the second noise reduction signal; uses the loss function to adjust the enhancement parameters of the first large model to generate a speech enhancement model for eliminating environmental noise in audio data. The present invention simplifies the noise reduction step by extracting the features of the audio data, thereby improving the efficiency of noise reduction, and adjusts the enhancement parameters of the first large model by the loss function to reduce or eliminate the edge effect of the signal in the time domain or frequency domain, making the signal smoother, thereby generating a mature speech enhancement model.
[0077] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0079] Figure 1 is a flowchart of a method for training a speech enhancement model provided by an embodiment of the present invention;
[0080] Figure 2 It is a structural schematic diagram of a training device for a speech enhancement model provided by an embodiment of the present invention;
[0081] Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. However, it can be understood by those skilled in the art that in the embodiments of the present invention, many technical details are proposed in order to enable readers to better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present invention can also be implemented. The division of the following embodiments is for the convenience of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined and referenced with each other without contradiction.
[0083] Adaptive filter technology is widely used in the existing speech noise reduction models on the market. This technology shows good results in strong noise backgrounds, especially adaptive filters based on adaptive noise cancellation. Adaptive filters enhance speech signals by adaptively adjusting filter parameters to adapt to the ever-changing noise environment. However, the computational requirements of traditional speech noise reduction models are usually very large, and they cannot achieve efficient noise elimination for weaker noise backgrounds, which seriously affects the accuracy of the model and thus cannot achieve good speech enhancement effects.
[0084] Based on this, the present invention proposes a training method, device, electronic device and storage medium for a speech enhancement model to solve the problem that the traditional speech noise reduction model has a large computational requirement and cannot achieve efficient noise elimination for weak noise backgrounds, which seriously affects the accuracy of the model and thus cannot achieve a good speech enhancement effect. Figure 1 , shows a flowchart of the steps of the training method of the speech enhancement model provided by an embodiment of the present invention, and the method may include:
[0085] Step 101 : collecting audio data, where the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise.
[0086] The present invention can pre-set some collection scenes when collecting audio data, such as office, street, cafe, home environment and other scenes, wherein the audio data of the office can include human voice, keyboard sound, mouse sound and air conditioner sound, the audio data of the street can include human voice and traffic noise, the audio data of the cafe can include human voice, coffee machine running sound and music sound, and the home environment can include human voice, pet sound, TV sound and electrical appliance running sound. Collecting audio data in these preset collection scenes can further process different types of background noise more accurately.
[0087] When collecting audio data, you need to ensure that the same recording equipment and parameter settings are used to collect audio data to avoid the impact of equipment differences on the data. You can also annotate the collected audio data in detail, such as including environment type, noise type, voice content, etc., for subsequent analysis and use. In addition, a sufficient amount of audio data needs to be collected when training the model. For example, you can collect 100 audio segments for each type of environmental noise, and each audio segment is 10 seconds long.
[0088] Collecting the first audio containing environmental noise and the second audio without environmental noise is a key step in building a high-quality audio data set. The present invention provides effective data support for tasks such as noise suppression, speech enhancement, and speech recognition by reasonably selecting the collection environment, setting up the recording equipment, annotating the data, and performing subsequent processing and analysis.
[0089] Since the collected audio data contains normal audio and environmental noise, the collected audio data needs to be purified after it is collected. The specific steps are as follows:
[0090] Separating a first audio containing environmental noise and a second audio not containing environmental noise from the audio data;
[0091] A negative sample data set is constructed using a first audio containing environmental noise, and a positive sample data set is constructed using a second audio not containing environmental noise, wherein the second audio in the positive sample data set corresponds one-to-one to the first audio in the negative sample data set.
[0092] The first audio in the negative sample data set is input into the preset second largest model for noise reduction, and the first noise reduction audio is output by the second largest model. In one embodiment of the present invention, the second largest model is an initial speech noise reduction model.
[0093] Determine the second audio corresponding to the first audio from the positive sample data set; since the positive sample data set and the negative sample data set of the present invention are one-to-one corresponding, the second audio corresponding to the first audio can be determined from the positive sample data set. After determining the second audio, compare the first noise reduction audio and the second audio to obtain a comparison result, and generate a first loss function based on the comparison result; the second audio is audio data that does not contain environmental noise separated from the original audio data, that is, the second audio can be regarded as pure audio, and the first noise reduction audio is infinitely close to the pure audio, therefore, the first loss function between the second audio and the first noise reduction audio can be calculated, and the noise reduction ability of the second largest model can be judged based on the convergence degree of the first loss function. The more convergent the first loss function is, the better the noise reduction ability of the second largest model is.
[0094] After calculating the first loss function, the denoising parameters of the second largest model are adjusted according to the first loss function, and the second largest model with adjusted denoising parameters is used to perform denoising on the first audio in the negative sample data set again. The new denoised audio generated after denoising is compared with the second audio corresponding to the first audio to generate a new loss function, and the convergence degree of the new loss function is evaluated. When the convergence degree of the new loss function reaches the expected value, the first largest model is generated.
[0095] Step 102: extract a first audio feature of the first audio and a second audio feature of the second audio.
[0096] The present invention requires feature extraction of audio data. Since the first audio containing environmental noise and the second audio not containing environmental noise are separated from the audio data after the audio data is collected in step 101, the first audio feature of the first audio and the second audio feature of the second audio can be extracted. The present invention does not specifically limit the feature extraction method.
[0097] Step 103: Use a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio.
[0098] The first large model generated in the above steps greatly improves the accuracy of speech denoising compared to the initial speech denoising model. However, the structure of the first large model is relatively complex and requires more computing time. Therefore, the present invention performs knowledge distillation on the first large model to generate a more streamlined speech enhancement model.
[0099] First, the first audio feature and the second audio feature are denoised using the first large model to obtain the first audio feature of the first audio and the second audio feature of the second audio, specifically including the following steps:
[0100] The first audio and the second audio are divided into frames to obtain a plurality of first audio segments of the first audio and a plurality of second audio segments of the second audio. First, the continuous audio data is divided into a plurality of short audio segments so that each frame can be processed independently. Before the division of the frames, the frame length and the frame shift can be set first. For example, the frame length can be set to 25 to 50 ms, and the frame shift is usually 50% of the frame length, that is, it can be set to 12.5 to 25 ms.
[0101] The first audio segment is converted into a first audio sub-signal, and the second audio segment is converted into a second audio sub-signal. For each frame of the audio segment, its time domain signal is extracted, so as to convert the audio segment into a corresponding audio sub-signal, that is, the first time domain signal of the first audio segment is extracted, so as to convert the first audio segment into a corresponding first audio sub-signal, and the second time domain signal of the second audio segment is extracted, so as to convert the second audio segment into a corresponding second audio sub-signal.
[0102] The first audio sub-signal and the second audio sub-signal are respectively multiplied by a preset window function to obtain a third audio sub-signal corresponding to the first audio sub-signal and a fourth audio sub-signal corresponding to the second audio sub-signal. The present invention can weight the first audio sub-signal and the second audio sub-signal of each frame by a window function to reduce the edge effect of the first audio sub-signal and the second audio sub-signal. Common window functions include Hanning Window, Hamming Window, Blackman Window and Kaiser Window.
[0103] The calculation formula of the Hamming window is:
[0104]
[0105] Among them, w 1 (n) is the value of the Hamming window, N is the length of the window, and n is a sample point in the window.
[0106] The calculation formula of the Hanning window is:
[0107]
[0108] Among them, w 1 (n) is the value of the Hanning window, N is the length of the window, and n is a sample point within the window.
[0109] The third audio sub-signal and the fourth audio sub-signal are forwardly transformed in time-frequency mode to obtain a first spectrum corresponding to the third audio sub-signal and a second spectrum corresponding to the fourth audio sub-signal. The forward time-frequency transform of the present invention refers to converting a signal from a time domain to a time-frequency domain. When the third audio sub-signal and the fourth audio sub-signal are forwardly transformed in time-frequency mode, an audio feature with noise can be obtained by any of the following methods: Fourier transform, Laplace transform, z transform, discrete cosine transform. For example, the Fourier transform calculation formula is as follows:
[0110]
[0111] Where x(h) is the signal in the time domain, x(k) is the signal in the frequency domain, H is the length of the signal, and j is the imaginary unit.
[0112] A first audio feature of the first audio is extracted from the first spectrum, and a second audio feature of the second audio is extracted from the second spectrum. After the forward time-frequency conversion is completed to obtain the first spectrum and the second spectrum, feature extraction can be performed on the first spectrum and the second spectrum to obtain the first audio feature of the first audio and the second audio feature of the second audio, respectively. The first audio feature and the second audio feature are essentially the spectrum features of the first audio and the spectrum features of the second audio. The spectrum features of the present invention at least include amplitude and phase.
[0113] In an embodiment of the present invention, the third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and forward time-frequency transformation is performed on the third audio sub-signal and the fourth audio sub-signal, including:
[0114] The third audio sub-signal and the fourth audio sub-signal are converted from the time domain to the time-frequency domain to obtain the first real part and the first imaginary part corresponding to the first audio point, and the second real part and the second imaginary part corresponding to the second audio point. The amplitude and phase of the entire spectrum can be calculated by the real part and the imaginary part, and the calculation formula is as follows:
[0115]
[0116] Therefore, for the third audio sub-signal:
[0117]
[0118] For the fourth audio sub-signal:
[0119]
[0120] A first spectrum corresponding to the third audio sub-signal is determined according to the first real part and the first imaginary part, and a second spectrum corresponding to the fourth audio sub-signal is determined according to the second real part and the second imaginary part.
[0121] Step 104, converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio, the specific steps are as follows:
[0122] Acquire first frequency domain information of a first noise reduction feature according to the first frequency spectrum, and acquire second frequency domain information of a second noise reduction feature according to the second frequency spectrum;
[0123] The first frequency domain information and the second frequency domain information are inversely transformed in time-frequency manner to obtain a first time domain frame of the first noise reduction feature and a second time domain frame of the second noise reduction feature, wherein the first time domain frame and the second time domain frame respectively represent a signal of the first noise reduction feature in the time domain and a signal of the second noise reduction feature in the time domain.
[0124] A frame shift is obtained when the first audio and the second audio are divided into frames. The frame shift is used to represent the overlapping distance between two consecutive frames, which is usually half of the frame length.
[0125] The time domain frames are superimposed according to the frame shift to obtain the noise-reduced signal, and the amplitude and phase of the noise-reduced signal are adjusted. The specific superposition process in the present invention is as follows:
[0126] The first superimposed signal is obtained by superimposing the first time domain frame according to the frame shift, and the first amplitude and the first phase of the first superimposed signal are adjusted to obtain the first noise reduction signal. The present invention superimposes the first time domain frame on the signal of the previous frame according to the frame shift to obtain the first superimposed signal, and adjusts the amplitude and the phase of the first superimposed signal to ensure a smooth transition of the signal. For example, the amplitude and the phase are adjusted by weighted averaging or a filter.
[0127] The second superimposed signal is obtained by superimposing the second time domain frame according to the frame shift, and the second amplitude and the second phase of the second superimposed signal are adjusted to obtain the second noise reduction signal. Similarly, the second time domain frame is superimposed on the signal of the previous frame according to the frame shift to obtain the second superimposed signal, and the amplitude and the phase of the second superimposed signal are adjusted to ensure a smooth transition of the signal.
[0128] Step 105, calculating the loss function between the first noise reduction signal and the second noise reduction signal, the specific steps are as follows:
[0129] The same number of acquisition points are selected from the first denoised signal and the second denoised signal respectively, and the first denoised signal and the second denoised signal are compared according to these acquisition points to obtain a second comparison result, and a second loss function is generated according to the comparison result. The formula for calculating the second loss function can be:
[0130]
[0131] Wherein, MSE is the error value between the first denoised signal and the second denoised signal, R is the number of acquisition points in the first denoised signal and the second denoised signal, and y i is the noise value corresponding to the second acquisition point on the second noise reduction signal, Y i is the noise value corresponding to the first acquisition point on the first denoised signal.
[0132] In another embodiment of the present invention, the second loss function may also be calculated by the following formula:
[0133] FL(P t )=-α t ·(1-P t ) γ ·log(P t )
[0134] Among them, P tis the predicted probability for a certain category; α t It is a factor that balances positive and negative samples and is used to adjust the weights of the positive and negative sample datasets. γ is an adjustment factor that can also adjust the weights of the positive and negative sample datasets.
[0135] Step 106, using the loss function to adjust the enhancement parameters of the first large model to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.
[0136] After generating the second loss function, the second loss function may be used to adjust the enhancement parameters of the first large model to generate a speech enhancement model.
[0137] After the speech enhancement model is generated, the audio to be processed that needs to be denoised can be input into the speech enhancement model and processed using the speech enhancement model of the present invention to obtain the target audio.
[0138] The training method of the speech enhancement model provided by the present invention collects audio data including at least a first audio containing environmental noise and a second audio not containing environmental noise, extracts a first audio feature of the first audio and a second audio feature of the second audio; uses a preset first large model to reduce noise on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio; converts the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio; calculates a loss function between the first noise reduction signal and the second noise reduction signal; uses the loss function to adjust the enhancement parameters of the first large model to generate a speech enhancement model for eliminating environmental noise in audio data. The present invention simplifies the noise reduction step by extracting the features of the audio data, thereby improving the efficiency of noise reduction, and adjusts the enhancement parameters of the first large model by the loss function to reduce or eliminate the edge effect of the signal in the time domain or frequency domain, making the signal smoother, thereby generating a mature speech enhancement model.
[0139] Further, see Figure 2 , shows a schematic diagram of the structure of a training device for a speech enhancement model provided by an embodiment of the present invention, the device may include:
[0140] The audio collection module 201 is used to collect audio data, where the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise;
[0141] A feature extraction module 202, configured to extract a first audio feature of the first audio and a second audio feature of the second audio;
[0142] A first noise reduction module 203 is used to perform noise reduction on the first audio feature and the second audio feature by using a preset first large model to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio;
[0143] A time-frequency conversion module 204, configured to convert the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio;
[0144] A loss calculation module 205, used to calculate a loss function between the first noise reduction signal and the second noise reduction signal;
[0145] The model generation module 206 is used to adjust the enhancement parameters of the first large model using the loss function to generate a speech enhancement model, and the speech enhancement model is used to eliminate environmental noise in the audio data.
[0146] In one embodiment of the present invention, the audio acquisition module 201 includes:
[0147] An audio separation module, used to separate a first audio containing environmental noise and a second audio not containing environmental noise from the audio data;
[0148] The sample data set construction module is used to construct a negative sample data set using a first audio containing environmental noise, and to construct a positive sample data set using a second audio not containing environmental noise, wherein the second audio in the positive sample data set corresponds one-to-one to the first audio in the negative sample data set.
[0149] In one embodiment of the present invention, after the audio acquisition module 201, the device includes:
[0150] A first noise reduction audio generation module is used to input the first audio in the negative sample data set into a preset second largest model for noise reduction, and obtain a first noise reduction audio output by the second largest model;
[0151] A second audio determination module, used to determine a second audio corresponding to the first audio from the positive sample data set;
[0152] A first loss function generating module, used for comparing the first noise reduction audio and the second audio to obtain a comparison result, and generating a first loss function according to the comparison result;
[0153] The first large model generation module is used to adjust the noise reduction parameters of the second large model according to the first loss function to generate the first large model.
[0154] In one embodiment of the present invention, the loss calculation module 205 is specifically used to compare the first noise reduction signal with the second noise reduction signal to obtain a second comparison result, and generate a second loss function according to the comparison result;
[0155] The model generation module 206 is specifically used to adjust the enhancement parameters of the first large model using the second loss function.
[0156] In one embodiment of the present invention, the feature extraction module 202 includes:
[0157] An audio framing module, used to frame the first audio and the second audio to obtain a plurality of first audio segments of the first audio and a plurality of second audio segments of the second audio;
[0158] A first signal conversion module, configured to convert a first audio segment into a first audio sub-signal, and convert a second audio segment into a second audio sub-signal;
[0159] A first signal conversion module, configured to multiply the first audio sub-signal and the second audio sub-signal by a preset window function respectively, to obtain a third audio sub-signal corresponding to the first audio sub-signal and a fourth audio sub-signal corresponding to the second audio sub-signal;
[0160] a forward time-frequency transformation module, configured to perform forward time-frequency transformation on the third audio sub-signal and the fourth audio sub-signal to obtain a first frequency spectrum corresponding to the third audio sub-signal and a second frequency spectrum corresponding to the fourth audio sub-signal;
[0161] The feature extraction submodule is used to extract a first audio feature of the first audio from the first spectrum, and to extract a second audio feature of the second audio from the second spectrum.
[0162] In an embodiment of the present invention, the third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and the forward time-frequency conversion module includes:
[0163] a time domain conversion module, configured to convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain, and obtain a first real part and a first imaginary part corresponding to the first audio point, and a second real part and a second imaginary part corresponding to the second audio point;
[0164] The spectrum determination module is used to determine a first spectrum corresponding to the third audio sub-signal according to the first real part and the first imaginary part, and to determine a second spectrum corresponding to the fourth audio sub-signal according to the second real part and the second imaginary part.
[0165] In one embodiment of the present invention, the time-frequency conversion module 204 includes:
[0166] A frequency domain information acquisition module, used to acquire first frequency domain information of a first noise reduction feature according to the first spectrum, and to acquire second frequency domain information of a second noise reduction feature according to the second spectrum;
[0167] An inverse time-frequency transformation module, configured to perform an inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain a first time domain frame of the first noise reduction feature and a second time domain frame of the second noise reduction feature;
[0168] A frame shift acquisition module, used to acquire a frame shift when framing the first audio and the second audio, where the frame shift is used to represent an overlapping distance between two consecutive frames;
[0169] A first noise reduction signal generating module, configured to obtain a first superimposed signal by superimposing the first time domain frame according to the frame shift, and to adjust a first amplitude and a first phase of the first superimposed signal to obtain a first noise reduction signal;
[0170] The second noise reduction signal generating module is used to obtain a second superimposed signal by superimposing the second time domain frame according to the frame shift, and adjust the second amplitude and the second phase of the second superimposed signal to obtain a second noise reduction signal.
[0171] Reference Figure 3 , an embodiment of the present invention further provides an electronic device, such as Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.
[0172] A processor 301, a memory 303 for storing processor executable instructions;
[0173] The processor 301 is configured to execute instructions to implement the above training method of the speech enhancement model:
[0174] Collecting audio data, the audio data at least including a first audio containing environmental noise and a second audio not containing environmental noise;
[0175] Extracting a first audio feature of the first audio and a second audio feature of the second audio;
[0176] Using a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio;
[0177] Converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio;
[0178] Calculating a loss function between the first denoised signal and the second denoised signal;
[0179] The loss function is used to adjust the enhancement parameters of the first model to generate a speech enhancement model, which is used to eliminate environmental noise in audio data.
[0180] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0181] The communication interface is used for communication between the above terminal and other devices.
[0182] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0183] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0184] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the speech enhancement model in any of the above embodiments is implemented.
[0185] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.
[0186] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0187] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0188] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for training a speech enhancement model, characterized in that: The method comprises: Collecting audio data, where the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise; Extracting a first audio feature of the first audio and a second audio feature of the second audio; Using a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio; Converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio; Calculating a loss function between the first denoised signal and the second denoised signal; The loss function is used to adjust the enhancement parameters of the first large model to generate a speech enhancement model, where the speech enhancement model is used to eliminate environmental noise in the audio data.
2. The method according to claim 1, characterized in that After collecting the audio data, the method includes: Separating a first audio containing environmental noise and a second audio not containing environmental noise from the audio data; A negative sample data set is constructed using the first audio containing environmental noise, and a positive sample data set is constructed using the second audio not containing environmental noise, wherein the second audio in the positive sample data set corresponds one-to-one to the first audio in the negative sample data set.
3. The method according to claim 2, characterized in that After collecting the audio data, the method includes: Inputting the first audio in the negative sample data set into a preset second largest model for noise reduction, and obtaining a first noise-reduced audio output by the second largest model; Determine a second audio corresponding to the first audio from the positive sample data set; Comparing the first noise reduction audio and the second audio to obtain a comparison result, and generating a first loss function according to the comparison result; The denoising parameters of the second large model are adjusted according to the first loss function to generate the first large model.
4. The method according to claim 3, characterized in that The calculating a loss function between the first noise reduction signal and the second noise reduction signal includes: Comparing the first noise reduction signal with the second noise reduction signal to obtain a second comparison result, and generating a second loss function according to the comparison result; Using the loss function to adjust the enhancement parameters of the first large model includes: The second loss function is used to adjust the enhancement parameters of the first large model.
5. The method according to claim 1, characterized in that The extracting a first audio feature of the first audio and a second audio feature of the second audio includes: Divide the first audio and the second audio into frames to obtain a plurality of first audio segments of the first audio and a plurality of second audio segments of the second audio; Converting the first audio segment into a first audio sub-signal, and converting the second audio segment into a second audio sub-signal; multiplying the first audio sub-signal and the second audio sub-signal by a preset window function respectively to obtain a third audio sub-signal corresponding to the first audio sub-signal and a fourth audio sub-signal corresponding to the second audio sub-signal; Performing a forward time-frequency transform on the third audio sub-signal and the fourth audio sub-signal to obtain a first frequency spectrum corresponding to the third audio sub-signal and a second frequency spectrum corresponding to the fourth audio sub-signal; A first audio feature of the first audio is extracted from the first frequency spectrum, and a second audio feature of the second audio is extracted from the second frequency spectrum.
6. The method according to claim 5, characterized in that The third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and the forward time-frequency transformation of the third audio sub-signal and the fourth audio sub-signal includes: Convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain to obtain a first real part and a first imaginary part corresponding to the first audio point, and a second real part and a second imaginary part corresponding to the second audio point; A first frequency spectrum corresponding to the third audio sub-signal is determined according to the first real part and the first imaginary part, and a second frequency spectrum corresponding to the fourth audio sub-signal is determined according to the second real part and the second imaginary part.
7. The method according to claim 5, characterized in that Converting the first noise reduction feature and the second noise reduction feature into a frequency domain includes: Acquire first frequency domain information of the first noise reduction feature according to the first frequency spectrum, and acquire second frequency domain information of the second noise reduction feature according to the second frequency spectrum; Performing inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain a first time domain frame of the first noise reduction feature and a second time domain frame of the second noise reduction feature; Acquire a frame shift when framing the first audio and the second audio, where the frame shift is used to represent an overlapping distance between two consecutive frames; Superimposing the first time domain frame according to the frame shift to obtain a first superimposed signal, and adjusting a first amplitude and a first phase of the first superimposed signal to obtain a first noise reduction signal; The second time domain frame is superimposed according to the frame shift to obtain a second superimposed signal, and a second amplitude and a second phase of the second superimposed signal are adjusted to obtain a second noise reduction signal.
8. A training device for a speech enhancement model, characterized in that: The device comprises: An audio collection module, used for collecting audio data, wherein the audio data at least includes a first audio containing environmental noise and a second audio not containing environmental noise; A feature extraction module, configured to extract a first audio feature of the first audio and a second audio feature of the second audio; A first noise reduction module, configured to use a preset first large model to perform noise reduction on the first audio feature and the second audio feature to obtain a first noise reduction feature of the first audio and a second noise reduction feature of the second audio; A time-frequency conversion module, used for converting the frequency domains of the first noise reduction feature and the second noise reduction feature to obtain a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio; A loss calculation module, used to calculate a loss function between the first denoised signal and the second denoised signal; A model generation module is used to use the loss function to adjust the enhancement parameters of the first large model to generate a speech enhancement model, and the speech enhancement model is used to eliminate the environmental noise in the audio data.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to execute the instructions to implement the training method of the speech enhancement model as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the training method of the speech enhancement model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image enhancement method and device
CN112258401A
Voice noise reduction model training method and device, storage medium and electronic device
CN114974283A
Data recommendation method and device based on power grid information
CN116775962A
Speech enhancement method, training method of speech enhancement network and electronic equipment
CN116959471A
Audio noise reduction method and device, electronic equipment and storage medium
CN118486323A
Cited By
Noise reduction and audio enhancement system of wearable hearing aid device
CN121662066A