Training methods, devices, electronic equipment, and storage media for speech enhancement models

By collecting and processing audio data containing environmental noise, extracting features, and adjusting model parameters using a loss function, a highly efficient speech enhancement model was generated. This solved the problems of high computational requirements and poor performance in eliminating weak noise in traditional models, achieving a more efficient speech enhancement effect.

CN119920263BActive Publication Date: 2025-10-28CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907760.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-28
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Traditional speech denoising models require a large amount of computation and cannot efficiently eliminate weak background noise, affecting model accuracy and speech enhancement effects.

Method used

Audio data with and without environmental noise are collected, features are extracted and noise reduction is performed, and model parameters are adjusted through a loss function to generate a speech enhancement model.

Benefits of technology

The noise reduction process has been simplified, the noise reduction efficiency has been improved, the signal edge effect has been reduced, and a smoother speech enhancement model has been generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920263B_ABST
    Figure CN119920263B_ABST
Patent Text Reader

Abstract

This invention provides a training method, apparatus, electronic device, and storage medium for a speech enhancement model, relating to the field of speech enhancement technology. The method involves acquiring audio data, including at least a first audio signal containing environmental noise and a second audio signal not containing environmental noise; extracting first audio features from the first audio signal and second audio features from the second audio signal; using a pre-set first large model to denoise the first and second audio features, obtaining first denoised features of the first audio signal and second denoised features of the second audio signal; converting the frequency domains of the first and second denoised features to obtain a first denoised signal corresponding to the first audio signal and a second denoised signal corresponding to the second audio signal; calculating a loss function between the first and second denoised signals; and adjusting the enhancement parameters of the first large model using the loss function to generate a speech enhancement model for eliminating environmental noise in the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech enhancement technology, and in particular to a training method, apparatus, electronic device, and storage medium for a speech enhancement model. Background Technology

[0002] Environmental noise mainly refers to the sounds that disturb the surrounding living environment generated in industrial production, construction, transportation, and social life. It is a sound phenomenon that is processed by adaptive noise suppression technology. This technology is an algorithm in signal processing of a fully digital machine. By analyzing the sound signal, it distinguishes between noise and speech components, suppresses noise as needed, and enhances speech to improve speech intelligibility in noisy environments.

[0003] Currently, most existing speech denoising models on the market utilize adaptive filter technology, which shows good performance in strong noise backgrounds. In particular, adaptive filters based on adaptive noise cancellation enhance the speech signal by adaptively adjusting the filter parameters to adapt to the constantly changing noise environment. However, traditional speech denoising models typically have very large computational requirements and cannot achieve efficient noise elimination in weak noise backgrounds, which seriously affects the accuracy of the model and thus fails to achieve good speech enhancement results. Summary of the Invention

[0004] In view of this, the present invention aims to propose a training method, device, electronic device and storage medium for a speech enhancement model, which solves the problem that traditional speech denoising models have large computational requirements and cannot achieve efficient noise elimination in weak noise backgrounds, which seriously affects the accuracy of the model and thus cannot achieve good speech enhancement results.

[0005] According to a first aspect of the present invention, a method for training a speech enhancement model is provided, the method comprising:

[0006] Collect audio data, wherein the audio data includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise;

[0007] Extract the first audio features of the first audio and the second audio features of the second audio;

[0008] A preset first large model is used to denoise the first audio feature and the second audio feature to obtain the first denoised feature of the first audio and the second denoised feature of the second audio.

[0009] By converting the frequency domains of the first noise reduction feature and the second noise reduction feature, a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio are obtained;

[0010] Calculate the loss function between the first denoised signal and the second denoised signal;

[0011] The enhancement parameters of the first large model are adjusted using the loss function to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.

[0012] Optionally, after acquiring audio data, the method includes:

[0013] Separate a first audio audio containing ambient noise and a second audio audio not containing ambient noise from the audio data;

[0014] A negative sample dataset is constructed using the first audio containing environmental noise, and a positive sample dataset is constructed using the second audio that does not contain environmental noise, wherein the second audio in the positive sample dataset corresponds one-to-one with the first audio in the negative sample dataset.

[0015] Optionally, after acquiring audio data, the method includes:

[0016] The first audio from the negative sample dataset is input into a preset second large model for noise reduction, and the first noise-reduced audio frequency output by the second large model is obtained.

[0017] Determine the second audio corresponding to the first audio from the positive sample dataset;

[0018] The first noise-reduced audio frequency and the second audio frequency are compared to obtain a comparison result, and a first loss function is generated based on the comparison result;

[0019] The denoising parameters of the second largest model are adjusted based on the first loss function to generate the first largest model.

[0020] Optionally, calculating the loss function between the first denoised signal and the second denoised signal includes:

[0021] A second comparison result is obtained by comparing the first denoised signal with the second denoised signal, and a second loss function is generated based on the comparison result;

[0022] Adjusting the enhancement parameters of the first large model using the loss function includes:

[0023] The enhancement parameters of the first large model are adjusted using the second loss function.

[0024] Optionally, extracting the first audio feature of the first audio and the second audio feature of the second audio includes:

[0025] The first audio and the second audio are divided into frames to obtain several frames of the first audio segment and several frames of the second audio segment.

[0026] The first audio segment is converted into a first audio sub-signal, and the second audio segment is converted into a second audio sub-signal;

[0027] The first audio sub-signal and the second audio sub-signal are multiplied by a preset window function to obtain the third audio sub-signal corresponding to the first audio sub-signal and the fourth audio sub-signal corresponding to the second audio sub-signal.

[0028] A forward time-frequency transformation is performed on the third audio sub-signal and the fourth audio sub-signal to obtain a first spectrum corresponding to the third audio sub-signal and a second spectrum corresponding to the fourth audio sub-signal;

[0029] A first audio feature of the first audio is extracted from the first spectrum, and a second audio feature of the second audio is extracted from the second spectrum.

[0030] Optionally, the third audio sub-signal includes a plurality of first audio points, and the fourth audio sub-signal includes a plurality of second audio points. The forward time-frequency transformation of the third and fourth audio sub-signals includes:

[0031] The third audio sub-signal and the fourth audio sub-signal are converted from the time domain to the time-frequency domain to obtain the first real part and the first imaginary part corresponding to the first audio point, and the second real part and the second imaginary part corresponding to the second audio point;

[0032] The first spectrum corresponding to the third audio sub-signal is determined based on the first real part and the first imaginary part, and the second spectrum corresponding to the fourth audio sub-signal is determined based on the second real part and the second imaginary part.

[0033] Optionally, converting the frequency domain of the first noise reduction feature and the second noise reduction feature includes:

[0034] The first frequency domain information of the first noise reduction feature is obtained based on the first spectrum, and the second frequency domain information of the second noise reduction feature is obtained based on the second spectrum;

[0035] Perform an inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain the first time domain frame of the first noise reduction feature and the second time domain frame of the second noise reduction feature;

[0036] The frame shift is obtained when the first audio and the second audio are framed, and the frame shift is used to characterize the overlap distance between two consecutive frames;

[0037] A first superimposed signal is obtained by superimposing the first time-domain frame according to the frame shift, and the first amplitude and first phase of the first superimposed signal are adjusted to obtain a first noise-reduced signal;

[0038] The second superimposed signal is obtained by superimposing the second time-domain frame according to the frame shift, and the second amplitude and second phase of the second superimposed signal are adjusted to obtain the second noise-reduced signal.

[0039] According to another aspect of the present invention, a training apparatus for a speech enhancement model is provided, the apparatus comprising:

[0040] An audio acquisition module is used to acquire audio data, wherein the audio data includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise;

[0041] The feature extraction module is used to extract the first audio features of the first audio and the second audio features of the second audio.

[0042] The first noise reduction module is used to perform noise reduction on the first audio feature and the second audio feature using a preset first large model to obtain the first noise reduction feature of the first audio and the second noise reduction feature of the second audio.

[0043] The time-frequency conversion module is used to convert the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain the first noise reduction signal corresponding to the first audio and the second noise reduction signal corresponding to the second audio.

[0044] The loss calculation module is used to calculate the loss function between the first denoised signal and the second denoised signal;

[0045] The model generation module is used to adjust the enhancement parameters of the first large model using the loss function to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.

[0046] Optionally, the audio acquisition module includes:

[0047] An audio separation module is used to separate a first audio signal containing ambient noise and a second audio signal not containing ambient noise from the audio data;

[0048] The sample dataset construction module is used to construct a negative sample dataset using the first audio containing environmental noise, and to construct a positive sample dataset using the second audio that does not contain environmental noise, wherein the second audio in the positive sample dataset corresponds one-to-one with the first audio in the negative sample dataset.

[0049] Optionally, after the audio acquisition module, the device includes:

[0050] The first noise reduction frequency generation module is used to input the first audio in the negative sample dataset into the preset second large model for noise reduction, and obtain the first noise reduction frequency output by the second large model.

[0051] The second audio determination module is used to determine the second audio corresponding to the first audio from the positive sample dataset;

[0052] The first loss function generation module is used to compare the first noise-reduced frequency and the second audio to obtain a comparison result, and generate a first loss function based on the comparison result;

[0053] The first major model generation module is used to adjust the denoising parameters of the second major model based on the first loss function to generate the first major model.

[0054] Optionally, the loss calculation module is specifically used to compare the first denoised signal and the second denoised signal to obtain a second comparison result, and generate a second loss function based on the comparison result;

[0055] The model generation module is specifically used to adjust the enhancement parameters of the first large model using the second loss function.

[0056] Optionally, the feature extraction module includes:

[0057] An audio framing module is used to segment the first audio and the second audio into frames to obtain several frames of the first audio segment and several frames of the second audio segment.

[0058] A first signal conversion module is used to convert the first audio segment into a first audio sub-signal and the second audio segment into a second audio sub-signal;

[0059] The first signal conversion module is used to multiply the first audio sub-signal and the second audio sub-signal by a preset window function respectively to obtain the third audio sub-signal corresponding to the first audio sub-signal and the fourth audio sub-signal corresponding to the second audio sub-signal;

[0060] A forward time-frequency conversion module is used to perform forward time-frequency conversion on the third audio sub-signal and the fourth audio sub-signal to obtain a first spectrum corresponding to the third audio sub-signal and a second spectrum corresponding to the fourth audio sub-signal;

[0061] The feature extraction submodule is used to extract a first audio feature from the first spectrum and a second audio feature from the second spectrum.

[0062] Optionally, the third audio sub-signal includes several first audio points, the fourth audio sub-signal includes several second audio points, and the forward time-frequency conversion module includes:

[0063] The time-domain conversion module is used to convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain to obtain the first real part and the first imaginary part corresponding to the first audio point, and the second real part and the second imaginary part corresponding to the second audio point.

[0064] The spectrum determination module is used to determine the first spectrum corresponding to the third audio sub-signal based on the first real part and the first imaginary part, and to determine the second spectrum corresponding to the fourth audio sub-signal based on the second real part and the second imaginary part.

[0065] Optionally, the time-frequency conversion module includes:

[0066] The frequency domain information acquisition module is used to acquire first frequency domain information of the first noise reduction feature based on the first spectrum, and to acquire second frequency domain information of the second noise reduction feature based on the second spectrum.

[0067] The inverse time-frequency transformation module is used to perform inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain the first time domain frame of the first noise reduction feature and the second time domain frame of the second noise reduction feature;

[0068] A frame shift acquisition module is used to acquire the frame shift when the first audio and the second audio are framed, wherein the frame shift is used to characterize the overlap distance between two consecutive frames.

[0069] The first noise reduction signal generation module is used to obtain a first superimposed signal by superimposing the first time domain frame according to the frame shift, and to adjust the first amplitude and first phase of the first superimposed signal to obtain a first noise reduction signal;

[0070] The second noise reduction signal generation module is used to obtain a second superimposed signal by superimposing the second time-domain frame according to the frame shift, and to adjust the second amplitude and second phase of the second superimposed signal to obtain a second noise reduction signal.

[0071] According to another aspect of the present invention, an electronic device is also provided, comprising:

[0072] processor;

[0073] Memory used to store the processor's executable instructions;

[0074] The processor is configured to execute the instructions to implement the training method for the speech enhancement model as described above.

[0075] According to another aspect of the present invention, a readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the training method for the speech enhancement model as described above.

[0076] The speech enhancement model training method provided by this invention involves acquiring audio data including at least a first audio signal containing environmental noise and a second audio signal not containing environmental noise; extracting first audio features from the first audio signal and second audio features from the second audio signal; using a pre-set first large model to denoise the first and second audio features, obtaining first denoised features of the first audio signal and second denoised features of the second audio signal; converting the frequency domains of the first and second denoised features to obtain a first denoised signal corresponding to the first audio signal and a second denoised signal corresponding to the second audio signal; calculating a loss function between the first and second denoised signals; and adjusting the enhancement parameters of the first large model using the loss function to generate a speech enhancement model for eliminating environmental noise in the audio data. This invention simplifies the denoising process by extracting features from the audio data, thereby improving denoising efficiency. Furthermore, by adjusting the enhancement parameters of the first large model using the loss function, it reduces or eliminates edge effects in the time or frequency domain, making the signal smoother, thus generating a mature speech enhancement model.

[0077] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0078] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0079] Figure 1 This is a flowchart illustrating the steps of a training method for a speech enhancement model provided in an embodiment of the present invention.

[0080] Figure 2 This is a schematic diagram of the structure of a training device for a speech enhancement model provided in an embodiment of the present invention;

[0081] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of the present invention to facilitate a better understanding of the invention. However, the technical solutions claimed in the present invention can be implemented even without these technical details and with various changes and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.

[0083] Currently, most existing speech denoising models on the market utilize adaptive filter technology, which shows good performance in strong noise backgrounds. In particular, adaptive filters based on adaptive noise cancellation enhance the speech signal by adaptively adjusting the filter parameters to adapt to the constantly changing noise environment. However, traditional speech denoising models typically have very large computational requirements and cannot achieve efficient noise elimination in weak noise backgrounds, which seriously affects the accuracy of the model and thus fails to achieve good speech enhancement results.

[0084] Based on this, the present invention proposes a training method, apparatus, electronic device, and storage medium for a speech enhancement model, solving the problems of traditional speech denoising models having large computational requirements and being unable to achieve efficient noise removal in weak background noise, severely affecting the model's accuracy and thus failing to achieve good speech enhancement results. (Refer to...) Figure 1 The diagram illustrates a flowchart of the training method for a speech enhancement model provided in an embodiment of the present invention. The method may include:

[0085] Step 101: Collect audio data. The audio data includes at least a first audio recording containing ambient noise and a second audio recording not containing ambient noise.

[0086] This invention allows for the pre-setting of several audio data collection scenarios, such as offices, streets, coffee shops, and homes. Office audio data can include human voices, keyboard sounds, mouse clicks, and air conditioner noise; street audio data can include human voices and traffic noise; coffee shop audio data can include human voices, coffee machine sounds, and music; and home audio data can include human voices, pet noises, television sounds, and appliance sounds. Collecting audio data within these pre-set scenarios allows for more precise processing of different types of background noise.

[0087] When collecting audio data, it's crucial to ensure the same recording equipment and parameter settings are used throughout the process to avoid the impact of equipment differences on the data. Detailed annotation of the collected audio data is also essential, including information such as environment type, noise type, and speech content, for subsequent analysis and use. Furthermore, a sufficient amount of audio data needs to be collected during model training. For example, 100 audio segments could be collected for each type of environmental noise, with each segment lasting 10 seconds.

[0088] Acquiring a first audio recording containing ambient noise and a second audio recording without ambient noise are key steps in constructing a high-quality audio dataset. This invention provides effective data support for tasks such as noise suppression, speech enhancement, and speech recognition by appropriately selecting the acquisition environment, setting up the recording equipment, labeling the data, and performing subsequent processing and analysis.

[0089] Since the collected audio data contains both normal audio and environmental noise, it is necessary to purify the collected audio data after collection. The specific steps are as follows:

[0090] Separate the first audio audio containing ambient noise and the second audio audio not containing ambient noise from the audio data;

[0091] A negative sample dataset is constructed using a first audio audio that contains environmental noise, and a positive sample dataset is constructed using a second audio audio that does not contain environmental noise. The second audio audio in the positive sample dataset corresponds one-to-one with the first audio audio in the negative sample dataset.

[0092] The first audio signal from the negative sample dataset is input into a pre-set second large model for noise reduction, resulting in the first denoised audio signal output by the second large model. In one embodiment of the invention, the second large model is the initial speech denoising model.

[0093] The second audio corresponding to the first audio is determined from the positive sample dataset. Since the positive sample dataset and the negative sample dataset of this invention are in one-to-one correspondence, the second audio corresponding to the first audio can be determined from the positive sample dataset. After determining the second audio, the first noise-reduced frequency and the second audio are compared to obtain a comparison result, and a first loss function is generated based on the comparison result. The second audio is audio data separated from the original audio data that does not contain environmental noise. That is, the second audio can be regarded as pure audio, and the first noise-reduced frequency is infinitely close to pure audio. Therefore, the first loss function between the second audio and the first noise-reduced frequency can be calculated, and the noise reduction capability of the second model is judged based on the convergence degree of the first loss function. The more convergent the first loss function is, the better the noise reduction capability of the second model is.

[0094] After calculating the first loss function, the denoising parameters of the second model are adjusted based on the first loss function. The second model with adjusted denoising parameters is then used to denoise the first audio in the negative sample dataset again. The newly denoised audio is compared with the second audio corresponding to the first audio to generate a new loss function. The convergence of the new loss function is evaluated. When the convergence of the new loss function reaches the expected value, the first model is generated.

[0095] Step 102: Extract the first audio features of the first audio and the second audio features of the second audio.

[0096] This invention requires feature extraction from audio data. Since a first audio with environmental noise and a second audio without environmental noise are separated from the audio data after the audio data is collected in step 101, the first audio features of the first audio and the second audio features of the second audio can be extracted. This invention does not specifically limit the method of feature extraction.

[0097] Step 103: Use a preset first large model to denoise the first audio feature and the second audio feature to obtain the first denoised feature of the first audio and the second denoised feature of the second audio.

[0098] Compared with the initial speech denoising model, the first large model generated by the aforementioned steps greatly improves the accuracy of speech denoising. However, the structure of the first large model is relatively complex and requires more computation time. Therefore, this invention performs knowledge distillation on the first large model to generate a more concise speech enhancement model.

[0099] First, the first large model is used to denoise the first audio feature and the second audio feature to obtain the first audio feature of the first audio and the second audio feature of the second audio. The specific steps include the following:

[0100] The first and second audio audio segments are framed to obtain several frames of the first audio segment and several frames of the second audio segment. First, the continuous audio data is divided into several short audio segments so that each frame can be processed independently. Before framing, the frame length and frame shift can be set. For example, the frame length can be set to 25–50 ms, and the frame shift is usually 50% of the frame length, which can be set to 12.5–25 ms.

[0101] The first audio segment is converted into a first audio sub-signal, and the second audio segment is converted into a second audio sub-signal. For each frame of audio segment, its time-domain signal is extracted, thereby converting the audio segment into the corresponding audio sub-signal. That is, the first time-domain signal of the first audio segment is extracted, thereby converting the first audio segment into the corresponding first audio sub-signal, and the second time-domain signal of the second audio segment is extracted, thereby converting the second audio segment into the corresponding second audio sub-signal.

[0102] The first and second audio sub-signals are multiplied by a preset window function to obtain the third audio sub-signal corresponding to the first audio sub-signal and the fourth audio sub-signal corresponding to the second audio sub-signal. This invention can reduce edge effects between the first and second audio sub-signals in each frame by weighting them using a window function. Common window functions include the Hanning Window, Hamming Window, Blackman Window, and Kaiser Window.

[0103] The formula for calculating the size of a Haiming window is:

[0104]

[0105] Where w1(n) is the value of the Hamming window, N is the length of the window, and n is a sample point within the window.

[0106] The formula for calculating the size of a Hanning window is:

[0107]

[0108] Where w1(n) is the value of the Hanning window, N is the length of the window, and n is a sample point within the window.

[0109] A forward time-frequency transform is performed on the third and fourth audio sub-signals to obtain the first spectrum corresponding to the third audio sub-signal and the second spectrum corresponding to the fourth audio sub-signal. The forward time-frequency transform of this invention refers to converting the signal from the time domain to the time-frequency domain. When performing the forward time-frequency transform on the third and fourth audio sub-signals, noisy audio characteristics can be obtained by choosing any of the following methods: Fourier transform, Laplace transform, z-transform, or discrete cosine transform. For example, the Fourier transform calculation formula is as follows:

[0110]

[0111] Where x(h) is the signal in the time domain, x(k) is the signal in the frequency domain, H is the length of the signal, and j is the imaginary unit.

[0112] First audio features of a first audio signal are extracted from a first spectrum, and second audio features of a second audio signal are extracted from a second spectrum. After obtaining the first and second spectra through forward time-frequency conversion, feature extraction can be performed on the first and second spectra to obtain the first audio features of the first audio signal and the second audio features of the second audio signal, respectively. The first and second audio features are essentially the spectral features of the first and second audio signals, and the spectral features of this invention include at least amplitude and phase.

[0113] In one embodiment of the present invention, the third audio sub-signal includes a plurality of first audio points, and the fourth audio sub-signal includes a plurality of second audio points. A forward time-frequency transformation is performed on the third audio sub-signal and the fourth audio sub-signal, including:

[0114] Transform the third and fourth audio sub-signals from the time domain to the time-frequency domain to obtain the first real and first imaginary parts corresponding to the first audio point, and the second real and second imaginary parts corresponding to the second audio point. The amplitude and phase of the entire spectrum can then be calculated using the real and imaginary parts, as shown in the following formula:

[0115]

[0116] Therefore, for the third audio sub-signal:

[0117]

[0118] For the fourth audio sub-signal:

[0119]

[0120] The first spectrum corresponding to the third audio sub-signal is determined based on the first real part and the first imaginary part, and the second spectrum corresponding to the fourth audio sub-signal is determined based on the second real part and the second imaginary part.

[0121] Step 104: Convert the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain the first noise reduction signal corresponding to the first audio and the second noise reduction signal corresponding to the second audio. The specific steps are as follows:

[0122] First frequency domain information of the first noise reduction feature is obtained based on the first spectrum, and second frequency domain information of the second noise reduction feature is obtained based on the second spectrum;

[0123] The first frequency domain information and the second frequency domain information are subjected to inverse time-frequency transformation to obtain the first time domain frame of the first noise reduction feature and the second time domain frame of the second noise reduction feature. The first time domain frame and the second time domain frame represent the signal of the first noise reduction feature in the time domain and the signal of the second noise reduction feature in the time domain, respectively.

[0124] Obtain the frame shift when the first and second audio are framed. The frame shift is used to characterize the overlap distance between two consecutive frames and is usually half the frame length.

[0125] The denoised signal is obtained by superimposing time-domain frames based on frame shifts, and the amplitude and phase of the denoised signal are adjusted. The specific superposition process in this invention is as follows:

[0126] A first superimposed signal is obtained by superimposing a first time-domain frame onto the previous frame based on a frame shift. The amplitude and phase of the first superimposed signal are then adjusted to obtain a first denoised signal. This invention superimposes a first time-domain frame onto the signal of the previous frame based on a frame shift to obtain a first superimposed signal. The amplitude and phase of the first superimposed signal are then adjusted to ensure a smooth signal transition. For example, the amplitude and phase can be adjusted through weighted averaging or filtering.

[0127] A second superimposed signal is obtained by superimposing a second time-domain frame based on the frame shift, and the second amplitude and second phase of the second superimposed signal are adjusted to obtain a second noise-reduced signal. Similarly, the second time-domain frame is superimposed on the signal of the previous frame based on the frame shift to obtain a second superimposed signal, and the amplitude and phase of the second superimposed signal are adjusted to ensure a smooth transition of the signal.

[0128] Step 105: Calculate the loss function between the first denoised signal and the second denoised signal. The specific steps are as follows:

[0129] The same number of sampling points are selected from both the first and second denoised signals. A second comparison result is obtained by comparing the first and second denoised signals based on these sampling points. A second loss function is then generated based on this comparison result. The formula for calculating the second loss function is as follows:

[0130]

[0131] Where MSE is the error value between the first and second denoised signals, R is the number of sampling points in the first and second denoised signals, and y i Y is the noise value corresponding to the second sampling point on the second noise-reduced signal. i It is the noise value corresponding to the first acquisition point on the first noise-reduced signal.

[0132] In another embodiment of the present invention, the second loss function can also be calculated using the following formula:

[0133] FL(P t )=-α t ·(1-P t ) γ ·log(P t )

[0134] Among them, P tIt is the predicted probability for a certain category; α t γ is a factor that balances positive and negative samples, used to adjust the weights of the positive and negative sample datasets; γ is an adjustment factor that can also adjust the weights of the positive and negative sample datasets.

[0135] Step 106: Adjust the enhancement parameters of the first model using the loss function to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.

[0136] After generating the second loss function, the enhancement parameters of the first model can be adjusted using the second loss function to generate a speech enhancement model.

[0137] After generating the speech enhancement model, the audio to be processed that needs noise reduction can be input into the speech enhancement model and processed using the speech enhancement model of the present invention to obtain the target audio.

[0138] The speech enhancement model training method provided by this invention involves acquiring audio data including at least a first audio signal containing environmental noise and a second audio signal not containing environmental noise; extracting first audio features from the first audio signal and second audio features from the second audio signal; using a pre-set first large model to denoise the first and second audio features, obtaining first denoised features of the first audio signal and second denoised features of the second audio signal; converting the frequency domains of the first and second denoised features to obtain a first denoised signal corresponding to the first audio signal and a second denoised signal corresponding to the second audio signal; calculating a loss function between the first and second denoised signals; and adjusting the enhancement parameters of the first large model using the loss function to generate a speech enhancement model for eliminating environmental noise in the audio data. This invention simplifies the denoising process by extracting features from the audio data, thereby improving denoising efficiency. Furthermore, by adjusting the enhancement parameters of the first large model using the loss function, it reduces or eliminates edge effects in the time or frequency domain, making the signal smoother, thus generating a mature speech enhancement model.

[0139] Furthermore, refer to Figure 2 The diagram illustrates a structural schematic of a training device for a speech enhancement model provided in an embodiment of the present invention. The device may include:

[0140] Audio acquisition module 201 is used to acquire audio data, which includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise;

[0141] Feature extraction module 202 is used to extract first audio features of the first audio and second audio features of the second audio.

[0142] The first noise reduction module 203 is used to perform noise reduction on the first audio feature and the second audio feature using a preset first large model to obtain the first noise reduction feature of the first audio and the second noise reduction feature of the second audio.

[0143] The time-frequency conversion module 204 is used to convert the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain the first noise reduction signal corresponding to the first audio and the second noise reduction signal corresponding to the second audio.

[0144] The loss calculation module 205 is used to calculate the loss function between the first denoised signal and the second denoised signal;

[0145] The model generation module 206 is used to adjust the enhancement parameters of the first model using a loss function to generate a speech enhancement model, which is used to remove environmental noise in the audio data.

[0146] In one embodiment of the present invention, the audio acquisition module 201 includes:

[0147] An audio separation module is used to separate a first audio signal containing ambient noise and a second audio signal not containing ambient noise from audio data;

[0148] The sample dataset construction module is used to construct a negative sample dataset using a first audio file containing environmental noise, and to construct a positive sample dataset using a second audio file that does not contain environmental noise. The second audio file in the positive sample dataset corresponds one-to-one with the first audio file in the negative sample dataset.

[0149] In one embodiment of the present invention, the device includes, after the audio acquisition module 201:

[0150] The first noise reduction frequency generation module is used to input the first audio in the negative sample dataset into the preset second large model for noise reduction, and obtain the first noise reduction frequency output by the second large model.

[0151] The second audio determination module is used to determine the second audio corresponding to the first audio from the positive sample dataset;

[0152] The first loss function generation module is used to compare the first noise-reduced frequency and the second audio to obtain a comparison result, and generate the first loss function based on the comparison result;

[0153] The first major model generation module is used to adjust the denoising parameters of the second major model based on the first loss function to generate the first major model.

[0154] In one embodiment of the present invention, the loss calculation module 205 is specifically used to compare the first denoised signal and the second denoised signal to obtain a second comparison result, and generate a second loss function based on the comparison result;

[0155] The model generation module 206 is specifically used to adjust the enhancement parameters of the first large model using the second loss function.

[0156] In one embodiment of the present invention, the feature extraction module 202 includes:

[0157] The audio framing module is used to segment the first audio and the second audio into frames, resulting in several frames of the first audio segment and several frames of the second audio segment.

[0158] The first signal conversion module is used to convert the first audio segment into a first audio sub-signal and to convert the second audio segment into a second audio sub-signal.

[0159] The first signal conversion module is used to multiply the first audio sub-signal and the second audio sub-signal by a preset window function respectively to obtain the third audio sub-signal corresponding to the first audio sub-signal and the fourth audio sub-signal corresponding to the second audio sub-signal.

[0160] The forward time-frequency conversion module is used to perform forward time-frequency conversion on the third audio sub-signal and the fourth audio sub-signal to obtain the first spectrum corresponding to the third audio sub-signal and the second spectrum corresponding to the fourth audio sub-signal;

[0161] The feature extraction submodule is used to extract a first audio feature from a first spectrum and a second audio feature from a second spectrum.

[0162] In one embodiment of the present invention, the third audio sub-signal includes a plurality of first audio points, the fourth audio sub-signal includes a plurality of second audio points, and the forward time-frequency conversion module includes:

[0163] The time-domain conversion module is used to convert the third audio sub-signal and the fourth audio sub-signal from the time domain to the time-frequency domain, so as to obtain the first real part and the first imaginary part corresponding to the first audio point, and the second real part and the second imaginary part corresponding to the second audio point.

[0164] The spectrum determination module is used to determine the first spectrum corresponding to the third audio sub-signal based on the first real part and the first imaginary part, and to determine the second spectrum corresponding to the fourth audio sub-signal based on the second real part and the second imaginary part.

[0165] In one embodiment of the present invention, the time-frequency conversion module 204 includes:

[0166] The frequency domain information acquisition module is used to acquire first frequency domain information of the first noise reduction feature based on the first spectrum, and to acquire second frequency domain information of the second noise reduction feature based on the second spectrum.

[0167] The inverse time-frequency transformation module is used to perform inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain the first time domain frame of the first noise reduction feature and the second time domain frame of the second noise reduction feature.

[0168] The frame shift acquisition module is used to acquire the frame shift when the first audio and the second audio are framed. The frame shift is used to characterize the overlap distance between two consecutive frames.

[0169] The first noise reduction signal generation module is used to obtain a first superimposed signal based on the first time domain frame superimposed by the frame shift, and to adjust the first amplitude and first phase of the first superimposed signal to obtain the first noise reduction signal;

[0170] The second noise reduction signal generation module is used to obtain a second superimposed signal based on the frame shift and superposition of a second time-domain frame, and to adjust the second amplitude and second phase of the second superimposed signal to obtain a second noise reduction signal.

[0171] Reference Figure 3 The present invention also provides an electronic device, such as... Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0172] Processor 301, memory 303 for storing processor-executable instructions;

[0173] The processor 301 is configured to execute instructions to implement the training method of the speech enhancement model described above:

[0174] Collect audio data, which includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise;

[0175] Extract the first audio features of the first audio and the second audio features of the second audio.

[0176] A pre-set first large model is used to denoise the first audio feature and the second audio feature to obtain the first denoised feature of the first audio and the second denoised feature of the second audio.

[0177] By converting the frequency domain of the first noise reduction feature and the second noise reduction feature, a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio are obtained.

[0178] Calculate the loss function between the first and second denoised signals;

[0179] The enhancement parameters of the first model are adjusted using a loss function to generate a speech enhancement model, which is used to remove environmental noise from the audio data.

[0180] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0181] The communication interface is used for communication between the aforementioned terminal and other devices.

[0182] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0183] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0184] In another embodiment of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the training method of any of the speech enhancement models in the above embodiments.

[0185] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0186] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0187] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0188] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A training method for a speech enhancement model, characterized in that, The method includes: Collect audio data, wherein the audio data includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise; Extract the first audio features of the first audio and the second audio features of the second audio; A preset first large model is used to denoise the first audio feature and the second audio feature to obtain the first denoised feature of the first audio and the second denoised feature of the second audio. By converting the frequency domains of the first noise reduction feature and the second noise reduction feature, a first noise reduction signal corresponding to the first audio and a second noise reduction signal corresponding to the second audio are obtained; Calculate the loss function between the first denoised signal and the second denoised signal; The enhancement parameters of the first large model are adjusted using the loss function to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.

2. The method according to claim 1, characterized in that, After acquiring audio data, the method includes: Separate a first audio audio containing ambient noise and a second audio audio not containing ambient noise from the audio data; A negative sample dataset is constructed using the first audio containing environmental noise, and a positive sample dataset is constructed using the second audio that does not contain environmental noise, wherein the second audio in the positive sample dataset corresponds one-to-one with the first audio in the negative sample dataset.

3. The method according to claim 2, characterized in that, After acquiring audio data, the method includes: The first audio from the negative sample dataset is input into a preset second large model for noise reduction, and the first noise-reduced audio frequency output by the second large model is obtained. Determine the second audio corresponding to the first audio from the positive sample dataset; The first noise-reduced audio frequency and the second audio frequency are compared to obtain a comparison result, and a first loss function is generated based on the comparison result; The denoising parameters of the second largest model are adjusted based on the first loss function to generate the first largest model.

4. The method according to claim 3, characterized in that, The calculation of the loss function between the first denoised signal and the second denoised signal includes: A second comparison result is obtained by comparing the first denoised signal with the second denoised signal, and a second loss function is generated based on the comparison result; Adjusting the enhancement parameters of the first large model using the loss function includes: The enhancement parameters of the first large model are adjusted using the second loss function.

5. The method according to claim 1, characterized in that, The extraction of the first audio feature of the first audio and the second audio feature of the second audio includes: The first audio and the second audio are divided into frames to obtain several frames of the first audio segment and several frames of the second audio segment. The first audio segment is converted into a first audio sub-signal, and the second audio segment is converted into a second audio sub-signal; The first audio sub-signal and the second audio sub-signal are multiplied by a preset window function to obtain the third audio sub-signal corresponding to the first audio sub-signal and the fourth audio sub-signal corresponding to the second audio sub-signal. A forward time-frequency transformation is performed on the third audio sub-signal and the fourth audio sub-signal to obtain a first spectrum corresponding to the third audio sub-signal and a second spectrum corresponding to the fourth audio sub-signal; A first audio feature of the first audio is extracted from the first spectrum, and a second audio feature of the second audio is extracted from the second spectrum.

6. The method according to claim 5, characterized in that, The third audio sub-signal contains several first audio points, and the fourth audio sub-signal contains several second audio points. The forward time-frequency transformation of the third and fourth audio sub-signals includes: The third audio sub-signal and the fourth audio sub-signal are converted from the time domain to the time-frequency domain to obtain the first real part and the first imaginary part corresponding to the first audio point, and the second real part and the second imaginary part corresponding to the second audio point; The first spectrum corresponding to the third audio sub-signal is determined based on the first real part and the first imaginary part, and the second spectrum corresponding to the fourth audio sub-signal is determined based on the second real part and the second imaginary part.

7. The method according to claim 5, characterized in that, Converting the frequency domains of the first noise reduction feature and the second noise reduction feature includes: The first frequency domain information of the first noise reduction feature is obtained based on the first spectrum, and the second frequency domain information of the second noise reduction feature is obtained based on the second spectrum; Perform an inverse time-frequency transformation on the first frequency domain information and the second frequency domain information to obtain the first time domain frame of the first noise reduction feature and the second time domain frame of the second noise reduction feature; The frame shift is obtained when the first audio and the second audio are framed, and the frame shift is used to characterize the overlap distance between two consecutive frames; A first superimposed signal is obtained by superimposing the first time-domain frame according to the frame shift, and the first amplitude and first phase of the first superimposed signal are adjusted to obtain a first noise-reduced signal; The second superimposed signal is obtained by superimposing the second time-domain frame according to the frame shift, and the second amplitude and second phase of the second superimposed signal are adjusted to obtain the second noise-reduced signal.

8. A training device for a speech enhancement model, characterized in that, The device includes: An audio acquisition module is used to acquire audio data, wherein the audio data includes at least a first audio audio containing ambient noise and a second audio audio not containing ambient noise; The feature extraction module is used to extract the first audio features of the first audio and the second audio features of the second audio. The first noise reduction module is used to perform noise reduction on the first audio feature and the second audio feature using a preset first large model to obtain the first noise reduction feature of the first audio and the second noise reduction feature of the second audio. The time-frequency conversion module is used to convert the frequency domain of the first noise reduction feature and the second noise reduction feature to obtain the first noise reduction signal corresponding to the first audio and the second noise reduction signal corresponding to the second audio. The loss calculation module is used to calculate the loss function between the first denoised signal and the second denoised signal; The model generation module is used to adjust the enhancement parameters of the first large model using the loss function to generate a speech enhancement model, which is used to eliminate environmental noise in the audio data.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the training method of the speech enhancement model as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, which, when executed by a processor, implements the training method for the speech enhancement model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image enhancement method and device

    CN112258401A

  • Voice noise reduction model training method and device, storage medium and electronic device

    CN114974283A