Speech enhancement model training method, speech enhancement method, device, medium and product
By using a mixed signal of speech and noise as training labels, the problem of oversuppression in speech enhancement models during training is solved, speech clarity and quality are improved, and the robustness of the model is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOERTEK INC
- Filing Date
- 2024-12-04
- Publication Date
- 2026-07-21
AI Technical Summary
Existing speech enhancement models tend to oversuppress speech during training, leading to a decrease in clarity and quality, and failing to effectively preserve subtle features and low-energy parts of speech.
A mixed signal of speech and noise is used as the training label. The speech enhancement model is trained using the input data and label data in the training dataset to ensure that the model retains some noise while maintaining near-clean speech, thereby reducing the amount of speech suppression.
It improves speech clarity and quality, reduces the likelihood of the model incorrectly removing subtle speech features or low-energy parts during training, and enhances the model's robustness and noise reduction effect.
Smart Images

Figure CN119479670B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and in particular to a speech enhancement model training method, speech enhancement method, device, medium and product. Background Technology
[0002] Speech enhancement (SE) refers to techniques for recovering useful speech signals from interfered speech signals, thereby improving the quality and clarity of the speech signal. With the rapid development of communication technologies and smart devices, the demand for clear, high-quality speech is increasing. However, in practical applications, speech signals are often affected by various noises and interferences, reducing speech quality and intelligibility. To address this problem, speech enhancement plays a crucial role in the field of digital signal processing.
[0003] Generally, speech enhancement models can be used for speech enhancement; therefore, training these models is crucial for improving the quality of speech denoising. Traditional speech enhancement model training methods first involve simulating the mixing of clean speech data from a clean speech database and noisy data from a noisy database to obtain a noisy speech database. For any noisy speech in this noisy speech database, there is a corresponding clean speech in the clean speech database. Then, the speech enhancement model is trained using the noisy speech database as input and the clean speech as labels, and the model is optimized based on the loss value. This process is iterated until a given training termination condition is met.
[0004] However, the conventional training method for speech enhancement models always uses clean speech as the training label; that is, the training objective of the speech enhancement model is always to make the output speech as close to clean speech as possible. During this training process, the model may mistakenly treat some subtle speech features or low-energy speech parts as noise to be removed, leading to over-suppression of speech and further degrading speech clarity and quality.
[0005] Therefore, how to reduce the amount of speech suppression by the speech enhancement model in order to improve the clarity and quality of speech is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] The main purpose of this application is to provide a speech enhancement model training method, speech enhancement method, device, medium and product, which aims to solve the technical problem of how to reduce the amount of speech suppression by the speech enhancement model in order to improve the clarity and quality of speech.
[0007] To achieve the above objectives, this application provides a method for training a speech enhancement model, which includes the following steps:
[0008] Obtain a training dataset and a speech enhancement model to be trained, wherein the training dataset includes at least one input data and label data corresponding to each input data, and at least one label data is a mixed signal containing speech and noise.
[0009] The speech enhancement model is trained by using the input data as input and the label data as labels, and the trained speech enhancement model is obtained.
[0010] In one embodiment, prior to the step of obtaining the training dataset and the speech enhancement model to be trained, the method further includes:
[0011] Obtain the original dataset, wherein the original dataset includes speech signals and noise signals;
[0012] The input data is obtained by mixing the speech signal and the noise signal at at least one first signal-to-noise ratio.
[0013] For each piece of input data, a second signal-to-noise ratio is determined based on a first signal-to-noise ratio of the input data, and the speech signal and the noise signal are mixed with the second signal-to-noise ratio to obtain the tag data corresponding to the input data, wherein the second signal-to-noise ratio is greater than or equal to the first signal-to-noise ratio;
[0014] The training dataset is obtained by combining the input data and the label data.
[0015] In one embodiment, the step of determining the second signal-to-noise ratio based on the first signal-to-noise ratio of the input data includes:
[0016] Obtain the mapping table, and based on the mapping table, find the second signal-to-noise ratio corresponding to the first signal-to-noise ratio of the input data; or...
[0017] The first signal-to-noise ratio corresponding to the input data is input into a preset mapping function to obtain the second signal-to-noise ratio.
[0018] In one embodiment, prior to the step of obtaining the mapping table, the method further includes:
[0019] At least one preset test data is input into a preset conventional speech enhancement model to obtain the output result corresponding to each preset test data, wherein the preset conventional speech enhancement model is the speech enhancement model pre-trained with speech signals as labels;
[0020] Obtain the input signal-to-noise ratio range to which the signal-to-noise ratio of each preset test data belongs;
[0021] For each input signal-to-noise ratio (SNR) interval, calculate the average value of the output results corresponding to all preset test data belonging to the input SNR interval, determine the label SNR interval based on the average value, and configure a correspondence between the input SNR interval and the label SNR interval, wherein the lower limit of the label SNR interval is greater than or equal to the average value.
[0022] By combining the correspondence between each input signal-to-noise ratio interval and each label signal-to-noise ratio interval, a mapping table is obtained.
[0023] In one embodiment, the raw dataset further includes room impulse responses, and after the step of obtaining the raw dataset, the method further includes:
[0024] The room impulse response is added to the speech signal to obtain a first speech signal, and the room impulse response is added to the noise signal to obtain a first noise signal;
[0025] The first speech signal is used as a new speech signal, and the first noise signal is used as a new noise signal. Based on the new speech signal and the new noise signal, the step of mixing the speech signal and the noise signal at at least one first signal-to-noise ratio to obtain input data is performed.
[0026] In one embodiment, after determining the second signal-to-noise ratio based on the first signal-to-noise ratio of the input data, the method further includes:
[0027] Obtain the training target of the speech enhancement model;
[0028] If the training objective indicates dereverberation, then the early reverberation in the room impulse response is added to the speech signal to obtain a second speech signal. The second speech signal is used as a new speech signal, and the step of mixing the speech signal and the noise signal with the second signal-to-noise ratio to obtain the label data corresponding to the input data is performed based on the new speech signal.
[0029] Furthermore, to achieve the above objectives, this application also provides a speech enhancement method, which includes the following steps:
[0030] Obtain the original signal to be enhanced and the target speech enhancement model, and input the original signal into the target speech enhancement model to obtain the speech enhancement result;
[0031] The target speech enhancement model is a speech enhancement model trained using the speech enhancement model training method described above.
[0032] In addition, to achieve the above objectives, this application also provides an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the speech enhancement model training method and / or speech enhancement method as described above.
[0033] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium storing a computer program that is executed by a processor to implement the steps of the speech enhancement model training method and / or speech enhancement method as described above.
[0034] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech enhancement model training method and / or speech enhancement method as described above.
[0035] One or more technical solutions proposed in this application have at least the following technical effects:
[0036] A training dataset and a speech enhancement model to be trained are obtained. The training dataset includes at least one input data and label data corresponding to each input data. At least one label data is a mixed signal containing both speech and noise. The speech enhancement model is trained using each input data as input and each label data as label, resulting in a trained speech enhancement model. Thus, this embodiment uses at least one mixed signal containing both speech and noise as the training label for the speech enhancement model, rather than always using clean speech. This allows the training objective of the speech enhancement model to retain some noise while approximating clean speech as much as possible. This reduces the likelihood of the model mistakenly removing subtle speech features or low-energy speech parts as noise during training, thereby reducing the amount of speech suppression by the speech enhancement model and improving speech clarity and quality. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the first embodiment of the speech enhancement model training method of this application;
[0040] Figure 2 This is a schematic diagram of the mapping function curve involved in an embodiment of the speech enhancement model training method of this application;
[0041] Figure 3 This is a schematic diagram of the room impulse response involved in an embodiment of the speech enhancement model training method of this application;
[0042] Figure 4 This is a simplified flowchart illustrating a specific embodiment of the speech enhancement model training method of this application.
[0043] Figure 5 This is a schematic diagram of the device structure of the speech enhancement model training device of this application;
[0044] Figure 6 This is a schematic diagram of the hardware operating environment of the speech enhancement model training method apparatus in the embodiments of this application.
[0045] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0046] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] With the widespread adoption of various smart devices, sound has gradually become one of the primary ways for humans to obtain information. However, most of the sounds humans receive daily contain various types of noise. For example, in remote suburbs, the average ambient noise level is 30 decibels; in more noisy environments, the noise level can reach 77 decibels. Therefore, voice interaction tasks in real-world scenarios face significant challenges.
[0048] Speech enhancement refers to the technique of recovering useful speech signals from distorted speech signals. Traditional signal processing-based speech enhancement methods treat noise reduction as a simple signal processing problem, establishing mathematical and statistical models and often requiring the manual addition of prior information. However, in complex real-world scenarios, the performance of traditional enhancement methods suffers significant degradation due to the large discrepancy between prior assumptions and the actual situation. In recent years, with the development of deep learning technology, deep neural networks have begun to be widely applied in speech enhancement. By using noisy speech-clean speech data pairs for large-scale supervised training, neural networks can learn the patterns of information related to clean speech. With continuous exploration by researchers, many advanced speech enhancement models based on the time domain and time-frequency domain have emerged, such as DCCRN (Deep Complex Convolution Recurrent Network), DPCRN (Dual-Path Convolution Recurrent Network), and Conv-TasNet (Convolutional Time-domain audioseparation Network).
[0049] In the aforementioned studies, regardless of the proportion of noise in the audio, the prediction target was always clean audio. This is often not the optimal target for enhancement models, especially lightweight models. When the audio signal-to-noise ratio (SNR) is high, the model can obtain clean audio by removing very little noise; however, when the SNR is low, the total amount of noise that needs to be removed may exceed the capabilities of lightweight models, causing the model to fail to converge to the optimal state during training. Ultimately, this results in over-suppression or excessive noise residue in the enhancement results. Speech enhancement models often suffer from over-suppression of speech under conditions of low parameter count and low computational resources.
[0050] Based on this, the main solution of this application is: to obtain a training dataset and a speech enhancement model to be trained, wherein the training dataset includes at least one input data and label data corresponding to each input data, and at least one label data is a mixed signal containing both speech and noise; to train the speech enhancement model using each input data as input to the speech enhancement model and each label data as label to the speech enhancement model, thereby obtaining the trained speech enhancement model.
[0051] This application trains the speech enhancement model by using at least one mixed signal of speech and noise as the training label, instead of always using clean speech as the training label. This allows the training objective of the speech enhancement model to retain some noise while being as close to clean speech as possible. This reduces the possibility that the model will mistakenly remove some subtle speech features or low-energy speech parts as noise during training, thereby reducing the amount of speech suppression by the speech enhancement model and improving speech clarity and quality.
[0052] It should be noted that the execution subject of the various embodiments of the speech enhancement model training method of this application can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of realizing the above functions, such as headphones, AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, AR helmets, VR helmets, smart glasses, etc. The various embodiments of the speech enhancement model training method of this application do not impose specific limitations in this regard.
[0053] Based on this, this application proposes a speech enhancement model training method according to the first embodiment, please refer to... Figure 1 As shown, the speech enhancement model training method includes the following steps S10 to S20:
[0054] Step S10: Obtain the training dataset and the speech enhancement model to be trained. The training dataset includes at least one input data and label data corresponding to each input data. At least one label data in each label data is a mixed signal containing speech and noise.
[0055] The speech enhancement model can be a pre-built model used for speech enhancement. Specifically, the speech enhancement model can be a deep learning model, such as the DCCRN model, the DPCRN model, the Conv-TasNet model, etc., but this embodiment does not impose any specific limitations on it.
[0056] It should be noted that, in this embodiment, each input data and each label data in the training dataset can be audio data.
[0057] The training dataset includes at least one input data and the label data corresponding to each input data. It is understood that, in order to improve the training effect of the model, multiple input data and the label data corresponding to each input data can be prepared in advance as the training dataset.
[0058] For each piece of input data, the label data corresponding to that piece of input data can be the expected output corresponding to the input data. In layman's terms, it is the ideal data that the speech enhancement model is expected to output after the input data is input into the speech enhancement model.
[0059] Each input data point can be a mixed signal that combines speech and noise. It should be noted that when there are multiple input data points, there may be one or more input data points that consist only of speech signals or only of noise signals, but there should be at least one input data point that is a mixed signal that combines speech and noise.
[0060] Furthermore, input data consisting only of speech or consisting only of noise is denoted as pure input data, and input data consisting of a mixture of speech and noise is denoted as mixed input data. The total number of pure input data is less than the total number of mixed input data, in order to improve the training effect of the model and enable the model to have a certain processing capability for pure speech signals and pure noise signals, thereby further improving the robustness of the model.
[0061] Furthermore, the first signal-to-noise ratio of the input data is less than or equal to the second signal-to-noise ratio of the corresponding label data, in order to ensure the noise removal effect of the speech enhancement model.
[0062] It should be noted that in this embodiment, speech signal (referred to as speech) specifically refers to sound signals emitted by humans, which can usually be expressed as speech, music, etc. Noise signal (referred to as noise) refers to signals that do not contain useful information and are undesirable signals, which can usually be expressed as environmental noise, thermal noise, electromagnetic interference noise, etc.
[0063] Step S20: Use each of the input data as input to the speech enhancement model and each of the label data as labels to train the speech enhancement model, and obtain the trained speech enhancement model.
[0064] After obtaining the training dataset, each input data point in the training dataset is used as the input to the speech enhancement model, and each label data point is used as the label for the speech enhancement model. It's easy to understand that for each input data point, the corresponding label data is the label used during model training.
[0065] Training termination conditions can be preset. During the training process of the speech enhancement model, if the training termination condition is met, the training of the speech enhancement model ends, and a trained speech enhancement model is obtained. If the training termination condition is not met, the speech enhancement model continues to be iteratively optimized until the training termination condition is met. The training termination condition can be a preset condition, such as reaching a predetermined number of iterations, the loss function value falling below a predetermined threshold, exhausting computing resources, reaching a time limit, or the accuracy exceeding a predetermined threshold. This embodiment does not impose specific restrictions on this.
[0066] It should be noted that the trained speech enhancement model can be the speech enhancement model obtained when the training ends, that is, the model parameters of the trained speech enhancement model are the model parameters obtained when training ends; or it can be the speech enhancement model with the best performance in the iteration process. Specifically, the model parameters obtained in each round of iteration can be recorded, and then the set of target model parameters with the best performance can be selected. The model parameters of the speech enhancement model can be set as the target model parameters to obtain the trained speech enhancement model.
[0067] Specifically, known model evaluation methods can be used to evaluate the model's performance, such as STOI (Short-Time Objective Intelligibility) evaluation and PESQ (Perceptual Evaluation of Speech Quality) evaluation. This implementation does not impose specific restrictions on this, nor will it elaborate on the specific evaluation process.
[0068] Furthermore, rules for evaluating the performance of the optimal model can be set in advance, such as the maximum value among the model performance capabilities being the optimal model performance capability. Based on these preset rules, the optimal model performance capability among all model performance capabilities can be determined, and the model parameters corresponding to the optimal model performance capability can be selected as the target model parameters.
[0069] After obtaining the trained speech enhancement model, it can be deployed to other devices so that other devices can perform speech enhancement processing based on the trained speech enhancement model.
[0070] This embodiment acquires a training dataset and a speech enhancement model to be trained. The training dataset includes at least one input data point and corresponding label data for each input data point. At least one label data point is a mixed signal containing both speech and noise. The speech enhancement model is trained using the input data points as inputs and the label data points as labels, resulting in a trained speech enhancement model. Thus, this embodiment uses at least one mixed signal containing both speech and noise as the training label for the speech enhancement model, rather than always using clean speech. This allows the training objective of the speech enhancement model to retain some noise while approximating clean speech as much as possible. This reduces the likelihood of the model mistakenly removing subtle speech features or low-energy speech components as noise during training, thereby reducing the amount of speech suppression by the speech enhancement model and improving speech clarity and quality.
[0071] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description and will not be repeated hereafter. Furthermore, before the step of obtaining the training dataset and the speech enhancement model to be trained, the method further includes:
[0072] Step A10: Obtain the original dataset, wherein the original dataset includes speech signals and noise signals;
[0073] The original dataset can be a dataset prepared in advance by relevant personnel. It should be noted that the original dataset includes speech signals and noise signals; specifically, the original dataset may include one or more speech signals and one or more noise signals.
[0074] Furthermore, each time, one speech signal and one noise signal can be randomly selected from the original dataset, or one speech signal and one noise signal can be selected from the original dataset based on predetermined rules, without any specific restrictions.
[0075] Step A20: Mix the speech signal and the noise signal at at least one first signal-to-noise ratio to obtain input data;
[0076] It should be noted that each time, one speech signal and one noise signal can be selected from the original dataset to execute steps A20 and A30 until the number of input data items meets the preset threshold. The preset threshold can be any positive value set in advance by relevant personnel, and this embodiment does not impose any specific restrictions on it.
[0077] The first signal-to-noise ratio can be one or more randomly generated signal-to-noise ratios. For example, in one specific embodiment, the first signal-to-noise ratio can be a signal-to-noise ratio randomly selected from the signal-to-noise ratio range of [-20, 20].
[0078] Step A30: For each piece of input data, a second signal-to-noise ratio is determined based on a first signal-to-noise ratio of the input data, and the speech signal and the noise signal are mixed with the second signal-to-noise ratio to obtain the tag data corresponding to the input data, wherein the second signal-to-noise ratio is greater than or equal to the first signal-to-noise ratio;
[0079] The corresponding input data and label data are obtained by mixing the same speech signal and the same noise signal at different signal-to-noise ratios (SNRs). For example, a first SNR A is randomly generated, and a second SNR A is determined based on this first SNR A. The speech signal A and the noise signal are mixed at the first SNR A to obtain input data A. Then, the speech signal A and the noise signal A are mixed at the second SNR A to obtain label data A. Label data A is the label data corresponding to input data A. This process is repeated until the label data corresponding to all input data is obtained.
[0080] It should be noted that an upper limit for the signal-to-noise ratio (SNR) can also be set. The upper limit represents the SNR when the signal consists entirely of speech. If the SNR is at the upper limit, the speech signal can be used as input data or label data. Specifically, if the first SNR is at the upper limit, the speech signal is used as input data; if the second SNR is at the upper limit, the speech signal is used as label data. Similarly, a lower limit for the SNR can also be set. The lower limit represents the SNR when the signal consists entirely of noise. If the SNR is at the lower limit, the noise signal can be used as input data or label data. Specifically, if the first SNR is at the lower limit, the noise signal is used as input data; if the second SNR is at the lower limit, the noise signal is used as label data.
[0081] Furthermore, when the first signal-to-noise ratio is the upper limit, the second signal-to-noise ratio can be equal to the first signal-to-noise ratio; when the first signal-to-noise ratio is not the upper limit, the second signal-to-noise ratio is greater than the first signal-to-noise ratio, so as to ensure the noise reduction effect of the speech enhancement model.
[0082] Furthermore, both the upper and lower limits can be predefined based on actual needs. They can be specific numerical values or custom strings. For example, the string "MAX" can be defined as the upper limit and the string "MIN" as the lower limit.
[0083] Step A40: Combine the input data and the label data to obtain the training dataset.
[0084] In this embodiment, the input data is obtained by mixing the speech signal and the noise signal at a first signal-to-noise ratio (SNR), and the label data is obtained by mixing the speech signal and the noise signal at a second SNR that is greater than or equal to the first SNR. Thus, it is easy to understand that the optimization goal of the speech enhancement model is to increase (or at least not decrease) the signal-to-noise ratio of the signal, thereby ensuring that the speech enhancement model can reduce noise residue and achieve a balance between noise suppression and speech preservation.
[0085] In one possible implementation, the step of determining the second signal-to-noise ratio based on the first signal-to-noise ratio of the input data includes:
[0086] Step B10: Obtain the mapping table, and based on the mapping table, find the second signal-to-noise ratio corresponding to the first signal-to-noise ratio of the input data; or,
[0087] Specifically, this mapping table can be a pre-set mapping table that stores the corresponding signal-to-noise ratio relationships.
[0088] Furthermore, the mapping table can be a mapping table containing multiple pairs of one-to-one correspondences of signal-to-noise ratios. However, considering the continuity of signal-to-noise ratios and their wide range of values, storing the correspondences in units of signal-to-noise ratios would require a large amount of data, resulting in a significant need for storage space. Therefore, as another implementation, the mapping table can be a mapping table containing multiple pairs of one-to-one correspondences of signal-to-noise ratio intervals to reduce data storage. For example, in one specific implementation, the mapping table can be the mapping table shown in Table 1 below, where the input signal-to-noise ratio interval represents the interval to which the signal-to-noise ratio of the input data belongs, the label signal-to-noise ratio interval represents the interval to which the signal-to-noise ratio of the label data belongs, and "Clean Label" represents the upper limit of the signal-to-noise ratio.
[0089] [-20,-15) [-8,-3) [-15,-10) [-3,2) [-10,-5) [3,8) [-5,0) [10,18) [0,20] Clean Label
[0090] Table 1
[0091] Furthermore, when the mapping table contains multiple pairs of one-to-one correspondences of signal-to-noise ratio (SNR) intervals, the target SNR interval to which the input data's SNR belongs can be determined first. Then, the corresponding SNR interval in the mapping table (referred to as the selected SNR interval for ease of subsequent explanation) is searched, and an SNR is selected from the selected SNR interval as the second SNR. For example, in one specific embodiment, an SNR can be randomly selected from the selected SNR interval as the second SNR.
[0092] Step B20: Input the first signal-to-noise ratio corresponding to the input data into a preset mapping function to obtain the second signal-to-noise ratio.
[0093] The preset mapping function is a pre-set mapping function that satisfies the condition that the value after mapping is greater than or equal to the value before mapping. In a preferred embodiment, the preset mapping function is f(snr) = 30 - 7e -(snr+20) Wherein, SNR represents the mapping variable, specifically the first signal-to-noise ratio of the input data in this embodiment. The mapping curve of the preset mapping function is as follows: Figure 2 As shown, Figure 2 The horizontal axis represents the first signal-to-noise ratio, and the vertical axis represents the second signal-to-noise ratio.
[0094] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Furthermore, before the step of obtaining the mapping relationship table, the method further includes:
[0095] Step C10: Input at least one preset test data into a preset conventional speech enhancement model to obtain the output result corresponding to each preset test data, wherein the preset conventional speech enhancement model is the speech enhancement model pre-trained with speech signals as labels;
[0096] The preset conventional model is a speech enhancement model pre-trained with speech signals as labels. In other words, the preset conventional speech enhancement model is a speech enhancement model trained using conventional model training methods.
[0097] The preset test data can be one or more pre-prepared test data sets. Generally, to improve the validity and accuracy of the test results, multiple test data sets are prepared in advance. For example, in one specific implementation, 1000 test data sets with different signal-to-noise ratio ranges are prepared in advance, that is, 1000 preset test data sets are prepared.
[0098] Each test data point is input into a preset conventional speech enhancement model, and the corresponding output result is obtained. It should be noted that the output result here can specifically be the signal-to-noise ratio (SNR) of the speech enhancement model's output data. For example, if preset test data A is input into the preset conventional speech enhancement model, and the output of the preset conventional speech enhancement model is output data A, then the output result corresponding to preset test data A is the SNR of output data A.
[0099] Step C20: Obtain the input signal-to-noise ratio range to which the signal-to-noise ratio of each preset test data belongs;
[0100] It should be noted that one or more input signal-to-noise ratio (SNR) intervals can be predefined. Each input SNR interval is a non-empty set, and different input SNR intervals may not overlap with each other. For example, in one specific implementation, the predefined input SNR intervals are [-20, -15), [-15, -10), [-10, -5), [-5, 0), [0, 5), [5, 10), [10, 20].
[0101] Based on the pre-defined input signal-to-noise ratio (SNR) intervals, the input SNR interval to which each preset prediction data belongs can be obtained. For example, assuming the SNR of a preset test data is 5, and the input SNR interval is as shown in the specific implementation described above, then the input SNR interval to which the preset test data belongs is [5, 10).
[0102] Step C30: For each input signal-to-noise ratio interval, calculate the average value of the output results corresponding to all preset test data belonging to the input signal-to-noise ratio interval, determine the tag signal-to-noise ratio interval based on the average value, and configure a correspondence between the input signal-to-noise ratio interval and the tag signal-to-noise ratio interval, wherein the lower limit of the tag signal-to-noise ratio interval is greater than or equal to the average value.
[0103] For each input signal-to-noise ratio (SNR) interval, calculate the average of the output results corresponding to all preset test data belonging to this input SNR interval. The input SNR interval to which the SNR of the preset test data belongs is also the input SNR interval to which the preset test data belongs. For example, for a certain input SNR interval 1, the predicted test data whose SNR belongs to this input SNR interval are [preset test data A, predicted test data B, preset test data C], and the corresponding output results are [output result A, output result B, output result C], respectively. Then the average value is (output result A + output result B + output result C) / 3.
[0104] The tag signal-to-noise ratio (SNR) range is determined based on the average value. Specifically, the lower limit of the tag SNR range is greater than or equal to the average value. For example, in one specific implementation, as shown in Table 2 below, the tag SNR range is a range centered on the average tag value after an increase of 10 dB, fluctuating by 1 dB above and below it.
[0105]
[0106] Table 2
[0107] Step C40: Combine the correspondence between each input signal-to-noise ratio interval and each label signal-to-noise ratio interval to obtain a mapping table.
[0108] The mapping table stores the correspondence between the input signal-to-noise ratio (SNR) intervals and the tag SNR intervals. For example, assuming the obtained input SNR intervals and tag SNR intervals are as shown in Table 2 above, the mapping table is shown in Table 3.
[0109]
[0110]
[0111] Table 3
[0112] In this embodiment, preset test data is input into a preset conventional speech enhancement model to obtain output results. The label signal-to-noise ratio (SNR) interval is determined based on the average value of the output results of each preset test data. Thus, the label SNR interval is set based on the output results of the conventional training method, and the lower limit of the label SNR interval is greater than the average value of the output results. This ensures that the SNR selected from the label SNR interval is greater than the average value, which means that the speech enhancement model trained from the label SNR interval has a better noise reduction effect.
[0113] Based on the first, second, and / or third embodiments of this application, in the fourth embodiment of this application, the content that is the same as or similar to the above-described embodiments one, two, and three can be referred to the above description and will not be repeated hereafter. In addition, the original dataset also includes the room impulse response. After the step of obtaining the original dataset, the method further includes:
[0114] Step D10: Add the room impulse response to the speech signal to obtain a first speech signal, and add the room impulse response to the noise signal to obtain a first noise signal;
[0115] Room impulse response (RIR) is an acoustic term used to describe the interaction of sound waves with surfaces such as walls and furniture as sound propagates within an enclosed space (such as a room). It reflects the process of sound waves reflecting, interfering with, and attenuating multiple times within the room after being emitted from a sound source. Specifically, room impulse response can be a time-domain function, describing the change in sound pressure received at any location within the room over time when a sound source emits a very short impulse sound.
[0116] Understandably, the room shock response can be pre-prepared; for example, relevant personnel can prepare the room shock response based on a publicly available room shock response dataset.
[0117] Adding the room impulse response to the signal (including speech and noise signals) can be achieved by convolving the room impulse response with the signal.
[0118] It should be noted that step D10 can be performed by selecting some or all of the speech and noise signals in the original dataset. That is, room impulse response can be added to some or all of the signals. The specific amount of speech and noise signals to be added to the room impulse response can be set by relevant personnel based on actual needs. This embodiment does not impose specific restrictions on this.
[0119] Furthermore, similar to speech signals and noise signals, the original dataset may include one or more sets of room impulse responses. For speech signals and noise signals to which room impulse responses are to be added, a set of room impulse responses can be selected from the original dataset each time and added to a pair of speech signals and noise signals. The first room impulse response (hereinafter referred to as the first response) in the set of room impulse responses is added to the speech signal, and the second room impulse response (hereinafter referred to as the second response) in the set of room impulse responses is added to the noise signal. For example, assuming the original dataset is [speech signal A, speech signal B, speech signal C, noise signal A, noise signal B, noise signal C, first set of room impulse responses, second set of room impulse responses], and room impulse responses are added to all speech and noise signals, then in one specific implementation, the first line of the first set of room impulse responses can be added to speech signal A, and the second line of the first set of room impulse responses can be added to noise signal A; the first line of the first set of room impulse responses can be added to speech signal B, and the second line of the first set of room impulse responses can be added to noise signal B; the first line of the second set of room impulse responses can be added to speech signal C, and the second line of the second set of room impulse responses can be added to noise signal C.
[0120] Reference Figure 3 As shown, each group of rooms simulates the same scenario in terms of impulse response (e.g., Figure 3 The office scene or lecture hall scene shown is illustrated, and the microphone is fixed in position within the scene; however, the two sound source locations are different, with the two sound sources used to simulate a speech source and a noise source, respectively. Therefore, each set of room impulse responses consists of two different room impulse responses, where the first RIR.1 describes the response at any location in the room when the speech source emits a very short pulse sound (RIR.1). Figure 3 The figure shows the change in sound pressure received at the microphone position over time. The second RIR.2 describes the change in sound pressure received at any location in the room when a noise source emits a very short pulse sound. Figure 3 The figure shows the change in sound pressure received at the microphone position over time.
[0121] Step D20: Using the first speech signal as a new speech signal and the first noise signal as a new noise signal, perform the step of mixing the speech signal and the noise signal at at least one first signal-to-noise ratio to obtain input data based on the new speech signal and the new noise signal.
[0122] In this embodiment, the room impulse response is added to the speech signal and the noise signal, that is, a reverberation effect is added to the speech signal and the noise signal. It is understood that signals in the real environment usually have a reverberation effect. In this way, the training dataset is constructed based on the signal with added reverberation effect, so that the signal in the training dataset is closer to the signal in the real environment, thereby further improving the training effect of the speech enhancement model.
[0123] In one possible implementation, after the step of determining the second signal-to-noise ratio based on the first signal-to-noise ratio of the input data, the method further includes:
[0124] Step E10: Obtain the training target of the speech enhancement model;
[0125] It should be noted that the training objective of the speech enhancement model can be set in advance by relevant personnel based on actual needs, and can specifically be either de-reverb or no-reverb. De-reverb indicates that reverb in the signal needs to be removed, while no-reverb indicates that reverb in the signal does not need to be removed.
[0126] Step E20: If the training objective indicates dereverberation, then the early reverberation in the room impulse response is added to the speech signal to obtain a second speech signal. The second speech signal is used as a new speech signal. Based on the new speech signal, the step of mixing the speech signal and the noise signal with the second signal-to-noise ratio to obtain the label data corresponding to the input data is performed.
[0127] It should be noted that step E20 can be performed only for voice signals that have added room impulse response.
[0128] If the training objective indicates dereverberation, then early reverberation from the room impulse response is added to the speech signal. Early reverberation is the portion of the room impulse response preceding the index of the element with the largest value. Early reverberation is a very short (typically 10ms) segment of the room impulse response after the direct sound, and it doesn't have a perceptible reverberation. Adding early reverberation to the speech signal ensures that the second speech signal has the same time delay as the first speech signal and the first noise signal. This, in turn, ensures that the time delay of the subsequently mixed input data is the same as that of the label data, aligning the label data and input data in the time dimension, which is beneficial for the learning of the speech enhancement model.
[0129] It should be noted that when the training objective is to de-reverberate, and label data is constructed based on the second speech signal, if the input data is constructed using the first noise data with added room impulse response, then the label data corresponding to the input data is obtained by mixing the second speech signal with the first noise signal. For example, assuming the original dataset is [speech signal A, speech signal B, speech signal C, noise signal A, noise signal B, noise signal C, first set of room impulse response, second set of room impulse response], the first set of room impulse response is added to speech signal A and noise signal A to obtain the first speech signal A and the first noise signal A. No room impulse response is added to speech signal B and noise signal B. Mixing the first speech signal A and the first noise signal A yields input data A, and mixing speech signal B and noise signal B yields input data B. If the training objective is to de-reverberate, then early reverberation of the first set of room impulse response is added to speech signal A to obtain the second speech signal A. Mixing the second speech signal A and the first noise signal A yields the label data A corresponding to input data A, and mixing speech signal B and noise signal B yields the label data B corresponding to input data B.
[0130] Furthermore, if the training objective indicates no dereverberation, the first speech signal and the first noise signal are mixed to obtain the label data, thus preserving the reverberation effect of the signal. For example, assuming the original dataset is [speech signal A, speech signal B, speech signal C, noise signal A, noise signal B, noise signal C, first set of room impulse responses, second set of room impulse responses], the second set of room impulse responses is added to speech signal C and noise signal C to obtain the first speech signal C and the first noise signal C. The first speech signal C and the second noise signal C are mixed to obtain the input data C. If the training objective indicates no dereverberation, the first speech signal C and the second noise signal C are mixed to obtain the label data C corresponding to the input data C.
[0131] In this embodiment, when the training objective indicates de-reverberation, the early reverberation in the room impulse response is added to the speech signal to obtain a second speech signal. Based on this second speech signal, label data is constructed, thereby ensuring that the trained speech enhancement model has a de-reverberation effect.
[0132] For example, to help understand the technical concept or principle of the speech enhancement model training method combined with the first, second, and third embodiments, please refer to... Figure 4 As shown, the training method for the speech enhancement model includes:
[0133] 1. Randomly select data. From the speech data of the original dataset ( Figure 4 A clean human voice (i.e., a speech signal) is randomly selected from the speech dataset shown in the image, and then compared with the noisy data in the original dataset. Figure 4A noise signal is randomly selected from the noise dataset shown in the image. This signal is then used to analyze the room impulse response data from the original dataset. Figure 4 A set of room shock responses (rir_list) is randomly selected from the room shock response dataset shown in the figure.
[0134] 2. Generate signal-to-noise ratio (SNR) (including input SNR and tag SNR). Three SNR generation methods are available.
[0135] Interval mapping: After randomly generating the first signal-to-noise ratio mix_snr from the signal-to-noise ratio range [-20, 20], a value is randomly sampled from the corresponding interval (based on the interval determined in Table 1) according to the interval where mix_snr is located, as the second signal-to-noise ratio label_snr.
[0136] Statistical mapping: After randomly generating the input signal-to-noise ratio mix_snr from the signal-to-noise ratio interval [-20, 20], a value is randomly sampled from the corresponding interval (based on the interval determined in Table 3) according to the interval where mix_snr is located, as the second signal-to-noise ratio label_snr.
[0137] Function mapping: The first signal-to-noise ratio mix_snr is randomly generated from the signal-to-noise ratio interval [-20, 20], and the second signal-to-noise ratio is calculated by the mapping function f(snr)=30-7*exp(-(snr+20)).
[0138] Figure 4 The diagram only shows the method of generating signal-to-noise ratio based on function mapping.
[0139] 3. Add reverberation as needed, i.e., add room impulse response as needed. First, add the second rir from rir_list to noise to obtain the first noise signal reverb_noise. If the model requires dreverberation capability, add the first rir from rir_list to clean to obtain the first speech signal reverb_clean. The first noise signal reverb_noise and the first speech signal reverb_clean are mixed at a certain first signal-to-noise ratio to obtain the input data mix. Clean is then added with early reverberation to obtain the second speech signal. The second speech signal is mixed with the first noise signal reverb_noise at a certain second signal-to-noise ratio to obtain the label data label.
[0140] When the model does not require dereverberation capability ( Figure 4(Not shown in the image), the first speech signal reverb_clean and the first noise signal reverb_noise are mixed at a certain first signal-to-noise ratio to obtain input data mix, and the first speech signal reverb_clean and the first noise signal reverb_noise are mixed at a certain second signal-to-noise ratio to obtain label data label.
[0141] 4. Train the model. Repeat steps 1 to 3 until the amount of input data and label data obtained meets certain requirements. Then, use the obtained input data as input and the obtained label data as labels to train the speech enhancement model.
[0142] It should be noted that the above examples are only used to help understand this application and do not constitute a limitation on the speech enhancement model training method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0143] Furthermore, this application also proposes a speech enhancement method, which includes the following steps:
[0144] Obtain the original signal to be enhanced and the target speech enhancement model, and input the original signal into the target speech enhancement model to obtain the speech enhancement result;
[0145] The target speech enhancement model is a speech enhancement model trained using the speech enhancement model training method described in any of the above embodiments.
[0146] The execution entity of each embodiment of the speech enhancement training method can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone; or an electronic device capable of performing the above functions, such as headphones, AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, AR headsets, VR headsets, smart glasses, etc. This embodiment does not impose specific limitations on this. Furthermore, the execution entity of the speech enhancement method and the speech enhancement model training method of this application can be the same or different, and no specific limitations are imposed here.
[0147] It is easy to understand that after obtaining the original signal to be enhanced, the original signal is input into the target speech enhancement model to obtain the speech enhancement result, which can be the output of the target speech enhancement model.
[0148] Furthermore, embodiments of this application also propose a speech enhancement model training device, referring to... Figure 5 As shown, the speech enhancement model training device includes:
[0149] The acquisition module 10 is used to acquire a training dataset and a speech enhancement model to be trained. The training dataset includes at least one input data and label data corresponding to each input data. At least one label data in each label data is a mixed signal containing speech and noise.
[0150] The training module 20 is used to train the speech enhancement model by using each of the input data as input to the speech enhancement model and each of the label data as labels to the speech enhancement model, so as to obtain the trained speech enhancement model.
[0151] In one embodiment, the speech enhancement model training device further includes a data preparation module, the data preparation module being used for:
[0152] Obtain the original dataset, wherein the original dataset includes speech signals and noise signals;
[0153] The input data is obtained by mixing the speech signal and the noise signal at at least one first signal-to-noise ratio.
[0154] For each piece of input data, a second signal-to-noise ratio is determined based on a first signal-to-noise ratio of the input data, and the speech signal and the noise signal are mixed with the second signal-to-noise ratio to obtain the tag data corresponding to the input data, wherein the second signal-to-noise ratio is greater than or equal to the first signal-to-noise ratio;
[0155] The training dataset is obtained by combining the input data and the label data.
[0156] In one embodiment, the data preparation module is further configured to:
[0157] Obtain the mapping table, and based on the mapping table, find the second signal-to-noise ratio corresponding to the first signal-to-noise ratio of the input data; or...
[0158] The first signal-to-noise ratio corresponding to the input data is input into a preset mapping function to obtain the second signal-to-noise ratio.
[0159] In one embodiment, the data preparation module is further configured to:
[0160] At least one preset test data is input into a preset conventional speech enhancement model to obtain the output result corresponding to each preset test data, wherein the preset conventional speech enhancement model is the speech enhancement model pre-trained with speech signals as labels;
[0161] Obtain the input signal-to-noise ratio range to which the signal-to-noise ratio of each preset test data belongs;
[0162] For each input signal-to-noise ratio (SNR) interval, calculate the average value of the output results corresponding to all preset test data belonging to the input SNR interval, determine the label SNR interval based on the average value, and configure a correspondence between the input SNR interval and the label SNR interval, wherein the lower limit of the label SNR interval is greater than or equal to the average value.
[0163] By combining the correspondence between each input signal-to-noise ratio interval and each label signal-to-noise ratio interval, a mapping table is obtained.
[0164] In one embodiment, the data preparation module is further configured to:
[0165] The room impulse response is added to the speech signal to obtain a first speech signal, and the room impulse response is added to the noise signal to obtain a first noise signal;
[0166] The first speech signal is used as the new speech signal, and the first noise signal is used as the new noise signal.
[0167] In one embodiment, the data preparation module is further configured to:
[0168] Obtain the training target of the speech enhancement model;
[0169] If the training objective indicates dereverberation, then the early reverberation in the room impulse response is added to the speech signal to obtain a second speech signal, and the second speech signal is used as the new speech signal.
[0170] Furthermore, embodiments of this application also propose an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the speech enhancement model training method and / or speech enhancement method as described above.
[0171] refer to Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing the embodiments of this application. The electronic devices in the embodiments of this application may also include, but are not limited to, mobile terminals such as headphones, AR glasses, VR glasses, AR headsets, VR headsets, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 6The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0172] like Figure 6 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0173] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0174] The electronic device provided in this application, employing the speech enhancement model training method described in the above embodiments, can solve the technical problem of how to reduce the amount of speech suppression by the speech enhancement model, thereby improving speech clarity and quality. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the speech enhancement model training method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0175] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0176] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0177] In addition, to achieve the above objectives, embodiments of this application also provide a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the speech enhancement model training method in the above embodiments.
[0178] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0179] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0180] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire a training dataset and a speech enhancement model to be trained, wherein the training dataset includes at least one input data and label data corresponding to each input data, and at least one of the label data is a mixed signal containing both speech and noise; train the speech enhancement model using each input data as input to the speech enhancement model and each label data as a label for the speech enhancement model, thereby obtaining the trained speech enhancement model.
[0181] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0183] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the modules themselves.
[0184] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described speech enhancement model training method and / or speech enhancement method. This addresses the technical problem of how to reduce the amount of speech suppression by the speech enhancement model to improve speech clarity and quality. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the speech enhancement model training method provided in the above embodiments, and will not be elaborated upon here.
[0185] Furthermore, embodiments of this application also propose a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech enhancement model training method and / or speech enhancement method as described above.
[0186] The specific implementation of the computer program product in this application is basically the same as the embodiments of the above-mentioned speech enhancement model training and / or speech enhancement method, and will not be repeated here.
[0187] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0188] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0189] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0190] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for training a speech enhancement model, characterized in that, The speech enhancement model training method includes the following steps: Obtain the original dataset, wherein the original dataset includes speech signals and noise signals; The input data is obtained by mixing the speech signal and the noise signal at at least one first signal-to-noise ratio. For each piece of input data, at least one set of preset test data is input into a preset conventional speech enhancement model to obtain the output result corresponding to each set of preset test data. The preset conventional speech enhancement model is the speech enhancement model pre-trained with speech signals as labels. Obtain the input signal-to-noise ratio range to which the signal-to-noise ratio of each preset test data belongs; For each input signal-to-noise ratio (SNR) interval, calculate the average value of the output results corresponding to all preset test data belonging to the input SNR interval, determine the label SNR interval based on the average value, and configure a correspondence between the input SNR interval and the label SNR interval, wherein the lower limit of the label SNR interval is greater than or equal to the average value. By combining the correspondence between each input signal-to-noise ratio interval and each label signal-to-noise ratio interval, a mapping table is obtained; Based on the mapping table, find the second signal-to-noise ratio corresponding to the first signal-to-noise ratio of the input data, and mix the speech signal and the noise signal with the second signal-to-noise ratio to obtain the label data corresponding to the input data; The training dataset is obtained by combining the input data and the label data. Obtain the speech enhancement model to be trained; The speech enhancement model is trained by using the input data as input and the label data as labels, and the trained speech enhancement model is obtained.
2. The speech enhancement model training method as described in claim 1, characterized in that, The original dataset also includes room impulse responses. Following the step of obtaining the original dataset, the method further includes: The room impulse response is added to the speech signal to obtain a first speech signal, and the room impulse response is added to the noise signal to obtain a first noise signal; The first speech signal is used as a new speech signal, and the first noise signal is used as a new noise signal. Based on the new speech signal and the new noise signal, the step of mixing the speech signal and the noise signal at at least one first signal-to-noise ratio to obtain input data is performed.
3. The speech enhancement model training method as described in claim 2, characterized in that, After the step of finding the second signal-to-noise ratio corresponding to the first signal-to-noise ratio of the input data based on the mapping table, the method further includes: Obtain the training target of the speech enhancement model; If the training objective indicates dereverberation, then the early reverberation in the room impulse response is added to the speech signal to obtain a second speech signal. The second speech signal is used as a new speech signal, and the step of mixing the speech signal and the noise signal with the second signal-to-noise ratio to obtain the label data corresponding to the input data is performed based on the new speech signal.
4. A speech enhancement method, characterized in that, The speech enhancement method includes the following steps: Obtain the original signal to be enhanced and the target speech enhancement model, and input the original signal into the target speech enhancement model to obtain the speech enhancement result; The target speech enhancement model is a speech enhancement model trained using the speech enhancement model training method as described in any one of claims 1 to 3.
5. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method as claimed in any one of claims 1 to 4.
6. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
7. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 4.