Speech enhancement method and system for AI noise reduction
By obtaining and dividing diverse vocals, noise and reverb samples in indoor scenes, filtering and muting operations, generating reverb vocals and target noise, the problem of low accuracy in the existing technology is solved and the accuracy and applicability of the speech enhancement method is improved.
Patent Information
- Application Number
- CN202510061773.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-06
AI Technical Summary
The existing speech enhancement methods have low accuracy in diverse scenarios and cannot effectively simulate reverb and signal processing collected by audio equipment in indoor scenarios, resulting in increased training difficulty for supervision tasks and poor results.
By obtaining different types of vocal samples, noise samples and reverb samples in indoor scenes, they are divided into long audio, short audio, burst noise, steady-state noise, white noise, light reverb, medium reverb and reverb. Then, randomly sampled for filtering, signal-to-noise ratio adjustment and mute operation to generate reverb vocals and target noise, simulate the audio collected by the sound pickup device, and is used for noise reduction model training.
It improves the accuracy of the speech enhancement method, enhances the performance performance of the noise reduction model in different scenarios, reduces the difficulty of supervised learning, improves applicability, and promotes the development of voice communication, audio processing and speech recognition.
Smart Images

Figure CN119943074A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech signal processing technology, and in particular to a speech enhancement method and system for AI noise reduction. Background Art
[0002] Currently, existing speech enhancement methods combine noisy human voice data and clean human voice data for supervision tasks, such as randomly combining noise and human voice to generate a variety of noisy human voices and clean human voices, thereby providing samples for supervision tasks.
[0003] However, simply using the method of adding noise to human voice to generate noisy human voice often ignores the diversity of existing scenarios. For example, there will be sound reflections indoors, a phenomenon also known as reverberation. Existing noise reduction models trained on speech data do not have the function of removing reverberation. Secondly, when audio equipment is picking up sound, various signal processing is often performed. Simply adding noise to human voice cannot simulate the audio collected by the sound pickup device, and thus it is difficult to provide more useful samples for supervision tasks. It will also increase the difficulty of training supervision tasks, and thus fail to achieve better results in specific scenarios.
[0004] With respect to the above-mentioned related technologies, the inventors have found that the existing speech enhancement methods have the problem of low precision in diverse scenarios. Summary of the invention
[0005] In order to improve the accuracy of the speech enhancement method, the present application provides a speech enhancement method and system for AI noise reduction.
[0006] In a first aspect, the present application provides a speech enhancement method for AI noise reduction.
[0007] This application is achieved through the following technical solutions:
[0008] A speech enhancement method for AI noise reduction comprises the following steps:
[0009] Obtain different types of vocal samples, noise samples, and reverberation samples in indoor scenes;
[0010] The human voice samples are divided into long audio and short audio, the noise samples are divided into burst noise, steady-state noise and white noise, and the reverberation samples are divided into light reverberation, medium reverberation and heavy reverberation;
[0011] Randomly extracting a to-be-processed human voice and a to-be-processed reverberation from the classified human voice samples and reverberation samples, subjecting the to-be-processed human voice to filtering and muting operations in sequence, and then performing a reverberation operation in combination with the to-be-processed reverberation to obtain a human voice with reverberation;
[0012] Randomly extract a noise to be processed from the classified noise samples, and subject the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations in sequence to obtain a target noise;
[0013] Based on the reverberant human voice and the target noise, a noisy human voice is generated as audio collected by a sound pickup device for noise reduction model training.
[0014] In a preferred example, the present application can be further configured as follows: after the human voice to be processed is subjected to filtering and muting operations in sequence, a reverberation operation is performed in combination with the reverberation to be processed to obtain a human voice with reverberation, and an adaptive reverberation simulation operation is performed on the human voice with reverberation, including:
[0015] According to different types of reverberation to be processed, calculating the time taken for the sound field with reverberated human voice to decay to a preset decibel value;
[0016] If the preset decibel value of the sound field attenuation is within the first time threshold range, attenuating the reverberated human voice for a first time length;
[0017] If the time taken for the sound field to decay to a preset decibel value is within a second time threshold, decaying the reverberant human voice for a second time length;
[0018] If the time taken for the sound field to attenuate the preset decibel value is within a third time threshold, attenuating the reverberated human voice for a third time length;
[0019] Among them, the first time threshold < the second time threshold < the third time threshold, the first duration < the second duration < the third duration;
[0020] Updated the vocals with reverb.
[0021] In a preferred example, the present application can be further configured as follows: after the human voice to be processed is subjected to filtering processing and muting operation in sequence, a reverberation operation is performed in combination with the reverberation to be processed, and the step of obtaining the human voice with reverberation includes:
[0022] The human voice to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain an initial human voice;
[0023] Using a muffler to mute the initial human voice with a preset probability ratio to obtain a target human voice;
[0024] The reverberation to be processed and the target vocal are convolved to generate the vocal with reverberation.
[0025] In a preferred example, the present application can be further configured as follows: the step of subjecting the noise to be processed to filtering, signal-to-noise ratio adjustment and mute operation in sequence to obtain the target noise includes:
[0026] The noise to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain initial noise;
[0027] Randomly adjusting the volume of the initial noise according to the energy of the clean audio, and setting the signal-to-noise ratio of the clean audio and the initial noise;
[0028] The target noise is produced by using a silencer to mute the initial noise after the volume adjustment with a probability of a preset ratio.
[0029] In a preferred example, the present application may be further configured as follows: based on the reverberant human voice and the target noise, the step of generating the noisy human voice as the audio collected by the sound pickup device includes:
[0030] The reverberant human voice and the target noise are added to obtain the noisy human voice.
[0031] In a preferred example, the present application may be further configured as follows: after the step of generating the noisy human voice as the audio collected by the sound pickup device based on the human voice with reverberation and the target noise, the step further includes:
[0032] A volume calibration operation is performed on the noisy human voice.
[0033] In a preferred example, the present application can be further configured as follows: it also includes the following steps:
[0034] Repeatedly extracting the human voice to be processed and the reverberation to be processed to generate the human voice with reverberation, and repeatedly extracting the noise to be processed to generate the target noise, and using the generated human voice with reverberation and the generated target noise to generate the noisy human voice, until the number of samples of the noisy human voice reaches a preset threshold.
[0035] In a second aspect, the present application provides a speech enhancement system for AI noise reduction.
[0036] This application is achieved through the following technical solutions:
[0037] A speech enhancement system for AI noise reduction, comprising:
[0038] The sampling module is used to obtain different types of human voice samples, noise samples and reverberation samples in indoor scenes;
[0039] A classification module, used to classify the human voice samples into long audio and short audio, classify the noise samples into burst noise, steady-state noise and white noise, and classify the reverberation samples into light reverberation, medium reverberation and heavy reverberation;
[0040] A reverberation human voice generation module is used to randomly extract a human voice to be processed and a reverberation to be processed from the classified human voice samples and reverberation samples, subject the human voice to be processed to filtering and muting operations in sequence, and then perform a reverberation operation in combination with the reverberation to be processed to obtain a human voice with reverberation;
[0041] A target noise generation module is used to randomly extract a noise to be processed from the classified noise samples, and sequentially subject the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations to obtain a target noise;
[0042] The noisy human voice generation module is used to generate the noisy human voice as the audio collected by the sound pickup device based on the reverberant human voice and the target noise, for noise reduction model training.
[0043] In a preferred example, the present application can be further configured as follows: further comprising:
[0044] The adaptive reverberation module is used to perform an adaptive reverberation simulation operation on the human voice with reverberation, including calculating the time taken for the sound field of the human voice with reverberation to decay to a preset decibel value according to different types of reverberation to be processed; if the time taken for the sound field to decay to the preset decibel value is within a first time threshold range, decaying the human voice with reverberation for a first time length; if the time taken for the sound field to decay to the preset decibel value is within a second time threshold range, decaying the human voice with reverberation for a second time length; if the time taken for the sound field to decay to the preset decibel value is within a third time threshold range, decaying the human voice with reverberation for a third time length; wherein, the first time threshold < the second time threshold < the third time threshold, and the first time length < the second time length < the third time length; and updating the human voice with reverberation.
[0045] In a third aspect, the present application provides a computer device.
[0046] This application is achieved through the following technical solutions:
[0047] A computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the above-mentioned speech enhancement methods for AI noise reduction when executing the computer program.
[0048] In summary, compared with the prior art, the technical solution provided by this application has at least the following beneficial effects:
[0049] Different types of human voice samples, noise samples and reverberation samples in indoor scenes are obtained as the data basis; the human voice samples are divided into long audio and short audio to enhance the generalization ability of the noise reduction model when processing different people's speech; the noise samples are divided into burst noise, steady-state noise and white noise, and the reverberation samples are divided into light reverberation, medium reverberation and heavy reverberation, which is conducive to voice control in different proportions according to actual scenes and ensures the diversity of sample scenes; a human voice to be processed and a reverberation to be processed are randomly selected from the classified human voice samples and reverberation samples, and the human voice to be processed is filtered and muted in turn to simulate the signal processing operation after the audio device picks up the sound, improve the impact of different filter devices on the human voice collection and improve the performance of the AI model in the pure human voice scene; then the reverberation to be processed is combined for reverberation The method further improves the performance of the AI model in different noisy scenes by performing filtering, adjusting the signal-to-noise ratio and muting the noise to be processed, which is conducive to achieving better noise reduction results, improving the accuracy of speech enhancement methods, and promoting the development of speech communication, audio processing and speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of the main flow of a speech enhancement method for AI noise reduction provided as an exemplary embodiment of the present application.
[0051] Figure 2 An overall flow chart of a speech enhancement method for AI noise reduction provided as another exemplary embodiment of the present application.
[0052] Figure 3 An adaptive reverberation simulation flowchart of a speech enhancement method for AI noise reduction is provided as another exemplary embodiment of the present application.
[0053] Figure 4 A structural block diagram of a speech enhancement device for AI noise reduction provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0054] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make modifications to the present embodiment without any creative contribution as needed, but such modifications are protected by the patent law as long as they are within the scope of the claims of the present application.
[0055] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0056] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.
[0057] The AI noise reduction model trained with the existing open source datasets performs poorly in indoor scenes, mainly due to poor noise reduction effect, severe loss of human voice, and lack of dereverberation. This solution performs data augmentation and scene simulation based on the existing open source datasets, so that the AI noise reduction model can perform well in various scenes with poor effects.
[0058] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.
[0059] Reference Figure 1 , an embodiment of the present application provides a speech enhancement method for AI noise reduction, and the main steps of the method are described as follows.
[0060] S1: Obtain different types of human voice samples, noise samples and reverberation samples in indoor scenes;
[0061] S2: dividing the human voice samples to obtain long audio and short audio, dividing the noise samples to obtain burst noise, steady-state noise and white noise, and dividing the reverberation samples to obtain light reverberation, medium reverberation and heavy reverberation;
[0062] S3: randomly extracting a to-be-processed human voice and a to-be-processed reverberation from the classified human voice samples and reverberation samples, subjecting the to-be-processed human voice to filtering and muting operations in sequence, and then performing a reverberation operation in combination with the to-be-processed reverberation to obtain a human voice with reverberation;
[0063] S4: randomly extracting a noise to be processed from the classified noise samples, and subjecting the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations in sequence to obtain a target noise;
[0064] S5: Based on the reverberant human voice and the target noise, generate a noisy human voice as audio collected by a sound pickup device for noise reduction model training.
[0065] Specifically, different types of human voice samples, noise samples, and reverberation samples are obtained in indoor scenes. Among them, human voice samples, noise samples, and reverberation samples are from open source datasets. Human voice samples are clean human voices. There are many types of noise samples. For example, in the noise scene of a conference room, there are many types of noise. When constructing the dataset, it is necessary to collect a variety of noises, including knocking on the table, clapping, moving chairs, etc.
[0066] Divide open source vocal samples, noise samples, and reverb samples into preset categories.
[0067] The human voice samples include long audio and short audio. The classification of human voice length is mainly to enhance the generalization ability of the denoising model when processing different people speaking.
[0068] For example, in a meeting scenario, there will be multiple people discussing, and the noise reduction model has to process both the voice of A and the voice of B. If only long audio is used as a training sample, the noise reduction model will have a probabilistic degradation problem when the voice switches.
[0069] For example, in the human voice scene in the conference room, there are scenes such as simultaneous speaking, interrupting each other, and different people speaking at different distances. Conversations in meetings are basically short conversations, so short audio is more in line with the scene requirements; in addition, using short audio can ensure the integrity of a sentence, ensuring that the model can achieve better results when processing such scenes.
[0070] When different people speak at different distances, the volume of the voices picked up by the microphone will be different due to the distance. In this case, short audio can be used to simulate people speaking at different volumes.
[0071] For audio in simultaneous speaking scenarios, different audios with different volumes can be added together and the ratio of the two audios can be adjusted to ensure that scenarios of different distances and different people speaking at the same time can be simulated, thereby increasing the generalization of the model.
[0072] The noise samples include at least burst noise, steady-state noise and white noise.
[0073] Reverb samples include light reverb, medium reverb, and heavy reverb.
[0074] The human voice samples are divided into long audio and short audio, the noise samples are divided into burst noise, steady-state noise and white noise, and the reverberation samples are divided into light reverberation, medium reverberation and heavy reverberation. The human voice data, noise data and reverberation data are subdivided to simulate audio in different scenarios as training samples, thereby improving the model's noise reduction capabilities in various scenarios.
[0075] Next, based on the characteristics of the recorded audio, the recorded audio is simulated. A human voice to be processed and a reverberation to be processed are randomly selected from the classified human voice samples and reverberation samples, and the human voice to be processed is filtered by a filter and muted by a silencer in turn, and then reverberated in combination with the reverberation to be processed to obtain a human voice with reverberation, simulating the reverberation scene audio; and a noise to be processed is randomly selected from the classified noise samples, and the noise to be processed is filtered by a filter, the signal-to-noise ratio is adjusted, and the mute operation of the silencer is performed to obtain the target noise; the signal processing process of the pickup device is simulated to make the generated noisy human voice closer to the audio collected by the pickup device, improve the impact of different filter devices on the collection of human voice / noise, and improve the performance of the AI model in pure human voice / pure noise scenarios.
[0076] Finally, based on the reverberant human voice and the target noise, the noisy human voice is generated as the audio collected by the pickup device, that is, the audio collected by the microphone is simulated for noise reduction model training.
[0077] Reference Figure 2 In one embodiment, the steps of filtering and muting the human voice to be processed and then performing a reverberation operation on the human voice to be processed to obtain a human voice with reverberation include:
[0078] The human voice to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain an initial human voice;
[0079] Using a muffler to mute the initial human voice with a preset probability ratio to obtain a target human voice;
[0080] The reverberation to be processed and the target vocal are convolved to generate the vocal with reverberation.
[0081] Among them, filtering processing includes high-pass filtering and low-pass filtering, such as using high-pass and low-pass filters to enhance or suppress audio signals within a specific frequency range to achieve functions such as signal denoising and signal smoothing, so as to simulate the signal processing process of the pickup device and improve the impact of different filter devices on human voice collection. Filtering processing also includes a small amount of bandpass filtering, and the number of bandpass filtering is less than that of high-pass filtering or low-pass filtering, so as to achieve data augmentation and enhance the generalization of AI models. After the human voice is filtered, it will be probabilistically extinguished to solve the problem of poor human voice suppression in pure human voice scenes, so as to improve the problem of poor performance of AI models in pure human voice scenes.
[0082] By convolving the impulse response and the input audio, the input audio is given a sense of reverberation, and the simulated audio is more similar to the audio collected in the actual room. The main process of the convolution operation is to read the existing RIR audio (i.e. the reverberation to be processed) and the target human voice. The RIR audio is the room impulse response, including rooms with light, medium and heavy reverberation, rooms without reverberation, and rooms of different sizes to simulate various indoor reverberation scenes; the RIR audio and the target human voice are convolved to obtain a human voice with reverberation. Different types of RIR audio will cause different degrees of reverberation on the human voice.
[0083] In one embodiment, the filter order may be 2-4, the frequency band of the high-pass filter is mainly in the range of 0-300 Hz, the frequency band of the low-pass filter is mainly in the range of 7500 Hz-8000 Hz, and the probability ratio of the silencer may be 3%-5%.
[0084] In one embodiment, the filtering process also includes band-stop processing, etc., to achieve data augmentation and enhance the generalization of the AI model, which will not be repeated here.
[0085] Reference Figure 2 In one embodiment, the step of subjecting the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operation in sequence to obtain the target noise includes:
[0086] The noise to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain initial noise;
[0087] Randomly adjusting the volume of the initial noise according to the energy of the clean audio, and setting the signal-to-noise ratio of the clean audio and the initial noise;
[0088] The target noise is produced by performing a silencer on the initial noise after the volume adjustment with a probability of a preset ratio; for example, the target noise is obtained by performing a silencer on the initial noise after the volume adjustment with a probability of 5%.
[0089] Among them, filtering processing includes high-pass filtering and low-pass filtering to simulate the signal processing process of the sound pickup device and improve the impact of different filter devices on noise collection. Filtering processing also includes a small amount of bandpass filtering, which is less than high-pass filtering or low-pass filtering, to achieve data augmentation and enhance the generalization of AI models.
[0090] The signal-to-noise ratio adjustment randomly adjusts the volume of noise through the energy of clean audio, thereby indirectly adjusting the signal-to-noise ratio of clean audio and noise signals. The signal-to-noise ratio range can be between -8dB and 15dB, which can basically cover the existing noise scenarios, so that the performance of the noise reduction model in different noisy scenarios can be further improved.
[0091] After the noise is filtered and the signal-to-noise ratio is adjusted, the noise will be probabilistically muted, such as using a silencer to mute the initial noise after volume adjustment with a probability of 3%-5%, so as to improve the problem of poor performance of the AI model in pure noise scenarios.
[0092] In one embodiment, the step of generating a noisy human voice as audio collected by a sound pickup device based on the human voice with reverberation and the target noise includes:
[0093] The human voice with reverberation and the target noise are added to obtain the human voice with noise. The operation is simple and easy to implement.
[0094] In order to enhance the dereverberation capability of the noise reduction model, an adaptive reverberation simulation algorithm is designed for different reverberation scenarios. In one embodiment, after the human voice to be processed is subjected to filtering and muting operations in sequence, a reverberation operation is performed in combination with the reverberation to be processed to obtain a human voice with reverberation, and the step further includes performing an adaptive reverberation simulation operation on the human voice with reverberation, including:
[0095] According to different types of reverberation to be processed, calculating the time taken for the sound field with reverberated human voice to decay to a preset decibel value;
[0096] If the time taken for the sound field to attenuate to a preset decibel value is within a first time threshold, attenuating the reverberated human voice for a first time length;
[0097] If the time taken for the sound field to decay to a preset decibel value is within a second time threshold, decaying the reverberant human voice for a second time length;
[0098] If the time taken for the sound field to attenuate the preset decibel value is within a third time threshold, attenuating the reverberated human voice for a third time length;
[0099] Among them, the first time threshold < the second time threshold < the third time threshold, the first duration < the second duration < the third duration;
[0100] Updated the vocals with reverb.
[0101] Referring to Figure 3 , in this embodiment, the adaptive reverberation simulation algorithm determines the reverberation scene and then determines the reverberation length of the training sample by calculating the time RT60 for any reverberant human voice to attenuate by 60 dB according to different types of impulse responses, that is, each reverberation to be processed. For example, in a heavy reverberation scene, if RT60 ≤ 300 ms, then attenuate the reverberant human voice by [30 ms, 50 ms), and the reverberation length of the training label band is longer; in a medium reverberation scene, if 300 ms < RT60 ≤ 800 ms, then attenuate the reverberant human voice by [50 ms, 100 ms), and the reverberation length of the training label band is moderate; in a light reverberation scene, if RT60 > 800 ms, then attenuate the reverberant human voice by [100 ms, 150 ms), and at this time the reverberation length of the training label band is shorter. Compared with the prior art that sets the same length of reverberation labels in all types of rooms, this solution adopts a multi-segment processing method, sets different lengths of reverberation labels for different reverberant rooms, and ensures that the model can have better performance in rooms with different reverberations. Through practice, this method significantly improves the reverberation removal ability of the model, has a good performance in the voice restoration degree, and the noise reduction model has a significant effect on the reverberation removal scene.
[0102] In one embodiment, after the step of generating a noisy human voice based on the reverberant human voice and the target noise as the audio collected by the pickup device, it further includes,
[0103] Performing a volume calibration operation on the noisy human voice.
[0104] By calibrating the volume of the noisy human voice to a set volume range, such as [-55 dB, -15 dB], to obtain the training data set of the AI noise reduction model. Because this volume range is relatively wide, the AI noise reduction model has good performance in different volume scenarios, ensuring the generalization of the model and having better effects when processing different volumes, and thus being able to obtain a simulated audio closer to the audio collected by the pickup device, improving the damage and mis-cancellation problems that occur in the prior model when processing larger and smaller volumes.
[0105] Input the generated training data set into the AI noise reduction model. The training label of the AI noise reduction model uses the human voice without added noise, and train the AI noise reduction model until the output accuracy of the model reaches the preset requirement.
[0106] It has been verified that compared with open source data sets and the same model, the AI noise reduction model trained with the data set generated by this application has a stronger ability to restore voiceprints under low signal-to-noise ratios, reduces the ups and downs of human voices, and significantly improves the denoising effect; in conference room scenes with reverberation, the denoising effect of the model is stable, and the listening experience after processing is better; in pure noise scenes, the denoising ability of the model is greatly improved, and the residual noise is reduced. The training data set generated by this application greatly improves the model's noise reduction ability in various scenarios, is more conducive to the training of AI noise reduction models, and has better generalization ability and stronger applicability.
[0107] In one embodiment, the following steps are also included:
[0108] Repeatedly extract the human voice to be processed and the reverberation to be processed to generate the human voice with reverberation, and repeatedly extract the noise to be processed to generate the target noise, and use the generated human voice with reverberation and the generated target noise to generate the human voice with noisy, until the number of samples of the human voice with noisy reaches a preset threshold, so as to ensure the number of samples in the training data set of the denoising model and improve the denoising effect of the model.
[0109] In one embodiment, the proportions of different types of reverberation are adjusted according to actual needs, and then the human voice is convolved with the reverberation data of a specific proportion to obtain the human voice with reverberation; and the proportions of different types of noise are allocated according to actual needs; and then different proportions and types of human voices and noise can be extracted, processed and superimposed according to the needs of the actual scene. The reverberation ratio and noise ratio in the generated data set are controllable to accurately simulate the noisy human voice in a specific scene for model training, and also ensure the diversity of the scene.
[0110] In summary, a speech enhancement method for AI noise reduction obtains different types of human voice samples, noise samples and reverberation samples in indoor scenes as data basis; divides the human voice samples to obtain long audio and short audio to enhance the generalization ability of the noise reduction model when processing different people's speech; divides the noise samples to obtain burst noise, steady-state noise and white noise, and divides the reverberation samples to obtain light reverberation, medium reverberation and heavy reverberation, which is conducive to voice control of different proportions according to actual scenes, ensuring the diversity of sample scenes; randomly extracts a human voice to be processed and a reverberation to be processed from the classified human voice samples and reverberation samples, and sequentially filters and mutes the human voice to be processed to simulate the signal processing operation after the audio device picks up the sound, improves the influence of different filter devices on the human voice collection and improves the performance of the AI model in the pure human voice scene; then performs reverberation operation in combination with the reverberation to be processed to obtain a human voice with reverberation to simulate the reverberation scene audio;
[0111] A noise to be processed is randomly selected from the classified noise samples, and the noise to be processed is filtered, adjusted in signal-to-noise ratio and muted in sequence to simulate the signal processing operation after the audio device picks up the sound, improve the impact of different filter devices on noise collection and improve the performance of the AI model in pure noise scenarios. The performance of the model in different noisy scenarios is further improved; finally, based on the reverberant human voice and target noise, a noisy human voice is generated to approximate the audio of the real scene. The simulated noisy human voice is closer to the audio collected by the pickup device, and the difficulty of supervised learning is greatly reduced, which is conducive to achieving better noise reduction results, improving the accuracy of the speech enhancement method, improving the applicability, and promoting the development of voice communication, audio processing, speech recognition and other fields.
[0112] A speech enhancement method for AI noise reduction can simulate various existing indoor scene audio, so that the AI noise reduction model can perform better in various scenes with poor effects.
[0113] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0114] Reference Figure 4 The embodiment of the present application also provides a speech enhancement system for AI noise reduction, which corresponds one-to-one to a speech enhancement method for AI noise reduction in the above embodiment. The speech enhancement system for AI noise reduction includes:
[0115] The sampling module is used to obtain different types of human voice samples, noise samples and reverberation samples in indoor scenes;
[0116] A classification module, used to classify the human voice samples into long audio and short audio, classify the noise samples into burst noise, steady-state noise and white noise, and classify the reverberation samples into light reverberation, medium reverberation and heavy reverberation;
[0117] A reverberation human voice generation module is used to randomly extract a human voice to be processed and a reverberation to be processed from the classified human voice samples and reverberation samples, subject the human voice to be processed to filtering and muting operations in sequence, and then perform a reverberation operation in combination with the reverberation to be processed to obtain a human voice with reverberation;
[0118] A target noise generation module is used to randomly extract a noise to be processed from the classified noise samples, and sequentially subject the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations to obtain a target noise;
[0119] The noisy human voice generation module is used to generate the noisy human voice as the audio collected by the sound pickup device based on the reverberated human voice and the target noise, for noise reduction model training.
[0120] A speech enhancement system for AI noise reduction also includes:
[0121] The adaptive reverberation module is used to perform an adaptive reverberation simulation operation on the human voice with reverberation, including calculating the time taken for the sound field of the human voice with reverberation to decay to a preset decibel value according to different types of reverberation to be processed; if the time taken for the sound field to decay to the preset decibel value is within a first time threshold range, decaying the human voice with reverberation for a first time length; if the time taken for the sound field to decay to the preset decibel value is within a second time threshold range, decaying the human voice with reverberation for a second time length; if the time taken for the sound field to decay to the preset decibel value is within a third time threshold range, decaying the human voice with reverberation for a third time length; wherein, the first time threshold < the second time threshold < the third time threshold, and the first time length < the second time length < the third time length; and updating the human voice with reverberation.
[0122] For the specific definition of a speech enhancement system for AI noise reduction, please refer to the definition of a speech enhancement method for AI noise reduction in the above text, which will not be repeated here.
[0123] Each module in the above-mentioned speech enhancement system for AI noise reduction can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0124] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, any of the above-mentioned speech enhancement methods for AI noise reduction is implemented.
[0125] In one embodiment, a computer-readable storage medium is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, any one of the above-mentioned speech enhancement methods for AI noise reduction is implemented.
[0126] In one embodiment, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, any one of the above-mentioned speech enhancement methods for AI noise reduction is implemented.
[0127] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. When the computer program is executed, it may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0128] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A speech enhancement method for AI noise reduction, characterized in that: The following steps are included: Obtain different types of vocal samples, noise samples, and reverberation samples in indoor scenes; The human voice samples are divided into long audio and short audio, the noise samples are divided into burst noise, steady-state noise and white noise, and the reverberation samples are divided into light reverberation, medium reverberation and heavy reverberation; Randomly extracting a to-be-processed human voice and a to-be-processed reverberation from the classified human voice samples and reverberation samples, subjecting the to-be-processed human voice to filtering and muting operations in sequence, and then performing a reverberation operation in combination with the to-be-processed reverberation to obtain a human voice with reverberation; Randomly extract a noise to be processed from the classified noise samples, and subject the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations in sequence to obtain a target noise; Based on the reverberant human voice and the target noise, a noisy human voice is generated as audio collected by a sound pickup device for noise reduction model training.
2. The speech enhancement method for AI noise reduction according to claim 1, characterized in that: After the human voice to be processed is subjected to filtering and muting operations in sequence, a reverberation operation is performed in combination with the reverberation to be processed to obtain a human voice with reverberation, and the method further includes performing an adaptive reverberation simulation operation on the human voice with reverberation, including: According to different types of reverberation to be processed, calculating the time taken for the sound field with reverberated human voice to decay to a preset decibel value; If the time taken for the sound field to attenuate the preset decibel value is within the first time threshold range, attenuating the reverberated human voice for a first time length; If the time taken for the sound field to decay to a preset decibel value is within a second time threshold, decaying the reverberant human voice for a second time length; If the time taken for the sound field to attenuate the preset decibel value is within a third time threshold, attenuating the reverberated human voice for a third time length; Among them, the first time threshold < the second time threshold < the third time threshold, the first duration < the second duration < the third duration; Updated the vocals with reverb.
3. The speech enhancement method for AI noise reduction according to any one of claims 1 to 2, characterized in that: The steps of filtering and muting the human voice to be processed in sequence and then performing a reverberation operation in combination with the reverberation to be processed to obtain a human voice with reverberation include: The human voice to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain an initial human voice; Using a muffler to mute the initial human voice with a preset probability ratio to obtain a target human voice; The reverberation to be processed and the target vocal are convolved to generate the vocal with reverberation.
4. The speech enhancement method for AI noise reduction according to claim 3, characterized in that: The step of subjecting the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operation in sequence to obtain the target noise includes: The noise to be processed is passed through high-pass filters, low-pass filters and a small number of band-pass filters of different frequency bands to obtain initial noise; Randomly adjusting the volume of the initial noise according to the energy of the clean audio, and setting the signal-to-noise ratio of the clean audio and the initial noise; The target noise is produced by using a silencer to mute the initial noise after the volume adjustment with a probability of a preset ratio.
5. The speech enhancement method for AI noise reduction according to claim 4, characterized in that: Based on the reverberant human voice and the target noise, the step of generating the noisy human voice as the audio collected by the sound pickup device comprises: The reverberant human voice and the target noise are added to obtain the noisy human voice.
6. The speech enhancement method for AI noise reduction according to claim 5, characterized in that: After the step of generating a noisy human voice as audio collected by a sound pickup device based on the human voice with reverberation and the target noise, the method further includes: A volume calibration operation is performed on the noisy human voice.
7. The speech enhancement method for AI noise reduction according to claim 6, characterized in that: The following steps are also included: Repeatedly extracting the human voice to be processed and the reverberation to be processed to generate the human voice with reverberation, and repeatedly extracting the noise to be processed to generate the target noise, and using the generated human voice with reverberation and the generated target noise to generate the noisy human voice, until the number of samples of the noisy human voice reaches a preset threshold.
8. A speech enhancement system for AI noise reduction, characterized in that: include, The sampling module is used to obtain different types of human voice samples, noise samples and reverberation samples in indoor scenes; A classification module, used to classify the human voice samples into long audio and short audio, classify the noise samples into burst noise, steady-state noise and white noise, and classify the reverberation samples into light reverberation, medium reverberation and heavy reverberation; A reverberation human voice generation module is used to randomly extract a human voice to be processed and a reverberation to be processed from the classified human voice samples and reverberation samples, subject the human voice to be processed to filtering and muting operations in sequence, and then perform a reverberation operation in combination with the reverberation to be processed to obtain a human voice with reverberation; A target noise generation module is used to randomly extract a noise to be processed from the classified noise samples, and sequentially subject the noise to be processed to filtering, signal-to-noise ratio adjustment and muting operations to obtain a target noise; The noisy human voice generation module is used to generate the noisy human voice as the audio collected by the sound pickup device based on the reverberated human voice and the target noise, for noise reduction model training.
9. The speech enhancement system for AI noise reduction according to claim 8, characterized in that: Also includes, An adaptive reverberation module, used to perform an adaptive reverberation simulation operation on the reverberated human voice, including calculating the time taken for the sound field of the reverberated human voice to decay to a preset decibel value according to different types of reverberations to be processed; If the time taken for the sound field to attenuate the preset decibel value is within the first time threshold range, attenuating the reverberated human voice for a first time length; If the time taken for the sound field to attenuate the preset decibel value is within the second time threshold range, the human voice with reverberation is attenuated for a second time length; if the time taken for the sound field to attenuate the preset decibel value is within the third time threshold range, the human voice with reverberation is attenuated for a third time length; wherein, the first time threshold < the second time threshold < the third time threshold, and the first time length < the second time length < the third time length; the human voice with reverberation is updated.
10. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.