Multi-channel adaptive noise suppression method and electronic equipment thereof
By applying adaptive filtering and vocal active interval marking models in the microphone array, the noise suppression is dynamically adjusted, which solves the problem of poor results of existing noise reduction methods and achieves more efficient and accurate noise suppression.
Patent Information
- Application Number
- CN202510254393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
AI Technical Summary
Existing noise reduction methods mostly rely on static noise models or simple noise filtering technology, and cannot effectively distinguish between noise and useful voice signals, resulting in poor noise reduction and may even lose important voice information.
Audio data is collected through preset microphone arrays, and the adaptive filtering algorithm and the vocal active interval marking model are used to determine the noise reference signal and vocal active interval, dynamically adjust the suppression interval and noise suppression intensity, and perform multi-channel adaptive noise suppression processing.
It effectively reduces the interference of environmental noise on audio data, improves the accuracy of sound processing and noise reduction capabilities, and avoids loss of vocal signals.
Smart Images

Figure CN120148460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sound processing, and particularly to a multi-channel adaptive noise suppression method, device, system, electronic device and its storage medium. Background Art
[0002] With the rapid development of intelligent devices and speech recognition technology, the importance of audio processing in various applications has been increasing day by day. However, in the real environment, audio data is usually interfered by various noise sources, resulting in the degradation of audio quality and affecting the accuracy and effectiveness of subsequent speech recognition, speech interaction and other functions. Especially in complex environments, such as traffic, office or crowded places, the noise problem is particularly prominent.
[0003] Traditional noise reduction methods mostly rely on static noise models or simple noise filtering techniques. These methods often cannot effectively distinguish noise from useful speech signals, resulting in poor noise reduction effects and even possible loss of important speech information. Summary of the Invention
[0004] Embodiments of the present invention provide a multi-channel adaptive noise suppression method to solve the problem that existing noise reduction methods mostly rely on static noise models or simple noise filtering techniques.
[0005] In a first aspect, embodiments of the present invention provide a multi-channel adaptive noise suppression method, the method comprising the following steps: Collect audio data in the current environment through a preset microphone array; Process the audio data through an adaptive filtering algorithm and a vocal activity interval marking model to determine a noise reference signal and a vocal activity interval; Based on the noise reference signal and the vocal activity interval, determine a suppression interval and a corresponding noise suppression intensity; Based on the suppression interval and the corresponding noise suppression intensity, perform suppression processing on the sound data in the current environment and output denoised audio data.
[0006] Optionally, the collecting audio data in the current environment through a preset microphone array includes: Determine the sound collection feature data of two microphones at opposite positions in the preset microphone array, the sound collection feature data including clarity and effective audio duration; Based on the clarity and effective audio duration of the two microphones at opposite positions, determine four corresponding execution microphones in the preset microphone array; Collect sound in the current environment through the four corresponding execution microphones to determine the audio data in the current environment.
[0007] Optionally, the method of receiving sound from the current environment by the four corresponding execution microphones to determine audio data in the current environment further includes: Determine the audio data in the current environment received by the execution microphone; Perform time-domain and frequency-domain alignment processing on the audio data in the current environment received by the execution microphone to obtain multiple pieces of valid audio data in the current environment; Integrate the multiple pieces of valid audio data in the current environment according to a preset granularity to obtain the audio data in the current environment.
[0008] Optionally, the method of processing the audio data by an adaptive filtering algorithm to determine a noise reference signal includes: Obtain the number of channels and the channel positions of the preset microphone array; Based on the number of channels, determine the noise reception intensity of the corresponding audio data; Based on the channel positions, determine the noise direction of the corresponding audio data; Based on the noise reception intensity and the noise direction, perform noise extraction on the audio data to determine the corresponding noise reference signal.
[0009] Optionally, the vocal activity interval marking model includes a DNN model and a VAD marking model. The method of processing the audio data by the vocal activity interval marking model to determine the vocal activity interval includes: Perform vocal prediction on the audio data by the DNN model to determine the vocal probability of the corresponding audio data in the current interval; Mark the vocal probability of the corresponding audio data in the current interval by the VAD marking model to determine at least one audio interval with vocals; Determine the vocal activity interval according to the at least one audio interval with vocals.
[0010] Optionally, the method of determining the suppression interval and the corresponding noise suppression intensity based on the noise reference signal and the vocal activity interval includes: Based on the noise reference signal, determine at least one suppression interval among the multiple vocal activity intervals; According to the intensity of the noise reference signal, calculate the noise suppression intensity for suppressing the current noise reference signal within the at least one suppression interval.
[0011] Optionally, the method of suppressing the sound data in the current environment based on the noise suppression intensity and the suppression interval and outputting the denoised audio data includes: Match the suppression interval with the sound data in the current environment to obtain the sound segments to be suppressed; Process the sound segments to be suppressed according to the noise suppression intensity to obtain the denoised audio data.
[0012] In a second aspect, an embodiment of the present invention further provides a multi-channel adaptive noise suppression device, which includes: A first collection module, configured to collect audio data in the current environment through a preset microphone array; A first processing module, configured to process the audio data through an adaptive filtering algorithm and a vocal activity interval marking model to determine a noise reference signal and a vocal activity interval; A first determination module, configured to determine a suppression interval and a corresponding noise suppression intensity based on the noise reference signal and the vocal activity interval; A first output module, configured to perform suppression processing on the sound data in the current environment based on the suppression interval and the corresponding noise suppression intensity, and output the denoised audio data.
[0013] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the multi-channel adaptive noise suppression method provided by the embodiment of the present invention are implemented.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the multi-channel adaptive noise suppression method provided by the embodiment of the invention are implemented.
[0015] In the embodiment of the present invention, audio data in the current environment is collected through a preset microphone array; the audio data is processed through an adaptive filtering algorithm and a vocal activity interval marking model to determine a noise reference signal and a vocal activity interval; a suppression interval and a corresponding noise suppression intensity are determined based on the noise reference signal and the vocal activity interval; suppression processing is performed on the sound data in the current environment based on the suppression interval and the corresponding noise suppression intensity, and the denoised audio data is output. By processing the environmental audio data through adaptive filtering and a vocal activity interval marking model, noise is suppressed and the denoised audio data is output, reducing the interference of environmental noise on the audio data and improving the accuracy of sound processing and the noise reduction ability. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 is a flowchart of a multi-channel adaptive noise suppression method provided by an embodiment of the present invention; Figure 2 is a schematic structural diagram of another multi-channel adaptive noise suppression device provided by an embodiment of the present invention; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0019] As Figure 1 shown, Figure 1 is a flowchart of a multi-channel adaptive noise suppression method provided by an embodiment of the present invention. The multi-channel adaptive noise suppression method includes the following steps: 101. Collect audio data in the current environment through a preset microphone array.
[0020] In the embodiment of the present invention, the above multi-channel adaptive noise suppression method can be applied to a multi-channel adaptive noise suppression platform. The above multi-channel adaptive noise suppression platform has functions such as noise suppression data processing, noise suppression data transceiver, and noise suppression data memory storage, and can be constructed based on a server or a server cluster. The above server or server cluster can be an electronic device with noise suppression data processing capabilities.
[0021] The above-mentioned preset microphone array can be a system with multiple microphones arranged and combined. Specifically, when the above-mentioned multiple microphones are arranged and combined, their positions can be correspondingly set. For example, when using a five-for-four microphone mode, the microphones can be paired up and down and left and right to form microphone combinations at corresponding positions, and the set microphone combinations can be respectively used for sound collection to obtain the sound collection data of the corresponding combined microphones. Moreover, the sound collection data of the combined microphones can be analyzed, and then the corresponding noise reduction intensity can be selected or the corresponding sound collection weight can be adjusted according to the combination selection to achieve the purpose of dynamic noise reduction.
[0022] The above-mentioned five-for-four mode can be to set four microphones at corresponding positions of up, down, left, and right, and a spare microphone can be set according to requirements. The spare microphone can be set between the microphone combinations at corresponding positions to distinguish the microphone combinations and dynamically adjust the cooperation degree between the microphone combinations. For example, by increasing the sound collection weight of the spare microphone between the microphone combinations, the common sound collection effect between the microphone combinations can be split to achieve the purpose of dynamically adjusting the cooperation degree between the microphone combinations.
[0023] The above-mentioned audio data can include but are not limited to ambient sound data and human voice data. The audio data around the earphone in the current environment can be collected through the above-mentioned preset microphone array to obtain audio data mixed with ambient sound, human voice, and noise, etc. In this embodiment, a microphone array in the five-for-four mode is preferably used, and only five microphones, that is, one spare microphone, can be used to achieve the effect of dynamically adjusting the cooperation degree between the microphone combinations and improving the sound collection accuracy of the microphones.
[0024] 102. Process the audio data through an adaptive filtering algorithm and a human voice active interval marking model to determine the noise reference signal and the human voice active interval.
[0025] In the embodiment of the present invention, the above-mentioned adaptive filtering algorithm can be used to automatically obtain the audio data of multiple channels in this combination according to the combination selection of the above-mentioned preset microphone array, mix the audio data collected at the corresponding positions, and extract the corresponding noise data therefrom to obtain the noise reference signal corresponding to the combination of the preset microphone array.
[0026] Specifically, the number of channels and the channel positions of the preset microphone array can be obtained. According to the selected number of channels, the noise reception intensity of the corresponding required audio data can be calculated. Through the channel positions, that is, the positions of the microphone combinations corresponding to the combination of the preset microphone array, the direction where the noise may occur in the corresponding audio data, that is, the noise direction, can be calculated. Finally, the final noise reference signal can be determined according to the noise reception intensity and the noise direction.
[0027] The above-mentioned vocal activity interval marking model may include, but is not limited to, a DNN model and a VAD marking model. Specifically, the above-mentioned DNN model can be used to predict the probability of the occurrence of vocal data under the current microphone array, that is, to predict the sound collection effect of the selected preset microphone array on vocal data. The above-mentioned VAD marking model can then perform interval marking on the predicted probability, mark the speech intervals with relatively large predicted probabilities, and obtain the corresponding vocal activity intervals. Through the above-mentioned vocal activity interval marking model, the obtained audio data can be processed to predict the vocal activity interval under the current preset microphone array, so that in the specific implementation process of this embodiment, mis-noise reduction operations will not be performed on the vocal activity interval, and the accuracy of noise reduction is improved.
[0028] 103. Based on the noise reference signal and the vocal activity interval, determine the suppression interval and the corresponding noise suppression intensity.
[0029] In the embodiment of the present invention, the above-mentioned suppression interval may be at least one sound interval other than the above-mentioned vocal activity interval. Specifically, there may be noisy audio data with a sufficiently large noise intensity in this suppression interval. Therefore, it is necessary to perform noise suppression on the above-mentioned suppression interval. Then it can be understood that for different suppression intervals, there are noise data with different noise intensities. In order to avoid wasting noise reduction resources, therefore, according to the suppression intervals with different noise intensities, strategies corresponding to the noise suppression intensity can be selected for noise reduction.
[0030] It should be noted that the above-mentioned noise suppression intensity can be dynamically set according to the size, quantity of the above-mentioned suppression interval, and the intensity of the noise reference signal in the corresponding interval. Generally speaking, when it is detected that the intensity of the noise reference signal in the above-mentioned suppression interval is greater, it means that the intensity of noise reduction required is higher. Therefore, the corresponding set noise suppression intensity is greater. It can be understood that since there are intervals with a small suppression interval but a large intensity of the noise reference signal, the noise suppression intensity can also be dynamically set according to the specific situation.
[0031] The above-mentioned noise suppression intensity can be understood as only being aimed at the noisy audio data that needs to be noise-reduced, and no noise reduction processing is performed on the required vocal audio data. However, in possible embodiments, the clarity of the vocal audio data can also be improved by means such as reducing distortion to strengthen the vocal audio data.
[0032] 104. Based on the suppression interval and the corresponding noise suppression intensity, perform suppression processing on the sound data in the current environment and output the noise-reduced audio data.
[0033] In the embodiment of the present invention, the above-mentioned suppression processing can be to process the noisy audio, or to perform fidelity processing on the audio data with vocals, so as to perform noise reduction on the audio data from two aspects.
[0034] Specifically, by comparing the noise reference signal with the human voice active intervals, it is also possible to dynamically determine the suppression intervals within multiple human voice active intervals. For intervals with strong noise interference, the suppression intensity is automatically adjusted according to the corresponding preset noise suppression intensity, and real-time optimization is performed according to the change of the environmental noise, achieving the effect of not losing the clarity of the human voice and being able to suppress the background noise.
[0035] In the embodiments of the present invention, audio data in the current environment is collected through a preset microphone array; the audio data is processed through an adaptive filtering algorithm and a human voice active interval marking model to determine the noise reference signal and the human voice active intervals; based on the noise reference signal and the human voice active intervals, the suppression intervals and the corresponding noise suppression intensities are determined; based on the suppression intervals and the corresponding noise suppression intensities, the sound data in the current environment is suppressed to output the denoised audio data. By processing the environmental audio data through the adaptive filtering and the human voice active interval marking model, the noise is suppressed and the denoised audio data is output, reducing the interference of the environmental noise on the audio data and improving the accuracy of sound processing and the noise reduction ability.
[0036] Optionally, in the step of collecting audio data in the current environment based on the preset microphone array, it is also possible to determine the sound collection characteristic data of two microphones at opposite positions in the preset microphone array; based on the clarity and the effective audio duration of the two microphones at opposite positions, four corresponding execution microphones are determined in the preset microphone array; the current environment is sound-collected through the four corresponding execution microphones to determine the audio data in the current environment.
[0037] In the embodiments of the present invention, the above-mentioned microphones at opposite positions can be, for example, in the above-mentioned preset microphone array setting, in the five-for-four mode, the other four microphones except the spare microphone, and the microphone combination with corresponding positions is obtained by calculating the distance between the two diagonal ends or other arbitrary position determination methods.
[0038] The above-mentioned sound collection characteristic data can include but is not limited to clarity and effective audio duration. Generally speaking, the clarity and the effective audio duration can be adjusted by adjusting the sound collection weight of the microphone. Conversely, the obtained clarity and effective audio duration of the microphone can also be analyzed, and the weight of the microphone can be dynamically adjusted according to the analysis result to enable it to obtain more clear and accurate sound collection characteristic data with effective duration.
[0039] The above-mentioned execution microphone is selected from the preset microphone array after analyzing the sound collection feature data of multiple microphones at opposite positions, and it is a microphone combination that is most suitable for collecting the audio data of the current environment. It can also be a single microphone, and the corresponding number of microphones can be determined according to the specific implementation scheme.
[0040] In a possible embodiment, by analyzing the sound collection feature data of multiple corresponding-position microphones with corresponding positional relationships in the preset microphone array, the execution microphone combination most suitable for sound collection in the current environment can be determined according to the analysis results, and the audio data in the current environment is collected through the selected execution microphones, and a series of preprocessing operations such as filtering, gain control, and noise suppression are performed to extract a clearer audio signal.
[0041] Optionally, in the step of collecting the audio data in the current environment through four execution microphones at corresponding positions, it further includes determining the audio data in the current environment received by the execution microphones; performing time-domain and frequency-domain alignment processing on the audio data in the current environment received by the execution microphones to obtain multiple pieces of effective audio data in the current environment; and integrating the multiple pieces of effective audio data in the current environment according to the preset granularity to obtain the audio data in the current environment.
[0042] In the embodiment of the present invention, the above-mentioned effective audio data can be the audio data with the same delay difference and frequency extracted after the above-mentioned alignment processing. This effective audio data facilitates the following noise reduction processing operations. It can be understood that noise reduction of audio data requires operating on audio data with the same frequency and delay. Therefore, after the above-mentioned alignment processing, the audio data with the same delay difference and frequency can be extracted as the above-mentioned effective audio data.
[0043] The above-mentioned preset granularity can be the minimum analysis unit for analyzing audio data, such as source separation or blind signal separation feature granularity, acoustic feature granularity, etc. Specifically, in this embodiment, it can be based on the audio duration of the time granularity, or the frequency band of the acoustic granularity, etc. The corresponding preset granularity can be selected according to the processing complexity of the audio data, so as to improve the accuracy of audio data processing and reduce its calculation amount.
[0044] The above-mentioned integration processing can be to merge the above-mentioned multiple pieces of effective audio data according to the preset granularity to generate a complete audio data set, which can ensure the coherence and integrity of the audio data and provide a high-quality data source for subsequent analysis.
[0045] The above alignment process can be understood as an operation of performing time-domain alignment and frequency-domain alignment on the above audio data. Specifically, the audio data can be synchronized according to the time axis, for example, aligning the audio signals captured by each microphone on the time axis to avoid time differences caused by transmission delays and other reasons. Frequency alignment adjustment can also be performed on the audio data according to the spectral characteristics of the audio signal to ensure the consistency of its frequency characteristics and improve the accuracy of subsequent analysis.
[0046] By collecting audio data in the environment through a microphone and performing time-domain and frequency-domain alignment processing on it, the consistency of audio signals from different microphones in terms of time and frequency is ensured, forming an effective audio data set. Based on a preset granularity (such as time or acoustic feature granularity), multiple effective audio data are integrated to obtain an effective audio data set of the current environmental audio data, so as to improve the efficiency of subsequent noise reduction processing.
[0047] Optionally, in the step of processing the audio data through an adaptive filtering algorithm to determine the noise reference signal, it further includes obtaining the number of channels and the channel positions of a preset microphone array; determining the noise reception intensity of the corresponding audio data based on the number of channels; determining the noise direction of the corresponding audio data based on the channel positions; and extracting noise from the audio data based on the noise reception intensity and the noise direction to determine the corresponding noise reference signal.
[0048] In the embodiment of the present invention, there is audio input with multiple channels in the preset microphone array. Therefore, after confirming the number of access channels for multiple channels, the intensity of possible noise, that is, the noise reception intensity, is calculated for the microphone weights of each access channel.
[0049] The above channel positions can be determined according to the accessed channels and the corresponding microphone position settings in the preset microphone array. Generally speaking, the noise direction can be the position of the noise source relative to the microphone array, usually calculated through the distance and angle between microphones. The above noise direction information can be used to help identify from which direction the noise comes from the microphone array, thereby optimizing the noise extraction process.
[0050] By combining the noise reception intensity and the noise direction, noise can be extracted from the audio data of each channel. Specifically, the noise reception intensity is used to identify the relative strength of the noise, and then the specific source of the noise is determined through the noise direction, and a corresponding noise reference signal is generated. The above noise reference signal is the noise feature extracted based on the array channel data, which represents the non-speech interference signal in the environment and can be used for subsequent noise suppression or audio enhancement processing.
[0051] Optionally, in the step of processing the audio data through the voice activity interval marking model to determine the voice activity interval, it further includes predicting the voice of the audio data in the current interval through the DNN model to determine the voice probability of the corresponding audio data in the current interval; marking the voice probability of the corresponding audio data in the current interval through the VAD marking model to determine at least one audio interval with voice; and determining the voice activity interval according to at least one audio interval with voice.
[0052] In an embodiment of the present invention, the above voice activity interval marking model may include, but is not limited to, a DNN model and a VAD marking model. Specifically, the above DNN model can be used to predict the probability of the occurrence of voice data under the current microphone array, that is, to predict the sound collection effect of the selected preset microphone array on voice data. The above VAD marking model can then mark the interval of the predicted probability, mark the voice intervals with larger predicted probabilities, and obtain the corresponding voice activity intervals. Through the above voice activity interval marking model, the obtained audio data can be processed to predict the voice activity interval under the current preset microphone array, so that in the specific implementation process of this embodiment, incorrect noise reduction operations will not be performed on the voice activity interval, and the accuracy of noise reduction is improved.
[0053] Optionally, in the step of determining the suppression interval and the corresponding noise suppression intensity based on the noise reference signal and the voice activity interval, it further includes determining at least one suppression interval from multiple voice activity intervals based on the noise reference signal; and calculating the noise suppression intensity for suppressing the current noise reference signal within at least one suppression interval according to the intensity of the noise reference signal.
[0054] In an embodiment of the present invention, the above suppression interval may also refer to the part in the voice activity interval where there is noise interference and no valid voice information is included, usually the high-intensity period of background noise. Specifically, the amplitude, spectral distribution, and time-domain characteristics of the noise reference signal can be used, combined with the time period of the voice activity interval, to determine the noise interference interval.
[0055] The above noise suppression intensity may be the degree of attenuation of the noise signal in the noise interference area. Specifically, the calculation of the noise suppression intensity is referred to according to the amplitude of the noise reference signal (i.e., the noise intensity) and the stability of the noise during this period. For example, if the noise signal is strong and lasts for a long time, the suppression intensity will be increased, while when the noise is weak or appears intermittently, the suppression intensity will be appropriately reduced, so that the voice signal will not be distorted and lost during the noise reduction process.
[0056] In this embodiment, by identifying the noise interference interval and dynamically adjusting the noise suppression intensity, the influence of background noise is reduced. By combining the noise reference signal and the human voice active interval, the suppression interval can be determined, unnecessary signal processing is avoided, and the clarity of the voice signal and the efficiency of noise reduction are improved.
[0057] Optionally, in the step of suppressing the sound data in the current environment based on the noise suppression intensity and the suppression interval and outputting the denoised audio data, it further includes matching the suppression interval with the sound data in the current environment to obtain the sound segment to be suppressed; processing the sound segment to be suppressed according to the noise suppression intensity to obtain the denoised audio data.
[0058] In the embodiment of the present invention, by comparing the time stamps of the above suppression interval and the time axis of the above sound data, the segment containing noise is located in the audio data. It should be noted that it usually corresponds to the overlapping part of the high-intensity noise region determined in the noise reference signal and the human voice active interval.
[0059] After matching, the corresponding noise suppression intensity is calculated according to the intensity of the noise reference signal in the suppression interval. Specifically, a stronger noise signal corresponds to a higher suppression intensity, and a weaker noise signal uses a lower suppression intensity. The gain of the audio signal can be adjusted by an adaptive algorithm to attenuate the noise part in the suppression interval, reduce the interference of background noise on the voice signal, and at the same time retain the clarity of the human voice.
[0060] In another possible embodiment, the above multi-channel adaptive noise suppression method can set the corresponding earpad position according to the preset microphone array applied to the mobile earphone, and dynamically adjust the corresponding microphone weights, so as to ensure the consistency of the noise reduction performance.
[0061] For example, when the wireless earphone is in the normal wearing state, a conventional noise reduction mode is adopted, the microphone array in the current wearing state is adjusted and selected, and then the above audio collection and subsequent operations are completed using the microphone array in the current wearing state.
[0062] Through the corresponding microphone arrays in the above different wearing states, different microphone weights are selected to meet the sound collection requirements in the current environment and wearing state, ensuring that accurate denoised audio data can be obtained subsequently, achieving excellent noise reduction effects at various wearing angles, and applying DNN and VAD to make it have better and more accurate noise recognition and elimination effects.
[0063] As Figure 2 shown, the embodiment of the present invention also provides a multi-channel adaptive noise suppression device 200, and the multi-channel adaptive noise suppression device 200 includes: The first collection module 201 is configured to collect audio data in the current environment through a preset microphone array; The first processing module 202 is configured to process the audio data through an adaptive filtering algorithm and a voice activity interval marking model to determine a noise reference signal and a voice activity interval; The first determination module 203 is configured to determine an inhibition interval and a corresponding noise suppression intensity based on the noise reference signal and the voice activity interval; The first output module 204 is configured to perform an inhibition process on the sound data in the current environment based on the inhibition interval and the corresponding noise suppression intensity, and output denoised audio data.
[0064] Optionally, the above-mentioned first collection module 201 includes: The first determination sub-module is configured to determine the sound collection feature data of two microphones at opposite positions in the preset microphone array, and the sound collection feature data includes clarity and effective audio duration; The second determination sub-module is configured to determine four corresponding execution microphones in the preset microphone array based on the clarity and effective audio duration of the two microphones at opposite positions; The third determination sub-module is configured to perform a sound collection process on the current environment through the four corresponding execution microphones to determine the audio data in the current environment.
[0065] Optionally, the above-mentioned third determination sub-module includes: The first determination unit is configured to determine the audio data in the current environment received by the execution microphone; The second determination unit is configured to perform time-domain and frequency-domain alignment processing on the audio data in the current environment received by the execution microphone to obtain a plurality of effective audio data in the current environment; The third determination unit is configured to integrate the plurality of effective audio data in the current environment according to a preset granularity to obtain the audio data in the current environment.
[0066] Optionally, the above-mentioned first processing module 202 includes: The first acquisition sub-module is configured to acquire the number of channels and the channel positions of the preset microphone array; The fourth determination sub-module is configured to determine the noise reception intensity of the corresponding audio data based on the number of channels; The fifth determination sub-module is configured to determine the noise direction of the corresponding audio data based on the channel positions; The sixth determination sub-module is configured to perform noise extraction on the audio data based on the noise reception intensity and the noise direction to determine the corresponding noise reference signal.
[0067] Optionally, the above first processing module 202 further includes: A seventh determination sub-module, configured to perform a voice prediction on the audio data through the DNN model, and determine the voice probability of the corresponding audio data within the current interval; An eighth determination sub-module, configured to mark the voice probability of the corresponding audio data within the current interval through the VAD marking model, and determine at least one audio interval with voice; A ninth determination sub-module, configured to determine a voice active interval according to the at least one audio interval with voice.
[0068] Optionally, the above first determination module 203 includes: A tenth determination sub-module, configured to determine at least one suppression interval from the multiple voice active intervals based on the noise reference signal; An eleventh determination sub-module, configured to calculate a noise suppression intensity for suppressing the current noise reference signal within the at least one suppression interval according to the intensity of the noise reference signal.
[0069] Optionally, the above first output module 204 includes: A twelfth determination sub-module, configured to match the suppression interval with the sound data in the current environment to obtain a sound segment to be suppressed; A thirteenth determination sub-module, configured to process the sound segment to be suppressed according to the noise suppression intensity to obtain denoised audio data.
[0070] As Figure 3 shown, an embodiment of the present invention further provides an electronic device 300, including a processor, and the above processor can execute any one of the above multi-channel adaptive noise suppression methods.
[0071] Specifically, it includes a processor 301 and a memory 302, and a computer program for executing the multi-channel adaptive noise suppression method stored on the memory 302 and capable of running on the processor 301, where: The processor 301 runs the calculator program of the multi-channel adaptive noise suppression method stored in the memory 302 and executes the following steps: Collect audio data in the current environment through a preset microphone array; Process the audio data through an adaptive filtering algorithm and a voice active interval marking model to determine a noise reference signal and a voice active interval; Based on the noise reference signal and the voice active interval, determine a suppression interval and a corresponding noise suppression intensity; Based on the suppression interval and the corresponding noise suppression intensity, perform suppression processing on the sound data in the current environment, and output the denoised audio data.
[0072] Optionally, the processor 301 executes collecting the audio data in the current environment based on the preset microphone array, including: Determine the sound collection feature data of the microphones at two opposite positions in the preset microphone array, where the sound collection feature data includes clarity and effective audio duration; Based on the clarity and effective audio duration of the microphones at the two opposite positions, determine four corresponding execution microphones in the preset microphone array; Through the four corresponding execution microphones, perform sound collection processing on the current environment to determine the audio data in the current environment.
[0073] Optionally, when the processor 301 executes performing sound collection processing on the current environment through the four corresponding execution microphones to determine the audio data in the current environment, the method further includes: Determine the audio data in the current environment received by the execution microphones; Perform time-domain and frequency-domain alignment processing on the audio data in the current environment received by the execution microphones to obtain multiple pieces of effective audio data in the current environment; Integrate the multiple pieces of effective audio data in the current environment according to the preset granularity to obtain the audio data in the current environment.
[0074] Optionally, when the processor 301 executes processing the audio data through the adaptive filtering algorithm to determine the noise reference signal, it includes: Obtain the number of channels and the channel positions of the preset microphone array; Based on the number of channels, determine the noise reception intensity of the corresponding audio data; Based on the channel positions, determine the noise direction of the corresponding audio data; Based on the noise reception intensity and the noise direction, perform noise extraction on the audio data to determine the corresponding noise reference signal.
[0075] Optionally, the human voice active interval marking model executed by the processor 301 includes a DNN model and a VAD marking model. When processing the audio data through the human voice active interval marking model to determine the human voice active interval, it includes: Perform human voice prediction on the audio data through the DNN model to determine the human voice probability of the corresponding audio data in the current interval; Mark the voice probability of the corresponding audio data in the current interval through the VAD marking model to determine at least one audio interval with voice. Determine the voice active interval according to the at least one audio interval with voice.
[0076] Optionally, the processor 301 also executes to determine the suppression interval and the corresponding noise suppression intensity based on the noise reference signal and the voice active interval, including: Based on the noise reference signal, determine at least one suppression interval among the multiple voice active intervals; According to the intensity of the noise reference signal, calculate the noise suppression intensity for suppressing the current noise reference signal within the at least one suppression interval.
[0077] Optionally, the processor 301 also executes to perform suppression processing on the sound data in the current environment based on the noise suppression intensity and the suppression interval, and output the denoised audio data, including: matching the suppression interval with the sound data in the current environment to obtain the sound segment to be suppressed; Process the sound segment to be suppressed according to the noise suppression intensity to obtain the denoised audio data.
[0078] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes each process of the multi-channel adaptive noise suppression method or the application-side multi-channel adaptive noise suppression method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0079] Those of ordinary skill in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0080] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A multi-channel adaptive noise suppression method, characterized in that: include: Collect audio data in the current environment through a preset microphone array; The audio data is processed by an adaptive filtering algorithm and a human voice active interval marking model to determine a noise reference signal and a human voice active interval; Determine a suppression interval and a corresponding noise suppression intensity based on the noise reference signal and the human voice active interval; Based on the suppression interval and the corresponding noise suppression intensity, the sound data in the current environment is suppressed and the noise-reduced audio data is output.
2. The multi-channel adaptive noise suppression method according to claim 1, characterized in that: The method of collecting audio data in the current environment based on a preset microphone array includes: Determine sound collection feature data of two microphones at opposite positions in a preset microphone array, wherein the sound collection feature data includes clarity and effective audio duration; Based on the clarity and effective audio duration of the two microphones at opposite positions, determining four execution microphones at corresponding positions in a preset microphone array; The sound collection process of the current environment is performed through the execution microphones at the four corresponding positions to determine the audio data in the current environment.
3. The multi-channel adaptive noise suppression method according to claim 2, characterized in that: The method further comprises: performing sound collection processing on the current environment by using the execution microphones at the four corresponding positions to determine the audio data in the current environment: Determining audio data in a current environment received by the execution microphone; Performing time domain and frequency domain alignment processing on the audio data in the current environment received by the execution microphone to obtain multiple valid audio data in the current environment; According to a preset granularity, the valid audio data in the multiple current environments are integrated to obtain the audio data in the current environment.
4. The multi-channel adaptive noise suppression method according to claim 2, characterized in that: The step of processing the audio data by an adaptive filtering algorithm to determine a noise reference signal includes: Obtain the number of channels and channel positions of the preset microphone array; Based on the number of channels, determining a noise reception intensity corresponding to the audio data; Based on the channel position, determining a noise direction corresponding to the audio data; Based on the noise reception intensity and the noise direction, noise extraction is performed on the audio data to determine a corresponding noise reference signal.
5. The multi-channel adaptive noise suppression method according to claim 1, characterized in that: The vocal active interval marking model includes a DNN model and a VAD marking model. The audio data is processed by the vocal active interval marking model to determine the vocal active interval, including: Performing human voice prediction on the audio data by using the DNN model to determine the human voice probability of the corresponding audio data in the current interval; Marking the human voice probability of the audio data corresponding to the current interval by the VAD marking model, and determining at least one audio interval having human voice; A human voice active interval is determined according to the at least one audio interval having human voice.
6. The multi-channel adaptive noise suppression method according to claim 1, characterized in that: The determining, based on the noise reference signal and the human voice active interval, a suppression interval and a corresponding noise suppression intensity includes: Based on the noise reference signal, determining at least one suppression interval among the plurality of vocal active intervals; According to the strength of the noise reference signal, a noise suppression strength for suppressing the current noise reference signal is calculated within the at least one suppression interval.
7. The multi-channel adaptive noise suppression method according to claim 1, characterized in that: The suppressing process is performed on the sound data in the current environment based on the noise suppression intensity and the suppression interval, and outputting the noise-reduced audio data, including: Matching the suppression interval with the sound data in the current environment to obtain a sound segment to be suppressed; The sound segment to be suppressed is processed according to the noise suppression strength to obtain noise-reduced audio data.
8. A multi-channel adaptive noise suppression device, characterized in that: include: A first collection module, used to collect audio data in the current environment through a preset microphone array; A first processing module, configured to process the audio data by using an adaptive filtering algorithm and a human voice active interval marking model to determine a noise reference signal and a human voice active interval; A first determination module, configured to determine a suppression interval and a corresponding noise suppression intensity based on the noise reference signal and a human voice active interval; The first output module is used to suppress the sound data in the current environment based on the suppression interval and the corresponding noise suppression intensity, and output the noise-reduced audio data.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the multi-channel adaptive noise suppression method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the multi-channel adaptive noise suppression method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Method and system for testing noise reduction performance of intelligent dynamic diversion anti-noise earphone
CN121151785A