A noise suppression method, system, microphone and medium for microphone

By obtaining the background noise energy of the microphone input audio signal and the spectrum characteristics of the voice signal, performing speech quality analysis and segmentation processing, and dynamically adjusting the noise suppression intensity, the problem of poor noise suppression effect in the prior art is solved, and better noise suppression effect and voice signal fidelity are achieved.

CN119052692BActive Publication Date: 2025-05-16ENPING ART STAR ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411122934.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-05-16
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

Existing noise suppression techniques cannot adaptively adjust the suppression intensity when dealing with complex background noise, resulting in poor suppression effect or distortion.

Method used

By obtaining the background noise energy of the microphone input audio signal and the spectrum characteristics of the voice signal, performing speech quality analysis, segmentation processing and directional noise suppression analysis, dynamically adjusting the noise suppression intensity, and achieving adaptive segmented noise suppression.

Benefits of technology

It effectively improves the noise suppression effect, reduces distortion and recognition errors of voice signals, and improves the performance of the voice system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052692B_ABST
    Figure CN119052692B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of microphones, and in particular to a noise suppression method, system, microphone and medium for microphones. The present application first obtains the background noise energy and speech spectrum characteristics of the microphone input audio, and analyzes the quality of the speech signal according to the background noise energy and speech spectrum characteristics; then, according to the quality analysis results, the entire audio is segmented, and a directional noise suppression analysis is performed on the local audio segments with unqualified quality to obtain the corresponding directional suppression parameters; then, according to the directional suppression parameters, the noise suppression strength of the local audio segments is dynamically adjusted, and new suppression process parameters are generated and sent to the noise suppression module to achieve adaptive segmented noise suppression, which can achieve better noise suppression effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of microphones, and in particular to a noise suppression method, system, microphone and medium for a microphone. Background Art

[0002] With the continuous development of voice interaction technology, voice recognition systems and voice communication devices have been increasingly widely used in our daily lives. However, in actual application scenarios, the presence of background noise will seriously affect the quality of voice signals, resulting in increased voice recognition error rates or decreased communication quality. Therefore, effective noise suppression technology is crucial to the performance of voice systems.

[0003] When dealing with complex background noise, existing noise suppression technology cannot adaptively adjust the suppression strength according to changes in speech signals and the characteristics of noise, resulting in poor suppression effect or distortion. This situation needs to be further improved. Summary of the invention

[0004] In order to solve the problem of distortion in existing noise suppression technology when processing complex background noise, the present application provides a noise suppression method, system, microphone and medium for a microphone, which adopts the following technical solutions:

[0005] In a first aspect, the present application provides a noise suppression method for a microphone, comprising the following steps:

[0006] Obtaining background noise energy of microphone input audio signal and frequency spectrum characteristics of speech signal;

[0007] Analyzing the quality of the speech signal according to the background noise energy and the frequency spectrum characteristics to obtain a speech quality analysis result;

[0008] According to the speech quality analysis result, the microphone input audio signal is processed in segments, and a directional noise suppression analysis is performed on the local audio segment with unqualified speech quality to obtain corresponding directional suppression parameters;

[0009] The noise suppression strength of the local audio segment is adjusted according to the directional suppression parameter, and the suppression process parameters of the local audio segment are generated according to the adjusted noise suppression strength and sent to the noise suppression module, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameters.

[0010] By adopting the above technical solution, the presence of background noise will seriously affect the quality of the voice signal, resulting in communication distortion or recognition errors; the present application first obtains the background noise energy and voice spectrum characteristics of the microphone input audio, and analyzes the voice signal quality according to the background noise energy and voice spectrum characteristics; then, according to the quality analysis results, the entire audio is segmented, and a directional noise suppression analysis is performed on the local audio segment with unqualified quality to obtain the corresponding directional suppression parameters; then, according to the directional suppression parameters, the noise suppression intensity of the local audio segment is dynamically adjusted, and new suppression process parameters are generated and sent to the noise suppression module to achieve adaptive segmented noise suppression, which can achieve better noise suppression effect.

[0011] Optionally, analyzing the quality of the speech signal according to the background noise energy and the spectrum characteristics to obtain a speech quality analysis result specifically includes the following steps:

[0012] Calculating the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and the frequency spectrum characteristics;

[0013] Calculate the speech component ratio corresponding to each time-frequency point according to the signal-to-noise ratio;

[0014] Analyzing the distortion degree between the current speech and the preset clear speech according to the speech component proportion, and obtaining the single-point speech distortion value of the current time-frequency point;

[0015] The comprehensive speech distortion value of the entire speech signal is estimated according to the single-point speech distortion value, and the quality analysis of the speech signal is performed according to the comprehensive speech distortion value to obtain the quality analysis result of the speech signal.

[0016] By adopting the above technical solution, since the voice signal will be contaminated by various noises during the transmission and collection process, resulting in its quality degradation, it is necessary to accurately evaluate the quality of the voice signal to provide a basis for subsequent noise suppression; the present application first calculates the signal-to-noise ratio of each frequency band according to the background noise energy and the voice spectrum characteristics, and then derives the proportion of the voice component at each time-frequency point; then compares the current voice with the preset clear voice, evaluates the degree of distortion between the two, and obtains the single-point voice distortion value of each time-frequency point; finally, based on these single-point distortion values, estimates the comprehensive voice distortion value of the entire voice signal, and analyzes and judges the overall quality of the voice signal according to the distortion value, which helps to improve the accuracy and reliability of the analysis.

[0017] Optionally, the method further comprises:

[0018] Acquire the noise energy difference between adjacent audio segments, and acquire the speech amplitude difference and the noise amplitude value corresponding to the speech component proportion according to the noise energy difference;

[0019] Constructing a speech change curve corresponding to the noise energy difference according to the speech amplitude difference and the noise amplitude value;

[0020] Analyzing the attenuation trend of the speech according to the speech change curve to obtain the speech attenuation value of the current noise environment;

[0021] When the actual amplitude of the speech reaches the speech attenuation value, the gain compensation information is output to the noise suppression module.

[0022] By adopting the above technical solution, relying solely on noise suppression cannot completely preserve the quality and details of the speech signal. The present application first obtains the noise energy difference between adjacent audio segments, and derives the amplitude information of the speech and noise based on the noise energy difference; then uses the amplitude information to construct a speech change curve, analyzes the attenuation trend of the speech, and obtains the speech attenuation value in the current noise environment; once the actual speech amplitude drops to the attenuation value, the gain compensation information is output to the noise suppression module, thereby effectively compensating for the speech attenuation and ensuring the stability of the speech quality in a complex noise environment.

[0023] Optionally, according to the voice quality analysis result, the microphone input audio signal is processed in segments, and a directional noise suppression analysis is performed on a local audio segment with unqualified voice quality to obtain a corresponding directional suppression parameter, which specifically includes the following steps:

[0024] According to the speech quality analysis result, obtaining the speech quality difference between adjacent audio segments;

[0025] Segmentally process the microphone input audio signal according to the voice quality difference to generate local audio segments with the same voice quality;

[0026] Marking a local audio segment whose voice quality is lower than a preset quality threshold as unqualified, and performing noise analysis on the marked unqualified audio segment to obtain noise characteristic parameters of the unqualified audio segment;

[0027] According to the noise characteristic parameters, the noise suppression intensity corresponding to the unqualified audio segment and the directional suppression process corresponding to each frequency band are analyzed to obtain the directional suppression parameters of the unqualified audio segment.

[0028] By adopting the above technical solution, since the speech signal will be affected by noise to different degrees in different time periods, the effect of adopting a unified suppression strategy is not good. The present application first segments the entire audio signal according to the speech quality analysis results to obtain local audio segments of the same quality level; then marks the unqualified audio segments whose quality is lower than the threshold, analyzes their noise characteristics, derives the corresponding noise suppression strength and directional suppression process, and obtains the directional suppression parameters; finally, applies the directional suppression parameters to the noise suppression module, and implements targeted directional noise suppression on the low-quality audio segments, so as to better eliminate noise and avoid excessive distortion.

[0029] Optionally, adjusting the noise suppression strength of the local audio segment according to the directional suppression parameter, generating and sending the suppression process parameter of the local audio segment to the noise suppression module according to the adjusted noise suppression strength, specifically includes the following steps:

[0030] Acquire a noise frequency range and a corresponding noise energy of the current noise according to the directional suppression parameter, and analyze a current noise intensity of the current noise according to the noise frequency range and the noise energy;

[0031] Calculating a deviation value between the directional suppression parameter and the current noise intensity, and adjusting the noise suppression intensity of the local audio segment according to the deviation value;

[0032] Acquire the current speech feature of the local audio segment, analyze the suppression effect of each suppression frequency band on the current speech feature layer by layer, and obtain a suppression frequency sequence;

[0033] According to the noise suppression intensity and the suppression frequency sequence, parameters of an original noise suppression process of the local audio segment are adjusted to generate suppression process parameters corresponding to the local audio segment.

[0034] By adopting the above technical solution, the present application first analyzes the frequency range and intensity of the current noise according to the previously obtained directional suppression parameters, and calculates the deviation of the current noise intensity from the expected value; then dynamically adjusts the noise suppression intensity of the current audio segment according to the deviation value; then analyzes the characteristics of the current speech, evaluates the impact of each suppressed frequency band on the speech, and obtains a suppressed frequency sequence; finally, combines the adjusted suppression intensity and frequency sequence to generate suppression process parameters for the current audio segment, making the parameter adjustment more accurate.

[0035] Optionally, calculating a deviation value between the directional suppression parameter and the current noise intensity, and adjusting the noise suppression intensity of the local audio segment according to the deviation value specifically includes the following steps:

[0036] estimating, according to a deviation value between the directional suppression parameter and the current noise intensity, a degree of speech quality loss caused by the current noise intensity to the local audio segment without noise suppression processing;

[0037] Obtaining an actual speech attenuation degree of the local audio segment under current noise energy;

[0038] Comparing the voice quality loss degree with the actual voice attenuation degree, and calculating the difference between the voice quality loss degree and the actual voice attenuation degree;

[0039] Determining whether the current noise suppression strength is reasonable according to the difference, and predicting the noise suppression effect of the local audio segment;

[0040] According to the noise suppression effect, the noise suppression intensity of each suppressed frequency band of the local audio segment is adjusted.

[0041] By adopting the above technical solution, the present application first estimates the potential loss of speech quality that noise will cause to the voice without suppression based on the deviation between the current noise and the expected value; then obtains the attenuation degree of the actual voice under noise; then compares the quality loss and the actual attenuation degree. If there is a difference between the two, it means that the current suppression strength is unreasonable and needs to be adjusted; finally, the suppression effect is predicted based on the difference, and the suppression strength of each frequency band is dynamically adjusted to ensure that the noise suppression is always maintained in a better state, effectively avoiding excessive or insufficient suppression.

[0042] Optionally, according to the noise suppression intensity and the suppression frequency sequence, adjusting parameters of an original noise suppression process of the local audio segment to generate suppression process parameters corresponding to the local audio segment specifically includes the following steps:

[0043] Obtaining the phase distortion of each suppressed frequency band under the current ambient temperature, and calculating the superposition delay time of adjacent suppressed frequency bands according to the phase distortion and the corresponding noise suppression strength;

[0044] Calculating the superposition interference ratio of the adjacent suppression frequency bands according to the superposition delay time and the corresponding suppression frequency sequence;

[0045] adjusting the single-frequency suppression bandwidth corresponding to the adjacent suppression frequency band according to the superimposed interference ratio, and calculating the suppression balance coefficient for achieving frequency balance of the adjacent suppression frequency band according to the adjusted suppression bandwidth;

[0046] According to the suppression balance coefficient, parameters of the original noise suppression process of the local audio segment are adjusted to generate suppression process parameters corresponding to the local audio segment.

[0047] By adopting the above technical solution, due to the problem of mutual superposition and interference between different frequency bands, if the adjusted noise suppression strength and frequency sequence are simply directly applied to the original process, it may cause an imbalance between frequency bands, thereby affecting the overall effect of noise suppression; the application first obtains the phase distortion of each suppressed frequency band at the current ambient temperature, and calculates the superposition delay time of adjacent frequency bands according to the phase distortion and suppression strength; then, based on the delay time and frequency sequence, the superposition interference ratio of adjacent frequency bands is derived; then, the suppression bandwidth of each frequency band is adjusted according to the interference ratio to achieve frequency balance; finally, based on the balanced bandwidth, the various parameters of the original process are corrected to generate new suppression process parameters to reduce the distortion and noise caused by frequency band imbalance.

[0048] In a second aspect, the present application provides a noise suppression system for a microphone, comprising:

[0049] A data acquisition module, used to acquire the background noise energy of the microphone input audio signal and the frequency spectrum characteristics of the speech signal;

[0050] A quality analysis module, used to analyze the quality of the speech signal according to the background noise energy and the spectrum characteristics to obtain a speech quality analysis result;

[0051] A suppression analysis module, configured to perform segmented processing on the microphone input audio signal according to the voice quality analysis result, perform directional noise suppression analysis on the local audio segment with unqualified voice quality, and obtain corresponding directional suppression parameters;

[0052] The noise suppression module is used to adjust the noise suppression strength of the local audio segment according to the directional suppression parameter, generate and send the suppression process parameters of the local audio segment to the noise suppression module according to the adjusted noise suppression strength, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameters.

[0053] In a third aspect, the present application provides a microphone, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned noise suppression method for a microphone when executing the computer program.

[0054] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-mentioned noise suppression method for a microphone.

[0055] In summary, the present application includes at least one of the following beneficial technical effects:

[0056] 1. This application first obtains the background noise energy and speech spectrum characteristics of the microphone input audio, and analyzes the quality of the speech signal according to the background noise energy and speech spectrum characteristics; then, according to the quality analysis results, the entire audio is segmented, and a directional noise suppression analysis is performed on the local audio segment with unqualified quality to obtain the corresponding directional suppression parameters; then, according to the directional suppression parameters, the noise suppression strength of the local audio segment is dynamically adjusted, and a new suppression process parameter is generated and sent to the noise suppression module to achieve adaptive segmented noise suppression, which can achieve better noise suppression effect;

[0057] 2. This application first calculates the signal-to-noise ratio of each frequency band based on the background noise energy and the speech spectrum characteristics, and then derives the speech component ratio of each time-frequency point; then compares the current speech with the preset clear speech, evaluates the degree of distortion between the two, and obtains the single-point speech distortion value of each time-frequency point; finally, based on these single-point distortion values, estimates the comprehensive speech distortion value of the entire speech signal, and analyzes and judges the overall quality of the speech signal based on the distortion value, which helps to improve the accuracy and reliability of the analysis;

[0058] 3. This application first obtains the phase distortion of each suppressed frequency band under the current ambient temperature, and calculates the superposition delay time of adjacent frequency bands according to the phase distortion and suppression strength; then, according to the delay time and frequency sequence, the superposition interference ratio of adjacent frequency bands is derived; then, according to the interference ratio, the suppression bandwidth of each frequency band is adjusted to achieve frequency balance; finally, based on the balanced bandwidth, various parameters of the original process are corrected to generate new suppression process parameters to reduce distortion and noise caused by frequency band imbalance. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of a noise suppression method for a microphone according to an embodiment of the present application;

[0060] Figure 2 is a flow chart of step S200 in a noise suppression method for a microphone according to an embodiment of the present application;

[0061] Figure 3 is a flow chart of step S300 in a noise suppression method for a microphone according to an embodiment of the present application;

[0062] Figure 4 is a flow chart of step S400 in a noise suppression method for a microphone according to an embodiment of the present application;

[0063] Figure 5 is a flow chart of step S420 in a noise suppression method for a microphone according to an embodiment of the present application;

[0064] Figure 6is a flow chart of step S440 in a noise suppression method for a microphone according to an embodiment of the present application;

[0065] Figure 7 is another flow chart of a noise suppression method for a microphone according to an embodiment of the present application;

[0066] Figure 8 is a module schematic diagram of a noise suppression system for a microphone according to an embodiment of the present application;

[0067] Fig. 9 This is a diagram of the internal structure of a microphone according to an embodiment of the present application. DETAILED DESCRIPTION

[0068] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more listed items.

[0069] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.

[0070] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.

[0071] In a first aspect, the present application provides a noise suppression method for a microphone, referring to Figure 1 , including the following steps:

[0072] S100: Obtaining background noise energy of a microphone input audio signal and frequency spectrum characteristics of a speech signal.

[0073] The background noise energy refers to the sum of the acoustic energy released by various noise sources in the environment, which can be separated from the microphone input audio signal by spectrum analysis and other methods. In this embodiment, a time-frequency analysis method such as wavelet transform or short-time Fourier transform can be used to decompose the input audio and extract the noise energy component.

[0074] Specifically, a silent time window without speech input is pre-set, and the energy value of each frequency band in the silent time window is calculated as the estimated value of the background noise energy. The spectral characteristics of the speech signal describe the energy distribution law of the speech at different frequencies. In this embodiment, the amplitude spectrum of each frequency band is obtained as the spectral characteristics by performing short-time Fourier transform on the time window with speech input.

[0075] S200: Analyze the quality of the speech signal according to the background noise energy and spectrum characteristics to obtain a speech quality analysis result.

[0076] The speech quality analysis is based on the frequency domain characteristics of noise and speech, and evaluates the reliability and fidelity of the current speech signal in a noisy environment. In this embodiment, the background noise energy is compared with the speech spectrum characteristics to analyze the degree of masking of the noise on the speech in different frequency bands.

[0077] Specifically, the noise intensity of each frequency band is first estimated based on the background noise energy, and then compared with the amplitude value of the corresponding frequency band in the speech spectrum feature. When the noise intensity exceeds a certain ratio threshold of the speech amplitude, the speech quality of the frequency band is judged to be unqualified. Finally, a speech quality analysis matrix can be obtained, which reflects the quality distribution of the speech signal in the entire frequency band.

[0078] S300: According to the result of the speech quality analysis, the microphone input audio signal is processed in segments, and a directional noise suppression analysis is performed on the local audio segment with unqualified speech quality to obtain corresponding directional suppression parameters.

[0079] Among them, segmentation processing refers to dividing the entire audio signal into multiple local segments with the same voice quality attributes according to the quality analysis results; in this embodiment, the voice quality analysis matrix is ​​scanned along the time axis, and when the quality attributes of adjacent time windows change, the dividing point of a new segment is determined.

[0080] Specifically, for the local audio segment marked as unqualified, directional noise suppression analysis is performed to obtain the directional suppression parameters of the segment. By counting the frequency bands with poor voice quality in the segment, the corresponding noise frequency range and energy distribution, and combining the actual scene such as temperature, environment and other factors for comprehensive analysis, the noise suppression strength, target suppression frequency and corresponding process parameters for the segment are derived and output as directional suppression parameters.

[0081] S400, adjusting the noise suppression strength of the local audio segment according to the directional suppression parameter, generating a suppression process parameter of the local audio segment according to the adjusted noise suppression strength and sending it to the noise suppression module, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameter.

[0082] Specifically, the noise suppression strength of the local audio segment is adjusted according to the directional suppression parameter, and the suppression process parameters are generated according to the adjusted noise suppression strength, so that the noise suppression process is adjusted according to the suppression process parameters to control the noise.

[0083] In one embodiment, referring to Figure 2 In step S200, the quality of the speech signal is analyzed according to the background noise energy and spectrum characteristics to obtain a speech quality analysis result, which specifically includes the following steps:

[0084] S210: Calculate the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and spectrum characteristics.

[0085] The signal-to-noise ratio refers to the ratio of speech signal energy to background noise energy, and is used to evaluate the significance of the corresponding frequency band speech volume relative to the noise. In this embodiment, the signal-to-noise ratio value of each frequency band is calculated by comparing the energy value of the background noise energy spectrum with the speech spectrum feature in each frequency band.

[0086] Specifically, short-time Fourier transform is first performed on the background noise energy and speech spectrum features to obtain their energy distribution matrices in the time-frequency domain. Then, in each frequency band, the corresponding element of the speech energy matrix is ​​divided by the element of the noise energy matrix, and the logarithm is taken to obtain the signal-to-noise ratio value of the frequency band.

[0087] S220: Calculate the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and spectrum characteristics.

[0088] The speech component ratio refers to the importance of the speech signal at the corresponding time-frequency point, and is the basis for speech quality assessment. In this embodiment, the speech component ratio of each time-frequency is estimated based on the signal-to-noise ratio value of each frequency band and the integrated speech model.

[0089] Specifically, the signal-to-noise ratio value of each frequency band is normalized, and then the normalized value is mapped to the corresponding voice component ratio value according to the pre-established voice model. For example, the voice / noise probability distribution estimation method of the Gaussian mixture model is used, and the signal-to-noise ratio value is substituted into the model to obtain the voice ratio estimation value of the time-frequency point.

[0090] S230: Analyze the distortion degree between the current speech and the preset clear speech according to the speech component ratio, and obtain the single-point speech distortion value of the current time-frequency point.

[0091] The speech distortion value is used to quantify the distortion degree of the current speech relative to the ideal clear speech. In this embodiment, the speech component ratio is compared with the preset clear speech template, and the deviation value between the two at each time-frequency point is evaluated, which is the single-point speech distortion value.

[0092] Specifically, first establish a corresponding clear speech template, in which each time-frequency point corresponds to an ideal speech proportion reference value, and then compare the speech component proportion matrix of the current speech signal with the clear speech template, and the absolute value of the difference between the two is the single-point speech distortion value of the corresponding time-frequency point.

[0093] S240, estimating a comprehensive speech distortion value of the entire speech signal according to the single-point speech distortion value, performing quality analysis on the speech signal according to the comprehensive speech distortion value, and obtaining a quality analysis result of the speech signal.

[0094] The comprehensive speech distortion value is the overall distortion degree of the entire speech signal, which is used to finally evaluate whether the speech quality is qualified. In this embodiment, the single-point distortion values ​​of all time-frequency points are weighted and summed to obtain the comprehensive distortion value of the speech segment.

[0095] Specifically, the single-point distortion value of each time-frequency point is first weighted, and the weight is set based on the voice proportion or importance of each time-frequency point. Then the weighted distortion values ​​are accumulated to obtain a comprehensive distortion value. Finally, based on the comparison between the comprehensive distortion value and the preset threshold, the voice signal is graded to generate a quality analysis result.

[0096] In one embodiment, referring to Figure 3 In step S300, according to the result of speech quality analysis, the microphone input audio signal is processed in segments, and the local audio segment with unqualified speech quality is subjected to directional noise suppression analysis to obtain the corresponding directional suppression parameters, which specifically includes the following steps:

[0097] S310: Obtain the speech quality difference between adjacent audio segments according to the speech quality analysis result.

[0098] The difference in speech quality refers to the difference in speech quality scores between two adjacent audio segments on the time axis. In this embodiment, the observation window is slid along the time axis to calculate the variation in quality scores between adjacent windows.

[0099] Specifically, the speech quality analysis result is represented as a time series, and each point corresponds to the quality score of a time window; then a sliding window mechanism is used to slide on the sequence, and the quality difference between adjacent windows is calculated as the estimated value of the speech quality difference.

[0100] S320: Segment the microphone input audio signal according to the difference in voice quality to generate local audio segments with the same voice quality.

[0101] In this embodiment, the entire audio signal is divided into a plurality of local segments having the same speech quality attribute according to the change of the quality difference value.

[0102] Specifically, a quality difference threshold is set in advance. When the quality difference of adjacent windows exceeds the threshold, it is considered that this is a demarcation point of a new segment. The traversal starts from the beginning of the audio. When encountering a demarcation point, the current time point is recorded, and the audio before and after this point is divided into two local segments, where the voice quality difference within each segment is small; repeat this process until the entire audio is traversed, and multiple local audio segments with the same voice quality can be obtained.

[0103] S330: Mark the local audio segments whose voice quality is lower than a preset quality threshold as unqualified, and perform noise analysis on the marked unqualified audio segments to obtain noise characteristic parameters of the unqualified audio segments.

[0104] The unqualified audio segments refer to those local segments whose speech quality scores are lower than a preset threshold. In this embodiment, all segments are first screened according to the quality threshold, and those segments with unqualified quality are marked, and then noise characteristics are analyzed for them.

[0105] Specifically, the quality scores of all local audio segments can be traversed and compared with the preset threshold. If the score is lower than the threshold, it is marked as an unqualified segment. For these unqualified segments, it is necessary to extract the noise components and analyze the characteristic parameters such as the frequency range and energy distribution of the noise. The spectral characteristics of the noise can be detected by wavelet transform, spectral flatness test and other methods to obtain the noise characteristic parameter set of the segment.

[0106] S340: Analyze the noise suppression intensity corresponding to the unqualified audio segment and the directional suppression process corresponding to each frequency band according to the noise characteristic parameters to obtain the directional suppression parameters of the unqualified audio segment.

[0107] Among them, the noise characteristic parameters are used to tailor the best noise suppression strategy for each unqualified audio segment, namely, the directional suppression parameters.

[0108] Specifically, firstly, based on the frequency range and energy distribution in the noise characteristic parameters, the required noise suppression intensity is estimated, that is, how many decibels of noise energy is expected to be suppressed in this frequency band. Secondly, combined with actual usage scenarios such as temperature, environment and other factors, what type of suppression process should be performed in each frequency band is analyzed.

[0109] In one embodiment, referring to Figure 4 In step S400, the noise suppression strength of the local audio segment is adjusted according to the directional suppression parameter, and the suppression process parameter of the local audio segment is generated and sent to the noise suppression module according to the adjusted noise suppression strength, which specifically includes the following steps:

[0110] S410: Obtain a noise frequency range and corresponding noise energy of the current noise according to the directional suppression parameter, and analyze the current noise intensity of the current noise according to the noise frequency range and the noise energy.

[0111] The noise intensity refers to the total energy value of the current noise in its main frequency range. This embodiment first extracts the frequency range of the noise and the noise energy value of the corresponding frequency band from the directional suppression parameter, and then sums these energy values ​​to obtain the noise intensity estimation value of the current noise.

[0112] Specifically, the noise frequency bands recorded in the directional suppression parameters are traversed, and the noise energy value in each frequency band is calculated using methods such as short-time Fourier transform. Finally, the energy values ​​of all frequency bands are accumulated to obtain the total noise intensity.

[0113] S420: Calculate a deviation between the directional suppression parameter and the current noise intensity, and adjust the noise suppression intensity of the local audio segment according to the deviation.

[0114] In this embodiment, the difference between the current noise condition and the expected suppression strategy is evaluated, and the noise suppression strength of the local audio segment is adjusted accordingly.

[0115] Specifically, first take the expected noise suppression intensity value from the directional suppression parameter, then compare it with the current noise intensity value, and calculate the deviation between the two. If the deviation value is large, it means that the current noise condition is significantly different from the expected one, and the suppression intensity needs to be adjusted. The original suppression intensity value can be added or subtracted according to the positive or negative deviation value. The corrected suppression intensity value will guide the subsequent noise suppression process.

[0116] S430: Acquire the current speech feature of the local audio segment, analyze the suppression effect of each suppression frequency band on the current speech feature layer by layer, and obtain a suppression frequency sequence.

[0117] Specifically, the speech spectrum features of the local audio segment are extracted as the representation of the current speech features. Then all possible suppression frequency bands are traversed, each frequency band is virtually suppressed, the speech fidelity and noise suppression effect under the current speech features are evaluated, and a comprehensive score is given. According to these scores, all frequency bands are sorted from high to low to obtain a suppression frequency sequence.

[0118] S440: According to the noise suppression intensity and the suppression frequency sequence, the original noise suppression process of the local audio segment is adjusted in parameters to generate suppression process parameters corresponding to the local audio segment.

[0119] Specifically, according to the adjusted noise suppression strength value, the corresponding energy attenuation is set for the suppression process of each frequency band. Then, according to the suppression frequency sequence, from high priority to low, the suppression operator type such as notch filter, adaptive filter, etc. and the specific parameter value are assigned to each frequency band in turn. For different frequency bands, it is also necessary to estimate the phase delay and compensate the parameters accordingly to avoid phase distortion. Finally, all process parameters are packaged to generate the final suppression process parameter set, which is output to the noise suppression module for execution.

[0120] In one embodiment, referring to Figure 5 In step S420, the deviation between the directional suppression parameter and the current noise intensity is calculated, and the noise suppression intensity of the local audio segment is adjusted according to the deviation value, which specifically includes the following steps:

[0121] S421. Estimate, based on the deviation value between the directional suppression parameter and the current noise intensity, the degree of speech quality loss caused by the current noise intensity to the local audio segment without noise suppression processing.

[0122] The degree of speech quality loss refers to the degree to which noise reduces speech clarity. This embodiment estimates, based on the deviation between the directional suppression parameter and the actual noise intensity, to what extent the speech quality of the local audio will be reduced due to the current noise intensity if no suppression is performed.

[0123] Specifically, a mapping model between noise intensity and speech quality loss is established, and the current noise intensity deviation value is substituted into the mapping model to estimate the corresponding speech quality loss value. The mapping model can be obtained through training of a large amount of speech noise data, and contains the influence of different noise intensities on speech clarity.

[0124] S422: Obtain an actual speech attenuation degree of the local audio segment under current noise energy.

[0125] Specifically, the actual degree of clarity reduction of the local speech when the current noise exists is measured, including extracting the pure speech component of the audio segment and calculating the degree of distortion between the speech and the ideal clear speech under the current noise energy, that is, the actual speech attenuation degree.

[0126] S423: Compare the voice quality loss degree with the actual voice attenuation degree, and calculate the difference between the voice quality loss degree and the actual voice attenuation degree.

[0127] Specifically, the theoretical speech quality loss degree estimated in the above steps is compared with the actual speech attenuation degree measured, and the difference between the two is calculated. A large positive difference indicates that the actual attenuation degree is higher than the theoretical value, and the suppression needs to be enhanced; a large negative difference indicates that the actual attenuation degree is lower than the theoretical value, and the suppression strength can be appropriately reduced.

[0128] S424: Determine whether the current noise suppression intensity is reasonable according to the difference, and predict the noise suppression effect of the local audio segment.

[0129] Specifically, a difference interval threshold is pre-set. If the difference falls within the interval, the current suppression strength is considered reasonable. If the difference is too large or too small and exceeds the interval, the suppression strength is considered unreasonable. For reasonable situations, the degree of noise suppression and speech fidelity can be predicted after suppression is implemented at this strength. For unreasonable situations, the suppression strength needs to be further adjusted.

[0130] S425. According to the noise suppression effect, adjust the noise suppression intensity of each suppressed frequency band of the local audio segment.

[0131] Specifically, the suppression strength value needs to be increased or decreased according to the positive or negative difference; for frequency bands with reasonable suppression strength, the original strength is maintained unchanged. By fine-tuning the suppression strength of each frequency band, the final noise suppression strategy for the local audio segment is comprehensively constructed.

[0132] In one embodiment, referring to Figure 6 In step S440, according to the noise suppression intensity and the suppression frequency sequence, the original noise suppression process of the local audio segment is adjusted to generate the suppression process parameters corresponding to the local audio segment, which specifically includes the following steps:

[0133] S441. Obtain the phase distortion of each suppression frequency band under the current ambient temperature, and calculate the superposition delay time of adjacent suppression frequency bands according to the phase distortion and the corresponding noise suppression strength.

[0134] In the actual noise suppression process, filtering operations in different frequency bands will introduce certain phase distortion and delay. Therefore, it is necessary to calculate the phase distortion of each suppressed frequency band under the current ambient temperature, and estimate the delay superposition between adjacent frequency bands based on this.

[0135] Specifically, a lookup table is pre-established to store the phase distortion modes of various filters at different temperatures and frequencies; then the phase distortion value of each frequency band is obtained according to the current temperature lookup table; the phase distortion values ​​of adjacent frequency bands are respectively substituted into a delay model to calculate the delay superposition amount.

[0136] S442: Calculate the superimposed interference ratio of adjacent suppressed frequency bands according to the superimposed delay time and the corresponding suppressed frequency sequence.

[0137] Specifically, due to the superposition of delays between frequency bands, frequency interference will occur in the suppression process. Therefore, the priority suppression order of the frequency bands is determined based on the obtained suppression frequency sequence; then all adjacent frequency band pairs are traversed, and the corresponding interference ratio is calculated based on the superposition delay time between them and an interference model.

[0138] S443. Adjust the single-frequency suppression bandwidth corresponding to the adjacent suppression frequency band according to the superimposed interference ratio, and calculate the suppression balance coefficient for achieving frequency balance in the adjacent suppression frequency band according to the adjusted suppression bandwidth.

[0139] Specifically, first traverse all adjacent frequency band pairs, and according to the superimposed interference ratio between them, reduce the single-frequency suppression bandwidth of one of the frequency bands according to certain rules such as linear rules; then traverse all frequency bands again, calculate the frequency imbalance between each frequency band and other frequency bands under the adjusted bandwidth distribution, and obtain a balance coefficient sequence. The balance coefficient sequence describes how much energy each frequency band needs to increase or decrease relatively to achieve overall frequency balance under the current bandwidth distribution.

[0140] S444: According to the suppression balance coefficient, the original noise suppression process of the local audio segment is adjusted in parameters to generate suppression process parameters corresponding to the local audio segment.

[0141] Specifically, traverse each suppression frequency band of the local audio band, and modify the original suppression filter parameters according to the corresponding suppression balance coefficient, such as adjusting the passband / stopband range of the notch filter, adjusting the step size factor of the adaptive filter, etc. For frequency bands that need to be suppressed, the stopband attenuation of the filter can be increased; and for frequency bands that need to be suppressed, the stopband attenuation can be appropriately reduced. At the same time, consider maintaining phase continuity between different frequency bands. Finally, all the corrected filter parameters are packaged to generate a suppression process parameter set, which is output to the noise suppression module for use.

[0142] In one embodiment, referring to Figure 7 , the method further comprises:

[0143] S500: Obtain noise energy differences between adjacent audio segments, and obtain speech amplitude difference values ​​and noise amplitude values ​​corresponding to speech component proportions according to the noise energy differences.

[0144] In this embodiment, attention is paid to the change of noise energy between adjacent audio segments, because such change will directly lead to a change in the proportion of speech components, thereby affecting the amplitude of speech and noise.

[0145] Specifically, the audio is first divided into multiple sub-bands using time-frequency analysis methods such as wavelet decomposition; then the total noise energy is calculated in each sub-band, and the noise energy difference between adjacent sub-bands is obtained. Then, the lookup table pre-trained with a large amount of speech noise data is consulted, and the corresponding change in the proportion of speech components is found according to the current noise energy difference. Finally, the speech amplitude and noise amplitude values ​​are combined to estimate the speech amplitude difference and noise amplitude value.

[0146] S600: construct a speech change curve corresponding to the noise energy difference according to the speech amplitude difference and the noise amplitude value.

[0147] Specifically, the speech amplitude difference and the noise amplitude value are respectively used as the independent variable and the dependent variable, and the functional relationship between them is estimated by a curve fitting algorithm such as polynomial fitting to obtain an initial speech change curve.

[0148] S700: Analyze the attenuation trend of the speech according to the speech change curve to obtain the speech attenuation value of the current noise environment.

[0149] In this embodiment, the speech attenuation value is an important indicator for evaluating the influence of the current noise environment on speech.

[0150] Specifically, we first perform mathematical analysis on the function expression of the speech change curve to solve its limit value when the independent variable approaches infinity, that is, the stable convergence value of the speech amplitude, i.e., the speech attenuation value. Then, other characteristic parameters of the noise environment, such as noise spectrum distribution and noise sound pressure level, are introduced into the correction model to correct the preliminary value and obtain the final speech attenuation value estimate.

[0151] S800: When the actual amplitude of the speech reaches the speech attenuation value, output gain compensation information to the noise suppression module.

[0152] In this embodiment, when the actual amplitude of the speech signal drops to a certain level, gain compensation needs to be performed on it to prevent it from further attenuating and affecting the clarity.

[0153] Specifically, the short-time energy of the speech signal is detected in real time and compared with the speech attenuation value. Once it is detected that the short-time energy of the speech is less than or equal to the attenuation value, the gain compensation parameters, including the gain coefficient, the gain frequency band, etc., are immediately generated, and these parameters are sent to the adaptive gain control submodule of the noise suppression module. After receiving the compensation parameters, the submodule will apply the corresponding gain to the specific frequency band of the speech signal to increase its volume and avoid further attenuation of the speech. Among them, the calculation of the compensation parameters can use a pre-trained compensation model, which summarizes the mapping relationship between speech amplitude, noise characteristics and gain adjustment.

[0154] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0155] In a second aspect, the present application provides a noise suppression system for a microphone. The noise suppression system for a microphone of the present application is described below in conjunction with the above-mentioned noise suppression method for a microphone.

[0156] Reference Figure 8 , a noise suppression system for a microphone, comprising:

[0157] A data acquisition module, used to acquire the background noise energy of the microphone input audio signal and the frequency spectrum characteristics of the speech signal;

[0158] The quality analysis module is used to analyze the quality of the speech signal according to the background noise energy and spectrum characteristics to obtain the speech quality analysis result;

[0159] The suppression analysis module is used to process the microphone input audio signal in segments according to the voice quality analysis result, and perform directional noise suppression analysis on the local audio segment with unqualified voice quality to obtain the corresponding directional suppression parameters;

[0160] The noise suppression module is used to adjust the noise suppression strength of the local audio segment according to the directional suppression parameter, generate and send the suppression process parameters of the local audio segment to the noise suppression module according to the adjusted noise suppression strength, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameters.

[0161] In an optional embodiment, the quality analysis module includes:

[0162] A signal-to-noise ratio calculation unit, used to calculate the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and spectrum characteristics;

[0163] A voice component ratio calculation unit, used to calculate the voice component ratio corresponding to each time-frequency point according to the signal-to-noise ratio;

[0164] A single-point speech distortion value analysis unit is used to analyze the degree of distortion between the current speech and the preset clear speech according to the speech component ratio, and obtain the single-point speech distortion value of the current time-frequency point;

[0165] The quality analysis result acquisition unit is used to estimate the comprehensive speech distortion value of the entire speech signal according to the single-point speech distortion value, perform quality analysis on the speech signal according to the comprehensive speech distortion value, and obtain the quality analysis result of the speech signal.

[0166] In an optional embodiment, the inhibition analysis module includes:

[0167] A voice quality difference obtaining unit, used for obtaining the voice quality difference between adjacent audio segments according to the voice quality analysis result;

[0168] A local audio segment generating unit, used for segmenting the microphone input audio signal according to the difference in voice quality to generate local audio segments with the same voice quality;

[0169] A noise characteristic parameter acquisition unit is used to mark the local audio segment whose voice quality is lower than a preset quality threshold as unqualified, and perform noise analysis on the marked unqualified audio segment to obtain the noise characteristic parameters of the unqualified audio segment;

[0170] The directional suppression parameter acquisition unit is used to analyze the noise suppression intensity corresponding to the unqualified audio segment and the directional suppression process corresponding to each frequency band according to the noise characteristic parameters, and obtain the directional suppression parameters of the unqualified audio segment.

[0171] In an optional embodiment, the noise suppression module includes:

[0172] A current noise intensity acquisition unit, used to acquire a noise frequency range and a corresponding noise energy of the current noise according to the directional suppression parameter, and analyze the current noise intensity of the current noise according to the noise frequency range and the noise energy;

[0173] A noise suppression intensity acquisition unit, used to calculate a deviation value between the directional suppression parameter and the current noise intensity, and adjust the noise suppression intensity of the local audio segment according to the deviation value;

[0174] A suppression frequency sequence acquisition unit is used to acquire the current speech feature of the local audio segment, analyze the suppression effect of each suppression frequency band on the current speech feature layer by layer, and obtain a suppression frequency sequence;

[0175] The suppression process parameter generating unit is used to adjust the parameters of the original noise suppression process of the local audio segment according to the noise suppression intensity and the suppression frequency sequence, and generate the suppression process parameters corresponding to the local audio segment.

[0176] In an optional embodiment, the noise suppression strength acquisition unit includes:

[0177] A speech quality loss degree estimation subunit, used for estimating the speech quality loss degree caused by the current noise intensity to the local audio segment without noise suppression processing according to the deviation value between the directional suppression parameter and the current noise intensity;

[0178] The actual speech attenuation degree acquisition subunit is used to acquire the actual speech attenuation degree of the local audio segment under the current noise energy;

[0179] A difference calculation subunit, used for comparing the voice quality loss degree with the actual voice attenuation degree, and calculating the difference between the voice quality loss degree and the actual voice attenuation degree;

[0180] A noise suppression effect prediction subunit, used to determine whether the current noise suppression strength is reasonable according to the difference, and to predict the noise suppression effect of the local audio segment;

[0181] The noise suppression intensity adjustment subunit is used to adjust the noise suppression intensity of each suppression frequency band of the local audio segment according to the noise suppression effect.

[0182] In an optional embodiment, the suppression process parameter generating unit includes:

[0183] The superposition delay time calculation subunit is used to obtain the phase distortion of each suppression frequency band under the current ambient temperature, and calculate the superposition delay time of adjacent suppression frequency bands according to the phase distortion and the corresponding noise suppression strength;

[0184] A superposition interference ratio calculation subunit, used to calculate the superposition interference ratio of adjacent suppression frequency bands according to the superposition delay time and the corresponding suppression frequency sequence;

[0185] The suppression balance coefficient calculation subunit is used to adjust the single-frequency suppression bandwidth corresponding to the adjacent suppression frequency band according to the superimposed interference ratio, and calculate the suppression balance coefficient for the adjacent suppression frequency band to achieve frequency balance according to the adjusted suppression bandwidth;

[0186] The suppression process parameter generation subunit is used to adjust the parameters of the original noise suppression process of the local audio segment according to the suppression balance coefficient, and generate the suppression process parameters corresponding to the local audio segment.

[0187] In an optional embodiment, the system further includes:

[0188] A noise energy difference acquisition module is used to obtain the noise energy difference between adjacent audio segments, and obtain the speech amplitude difference and the noise amplitude value corresponding to the speech component proportion according to the noise energy difference;

[0189] A speech change curve construction module is used to construct a speech change curve corresponding to the noise energy difference according to the speech amplitude difference and the noise amplitude value;

[0190] The speech attenuation value determination module is used to analyze the attenuation trend of the speech according to the speech change curve to obtain the speech attenuation value of the current noise environment;

[0191] The gain compensation information output module is used to output gain compensation information to the noise suppression module when the actual amplitude of the speech reaches the speech attenuation value.

[0192] In one embodiment, the present application provides a microphone, the internal structure of which can be as shown in FIG. Fig. 9 As shown. The microphone includes a processor, a memory and a network interface connected through a system bus. The processor of the microphone is used to provide computing and control capabilities. The memory of the microphone includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the microphone is used to store data. The network interface of the microphone is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a noise suppression method for a microphone is implemented.

[0193] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the microphone to which the scheme of the present application is applied. The specific microphone may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0194] In one embodiment, a microphone is further provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0195] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0196] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.

Claims

1. A noise suppression method for a microphone, characterized in that: The steps include: Obtaining background noise energy of microphone input audio signal and frequency spectrum characteristics of speech signal; Calculating the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and the frequency spectrum characteristics; Calculate the speech component ratio corresponding to each time-frequency point according to the signal-to-noise ratio; Analyzing the distortion degree between the current speech and the preset clear speech according to the speech component proportion, and obtaining the single-point speech distortion value of the current time-frequency point; estimating a comprehensive speech distortion value of the entire speech signal according to the single-point speech distortion value, performing quality analysis on the speech signal according to the comprehensive speech distortion value, and obtaining a quality analysis result of the speech signal; According to the quality analysis result of the speech signal, the microphone input audio signal is processed in segments, and a directional noise suppression analysis is performed on the local audio segment with unqualified speech quality to obtain corresponding directional suppression parameters; Acquire a noise frequency range and a corresponding noise energy of the current noise according to the directional suppression parameter, and analyze a current noise intensity of the current noise according to the noise frequency range and the noise energy; Calculating a deviation value between the directional suppression parameter and the current noise intensity, and adjusting the noise suppression intensity of the local audio segment according to the deviation value; Acquire the current speech feature of the local audio segment, analyze the suppression effect of each suppression frequency band on the current speech feature layer by layer, and obtain a suppression frequency sequence; Obtaining the phase distortion of each suppressed frequency band under the current ambient temperature, and calculating the superposition delay time of adjacent suppressed frequency bands according to the phase distortion and the corresponding noise suppression strength; Calculating the superposition interference ratio of the adjacent suppression frequency bands according to the superposition delay time and the corresponding suppression frequency sequence; adjusting the single-frequency suppression bandwidth corresponding to the adjacent suppression frequency band according to the superimposed interference ratio, and calculating the suppression balance coefficient for achieving frequency balance of the adjacent suppression frequency band according to the adjusted suppression bandwidth; According to the suppression balance coefficient, the parameters of the original noise suppression process of the local audio segment are adjusted, and the suppression process parameters corresponding to the local audio segment are generated and sent to the noise suppression module, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameters.

2. The noise suppression method for a microphone according to claim 1, characterized in that: The method further comprises: Acquire the noise energy difference between adjacent audio segments, and acquire the speech amplitude difference and the noise amplitude value corresponding to the speech component proportion according to the noise energy difference; Constructing a speech change curve corresponding to the noise energy difference according to the speech amplitude difference and the noise amplitude value; Analyzing the attenuation trend of the speech according to the speech change curve to obtain the speech attenuation value of the current noise environment; When the actual amplitude of the speech reaches the speech attenuation value, the gain compensation information is output to the noise suppression module.

3. The noise suppression method for a microphone according to claim 1, characterized in that: According to the speech quality analysis result, the microphone input audio signal is processed in segments, and a directional noise suppression analysis is performed on the local audio segment with unqualified speech quality to obtain the corresponding directional suppression parameter, which specifically includes the following steps: According to the quality analysis result of the speech signal, obtaining the speech quality difference between adjacent audio segments; Segmentally process the microphone input audio signal according to the voice quality difference to generate local audio segments with the same voice quality; Marking a local audio segment whose voice quality is lower than a preset quality threshold as unqualified, and performing noise analysis on the marked unqualified audio segment to obtain noise characteristic parameters of the unqualified audio segment; According to the noise characteristic parameters, the noise suppression intensity corresponding to the unqualified audio segment and the directional suppression process corresponding to each frequency band are analyzed to obtain the directional suppression parameters of the unqualified audio segment.

4. The noise suppression method for a microphone according to claim 1, characterized in that: Calculating a deviation value between the directional suppression parameter and the current noise intensity, and adjusting the noise suppression intensity of the local audio segment according to the deviation value, specifically comprises the following steps: estimating, according to a deviation value between the directional suppression parameter and the current noise intensity, a degree of speech quality loss caused by the current noise intensity to the local audio segment without noise suppression processing; Obtaining an actual speech attenuation degree of the local audio segment under current noise energy; Comparing the voice quality loss degree with the actual voice attenuation degree, and calculating the difference between the voice quality loss degree and the actual voice attenuation degree; Determining whether the current noise suppression strength is reasonable according to the difference, and predicting the noise suppression effect of the local audio segment; According to the noise suppression effect, the noise suppression intensity of each suppressed frequency band of the local audio segment is adjusted.

5. A noise suppression system for a microphone, characterized in that include: A data acquisition module, used to acquire the background noise energy of the microphone input audio signal and the frequency spectrum characteristics of the speech signal; A quality analysis module, used to analyze the quality of the speech signal according to the background noise energy and the spectrum characteristics to obtain a speech quality analysis result; A suppression analysis module, configured to perform segmented processing on the microphone input audio signal according to the voice quality analysis result, perform directional noise suppression analysis on the local audio segment with unqualified voice quality, and obtain corresponding directional suppression parameters; a noise suppression module, configured to adjust the noise suppression strength of the local audio segment according to the directional suppression parameter, and generate and send a suppression process parameter of the local audio segment to the noise suppression module according to the adjusted noise suppression strength, so that the noise suppression module adjusts the noise suppression process of the local audio segment according to the suppression process parameter; The quality analysis module includes: A signal-to-noise ratio calculation unit, used to calculate the signal-to-noise ratio of each frequency band of the speech signal according to the background noise energy and the frequency spectrum characteristics; A speech component ratio calculation unit, used to calculate the speech component ratio corresponding to each time-frequency point according to the signal-to-noise ratio; A single-point speech distortion value analysis unit, used to analyze the degree of distortion between the current speech and the preset clear speech according to the speech component proportion, and obtain the single-point speech distortion value of the current time-frequency point; A quality analysis result acquisition unit, used to estimate the comprehensive speech distortion value of the entire speech signal according to the single-point speech distortion value, perform quality analysis on the speech signal according to the comprehensive speech distortion value, and obtain a quality analysis result of the speech signal; The noise suppression module comprises: a current noise intensity acquisition unit, configured to acquire a noise frequency range and a corresponding noise energy of the current noise according to the directional suppression parameter, and analyze a current noise intensity of the current noise according to the noise frequency range and the noise energy; a noise suppression intensity acquisition unit, configured to calculate a deviation value between the directional suppression parameter and the current noise intensity, and adjust the noise suppression intensity of the local audio segment according to the deviation value; A suppression frequency sequence acquisition unit, used to acquire the current speech feature of the local audio segment, and analyze the suppression effect of each suppression frequency segment on the current speech feature layer by layer to obtain a suppression frequency sequence; a suppression process parameter generating unit, configured to adjust parameters of an original noise suppression process of a local audio segment according to the noise suppression intensity and the suppression frequency sequence, and generate a suppression process parameter corresponding to the local audio segment; The suppression process parameter generating unit comprises: The superposition delay time calculation subunit is used to obtain the phase distortion of each suppression frequency band under the current ambient temperature, and calculate the superposition delay time of adjacent suppression frequency bands according to the phase distortion and the corresponding noise suppression strength; A superposition interference ratio calculation subunit, used to calculate the superposition interference ratio of adjacent suppression frequency bands according to the superposition delay time and the corresponding suppression frequency sequence; The suppression balance coefficient calculation subunit is used to adjust the single-frequency suppression bandwidth corresponding to the adjacent suppression frequency band according to the superimposed interference ratio, and calculate the suppression balance coefficient for the adjacent suppression frequency band to achieve frequency balance according to the adjusted suppression bandwidth; The suppression process parameter generation subunit is used to adjust the parameters of the original noise suppression process of the local audio segment according to the suppression balance coefficient, and generate the suppression process parameters corresponding to the local audio segment.

6. A microphone, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the noise suppression method for a microphone described in any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps for noise suppression of a microphone according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Noise elimination device and detection method thereof

    CN113889134A

  • Noise suppression method and device, equipment and storage medium

    CN118280381A

  • Noise suppression device and noise suppressing method

    US20180033448A1