Microphone audio noise reduction method, device and equipment and readable storage medium
Patent Information
- Application Number
- CN202610985200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-07-03
AI Technical Summary
[0005]本发明的目的在于提出了一种麦克风音频降噪方法、装置、设备及可读存储介质,旨在解决现有技术中高速骑行场景下音频降噪效果差的问题
[0010]The embodiments of the present invention have the following beneficial effects: Unlike existing technologies, this application acquires an audio signal to be processed; performs wind noise detection on the audio signal to be processed to determine its wind noise level; matches adaptive noise reduction parameters based on the wind noise level, and performs noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; uses a target deep neural network model to extract the spectral features of the audio signal to be processed, and generates a second noise-reduced signal based on the spectral features and the wind noise level; fuses the first and second noise-reduced signals to obtain a noise-reduced audio signal. By using both adaptive noise reduction and neural network model-based dual-path noise reduction based on the wind noise level of the audio signal, a stable and high-fidelity noise reduction effect is ensured for various wind noise scenarios under extreme wind noise conditions during high-speed cycling.
Smart Images

Figure CN122511278B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human audio noise reduction technology, and in particular to a microphone audio noise reduction method, apparatus, device, and readable storage medium. Background Technology
[0002] In high-speed riding scenarios (speeds exceeding 120 km / h), riders typically wear helmet headsets or walkie-talkies with an integrated single microphone. This microphone is continuously exposed to a high-intensity, non-stationary, and complex acoustic environment. Major noise sources include: high-intensity wind noise generated by turbulent airflow impacting the microphone, tire noise from tire-road friction, mechanical noise from frame vibration, engine noise, and traffic noise from surrounding vehicles. Among these, wind noise, due to its wide bandwidth and non-stationary characteristics, has the most severe impact on voice communication quality. At high speeds, wind noise energy is sufficient to completely drown out normal voice signals; at extreme motorcycle speeds (e.g., above 100 km / h), wind noise pressure levels can reach over 120 dB SPL, rendering walkie-talkie communication essentially ineffective.
[0003] The formation mechanism and acoustic characteristics of wind noise: When high-speed airflow impacts the microphone and surrounding structures, it generates turbulent boundary layer separation, forming random pressure pulses. In the time domain, wind noise manifests as a sudden high-amplitude pulse sequence, and in the frequency domain, it presents a broadband spectrum extending from low frequencies to mid-high frequencies. Wind noise energy is mainly concentrated in the low-frequency band (below 200Hz), but in high-speed scenarios, it can extend upwards to above 2kHz, directly intruding into the main energy frequency band of speech.
[0004] The core challenge of single-microphone noise reduction lies in the fact that separating the target speech from a single-channel mixed signal cannot suppress noise through spatial filtering like multi-microphone solutions; it must rely on signal processing algorithms to model time-frequency features. In high-speed cycling scenarios, the challenges are further amplified: wind noise is extremely intense, the signal-to-noise ratio is extremely low, noise characteristics dynamically change with speed, and intercom devices are typically resource-constrained embedded platforms, requiring algorithms to perform low-power, low-latency real-time processing. Summary of the Invention
[0005] The purpose of this invention is to provide a microphone audio noise reduction method, apparatus, device, and readable storage medium, aiming to solve the problem of poor audio noise reduction effect in high-speed cycling scenarios in the prior art.
[0006] To address the aforementioned technical problems, this application provides a microphone audio noise reduction method, the method comprising: acquiring an audio signal to be processed; performing wind noise detection on the audio signal to be processed to determine the wind noise level of the audio signal to be processed; matching adaptive noise reduction parameters according to the wind noise level, and performing noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; extracting spectral features of the audio signal to be processed using a target deep neural network model, and generating a second noise-reduced signal according to the spectral features and the wind noise level; and fusing the first noise-reduced signal and the second noise-reduced signal to obtain a noise-reduced audio signal.
[0007] To address the aforementioned technical problems, a second aspect of this application provides a microphone audio noise reduction device. The device includes: an acquisition module for acquiring an audio signal to be processed; a noise grading module for detecting wind noise in the audio signal to be processed and determining the wind noise level of the audio signal; a noise reduction module for matching adaptive noise reduction parameters according to the wind noise level, and performing noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; extracting spectral features of the audio signal to be processed using a target deep neural network model, and generating a second noise-reduced signal based on the spectral features and the wind noise level; and fusing the first noise-reduced signal and the second noise-reduced signal to obtain a noise-reduced audio signal.
[0008] To address the aforementioned technical problems, a third aspect of this application provides a computer device including a processor and a memory coupled to each other; the memory stores a computer program, and the processor executes the computer program to implement the steps of the method provided in the first aspect above.
[0009] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium storing program data, which, when executed by a processor, implements the steps of the method provided in the first aspect above.
[0010] The embodiments of the present invention have the following beneficial effects: Unlike existing technologies, this application acquires an audio signal to be processed; performs wind noise detection on the audio signal to be processed to determine its wind noise level; matches adaptive noise reduction parameters based on the wind noise level, and performs noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; uses a target deep neural network model to extract the spectral features of the audio signal to be processed, and generates a second noise-reduced signal based on the spectral features and the wind noise level; fuses the first and second noise-reduced signals to obtain a noise-reduced audio signal. By using both adaptive noise reduction and neural network model-based dual-path noise reduction based on the wind noise level of the audio signal, a stable and high-fidelity noise reduction effect is ensured for various wind noise scenarios under extreme wind noise conditions during high-speed cycling. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] in: Figure 1 This is a schematic flowchart of an embodiment of the microphone audio noise reduction method of this application; Figure 2 This is a schematic block diagram of an embodiment of the microphone audio noise reduction device of this application; Figure 3 This is a schematic flowchart of another embodiment of the microphone audio noise reduction method of this application; Figure 4 This is a schematic flowchart of another embodiment of the microphone audio noise reduction method of this application; Figure 5 This is a schematic block diagram of the structure of an embodiment of the computer device of this application; Figure 6 This is a schematic block diagram of an embodiment of a computer-readable storage medium of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Please see Figure 1, Figure 1 This is a schematic flowchart of an embodiment of the microphone audio noise reduction method of this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it with a similar method. Figure 1 The illustrated process sequence is limited. It includes the following steps S11~S15: S11: Acquire the audio signal to be processed.
[0015] The audio signal to be processed can be the original noisy voice signal of a cyclist during high-speed riding (speed can reach 120km / h~195km / h) picked up by a single microphone. The signal contains the cyclist's voice components as well as environmental noise components such as wind noise, tire noise, and mechanical noise, especially high-intensity, non-steady turbulent wind noise.
[0016] S12: Perform wind noise detection on the audio signal to be processed to determine the wind noise level of the audio signal to be processed.
[0017] Wind noise detection of the audio signal to be processed refers to dividing the current signal segment into different wind noise intensity levels, such as five levels from N0 (weak wind noise) to N4 (extremely strong wind noise), based on the acoustic characteristics of the audio signal itself, using a lightweight classifier. Each level corresponds to a different signal-to-noise ratio range.
[0018] The audio performance of the five wind noise levels N0 to N4 is as follows: N0 level (weak wind noise): No obvious wind noise, mainly stable environmental noise, with a high signal-to-noise ratio (>15dB). Level N1 (Medium-weak wind noise): Wind noise begins to appear, low-frequency energy is slightly increased, and the signal-to-noise ratio is moderate (10~15dB). N2 level (medium wind noise): significant wind noise, high proportion of low frequency energy, rapid autocorrelation decay, and low signal-to-noise ratio (5~10dB). N3 level (strong wind noise): high-intensity wind noise, interference across the entire frequency band, low signal-to-noise ratio (0~5dB). Level N4 (Extremely Strong Wind Noise): Extreme wind noise, speech is almost completely drowned out, signal-to-noise ratio is less than 0dB.
[0019] S13: Based on the wind noise level, match adaptive noise reduction parameters, and perform noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal.
[0020] Specifically, based on the wind noise level obtained from the classification, noise reduction control parameters matching the level can be selected from a preset parameter set or parameter mapping rules. These parameters include, but are not limited to, the time constant for noise spectrum estimation, the intensity of the over-subtraction factor, the fusion weight coefficient, and the gain adjustment range. Then, a traditional signal processing path (such as the Optimal Modified Logarithmic Spectral Amplitude Estimation (OMLSA) algorithm) is adopted, and the aforementioned adaptive noise reduction parameters are used to conditionally control the algorithm, generating a first-path noise-reduced signal with low computational complexity and good stability. In one specific embodiment, the adaptive noise reduction parameter is the time constant for noise spectrum estimation. When the wind noise level is low, a longer smoothing time constant is used to ensure estimation stability; when the wind noise level is high, a shorter time constant is used to quickly track wind noise changes.
[0021] S14: Using the target deep neural network model, extract the spectral features of the audio signal to be processed, and generate a second noise reduction signal based on the spectral features and the wind noise level.
[0022] A lightweight deep neural network model is used to extract time-spectral features from the audio signal to be processed. At the same time, the wind noise level is injected into the model as conditional information (such as by one-hot encoding and splicing with the spectral features), so that the model can dynamically adjust the noise reduction strategy according to the current wind noise intensity and generate a second noise reduction signal.
[0023] S15: The first noise-reduced signal and the second noise-reduced signal are fused together to obtain the noise-reduced audio signal.
[0024] Specifically, the fusion weights of the two noise-reduced signals can be dynamically determined according to the wind noise level. The two signals can be merged by weighted summation or other fusion strategies to achieve the complementary advantages of traditional adaptive noise reduction methods and neural network model noise reduction methods. That is, the advantages of human voice fidelity in neural network model noise reduction are brought into play when the wind noise is weak, while the robustness of traditional adaptive noise reduction methods is the main factor when the wind noise is strong.
[0025] This embodiment performs dual-path noise reduction based on the wind noise level of the audio signal. According to the wind noise level, the parameter configuration of the adaptive noise reduction path and the conditional injection of the neural network model noise reduction path are controlled respectively. The two noise reduction signals are then fused by matching the levels. This can achieve a stable and high-fidelity noise reduction effect in all scenarios from weak wind noise to extremely strong wind noise in high-speed cycling extreme wind noise environment. It avoids the problem of insufficient noise suppression under strong wind noise when using only adaptive noise reduction method, and overcomes the defects of poor generalization ability and possible uncontrollable distortion when using only neural network model under unknown extreme noise. It achieves noise reduction effect and voice fidelity in high-speed scenarios.
[0026] In one embodiment, motion state information or speed information corresponding to the audio signal to be processed is obtained, the wind noise level for the current or future period is predicted based on the motion state information or speed information, and the predicted wind noise level is used to replace the real-time detected wind noise level.
[0027] It should be noted that motion status information or speed information can be obtained in various ways. In cycling scenarios, the current cycling speed can be obtained using walkie-talkies or accelerometers, gyroscopes, and GPS modules built into helmets; real-time speed can be obtained through communication with the vehicle's CAN bus; or speed can be indirectly estimated using the Doppler effect in the audio signal collected by the microphone or engine speed harmonics. Wind noise level is strongly correlated with cycling speed; the higher the speed, the greater the wind noise intensity. Based on speed information, a speed-wind noise level mapping model can be established. For example, the correspondence between speed ranges and wind noise levels can be pre-calibrated using a large amount of experimental data: when the speed is below 30 km / h, it usually corresponds to level N0 weak wind noise; when the speed is between 30-60 km / h, it corresponds to levels N1 to N2; when the speed is between 60-100 km / h, it corresponds to levels N2 to N3; and when the speed exceeds 100 km / h, it corresponds to levels N3 to N4. The specific mapping can be implemented using piecewise linear interpolation or lookup tables. When real-time wind noise detection results are not yet stable or are lagging, speed information is used to predict the wind noise level in the current frame or a short future window, such as 200ms to 500ms. If a speed increase exceeding 20km / h is detected within 100ms, a higher wind noise level can be predicted in advance and used to replace the real-time detected level for subsequent adaptive noise reduction parameter matching, neural network condition injection, and fusion weight determination. This prediction can be weighted and fused with the real-time detection results, for example, with a prediction weight of 0.7 and a detection weight of 0.3, or it can be directly substituted. Speed prediction effectively compensates for the lag in response to wind noise level changes caused by computational delays or time window accumulation in pure audio wind noise detection methods.
[0028] This implementation incorporates physical sensor information such as cycling speed into the wind noise reduction system, enabling feedforward predictive control of wind noise levels. This allows the noise reduction system to adjust its strategy in advance when wind noise changes rapidly during high-speed cycling, significantly improving the system's response speed to rapidly changing wind noise. Furthermore, even in the event of a brief loss of GPS or sensor signals, inferences can still be made based on historical speed trends, enhancing the system's robustness.
[0029] In one embodiment, the degree of nonstationarity of wind noise in the audio signal to be processed is detected, and the update rate of the noise spectrum estimation is dynamically adjusted according to the degree of nonstationarity, wherein the higher the degree of nonstationarity, the faster the update rate.
[0030] Understandably, the degree of non-stationarity refers to the drastic change of wind noise over time, which can be quantified by calculating indicators such as the short-time energy variance of the audio signal, the volatility of the amplitude envelope, the decay rate of the spectral flux, or the autocorrelation function. Taking the short-time energy variance as an example, the energy of each frame of the audio signal is calculated frame by frame, with a frame length of 20-40ms and a frame shift of 10ms, and then the energy variance of 5-10 consecutive frames is calculated. The larger the variance, the more drastic the wind noise fluctuation. Wind noise under turbulent impact often manifests as a sudden high-amplitude pulse sequence, and its energy can fluctuate drastically in a very short time. By detecting the degree of non-stationarity in real time, the update rate of the noise spectrum estimation is dynamically adjusted. Specifically, a range of non-stationarity indicators can be set, such as normalizing the short-time energy variance to between 0 and 1. When the variance is less than 0.2, it is judged as stationary wind noise, and the recursive average coefficient α of the noise spectrum estimation is... s Setting it to 0.98 results in a slow update rate; when the variance is between 0.2 and 0.6, α... s Set it to 0.9; when the variance is greater than 0.6, α s Setting it to 0.7 provides a fast update rate. Slow updates during stable conditions ensure the stability of the noise spectrum estimation and prevent mistracking of speech components; rapid updates during drastic fluctuations allow the noise spectrum to reflect wind noise changes promptly, avoiding insufficient noise suppression. The above values can be fine-tuned based on the actual hardware platform and wind noise characteristics.
[0031] This implementation solves the problem of traditional noise spectrum estimation not tracking in a timely manner in scenarios with rapidly changing wind noise. It obtains a stable estimate when the wind noise is stable and tracks quickly when the wind noise fluctuates violently, achieving an automatic balance between the accuracy and response speed of noise spectrum estimation, and effectively improving the noise reduction quality of the first noise reduction signal in scenarios with strong wind noise and transient wind noise.
[0032] In one embodiment, the audio signal to be processed is divided into multiple frequency bands, wind noise levels are detected for different frequency bands, and different adaptive noise reduction parameters and different neural network processing strategies are matched for different frequency bands, wherein the wind noise level weight of the low frequency band is higher than that of the high frequency band.
[0033] In one possible implementation, frequency band division can be uniform or non-uniform based on auditory characteristics, such as dividing the frequency band into 8 or 16 bands according to the critical frequency band. For example, 0-8kHz can be divided into 5 bands: band 1 (0-200Hz), the main concentration area of wind noise; band 2 (200-500Hz); band 3 (500-1500Hz); band 4 (1500-4000Hz); and band 5 (4000-8000Hz), the speech harmonic region. Since wind noise energy is mainly concentrated in the low-frequency band, but can extend to the mid-high frequencies during high-speed riding, the wind noise intensity of different bands often varies. Wind noise level detection can be performed independently for each band. The wind noise level of each band can be estimated based on characteristics such as the energy proportion and spectral flatness of each band, or the level of each band can be inferred from the overall band level and the band energy distribution ratio. For example, when the full-band detection indicates strong wind noise at level N3, band 1 can be set to level N4, band 2 to level N3, band 3 to level N2, and bands 4 and 5 to level N1. Then, different adaptive noise reduction parameters are matched for different bands. Bands 1 and 2 have higher weights for wind noise levels and are matched with more aggressive time constants, such as a noise spectrum estimation time constant of 0.6-0.7, and an over-subtraction factor of 1.5-2.0. Band 3 uses moderate parameters, with a time constant of 0.8-0.9 and an over-subtraction factor of 1.2-1.5. Bands 4 and 5 are treated conservatively to protect the speech harmonic structure, with a time constant of 0.95 and an over-subtraction factor of 1.0-1.1. The higher the wind noise energy, the stronger the noise overestimation is needed to avoid residual wind noise, but an over-subtraction factor greater than 2.5 will cause speech distortion, so it is limited to within 2.0. Meanwhile, the neural network denoising path can also employ different processing strategies for different frequency bands. For example, a more complex sub-network, such as more convolutional layers or larger hidden layer dimensions, can be used for low-frequency bands, while lightweight processing can be used for high-frequency bands to reduce the number of channels. The processing results of each frequency band are then recombined in the frequency domain and inversely transformed to obtain the complete second denoised signal.
[0034] This implementation achieves fine-grained frequency domain control of wind noise reduction, avoiding excessive damage to high-frequency speech or insufficient suppression of low-frequency wind noise caused by global uniform processing. A stronger noise reduction strategy is adopted in the low-frequency band to suppress the main wind noise, while a more conservative strategy is adopted in the high-frequency band to protect the clarity and naturalness of speech. At the same time, it is convenient to allocate different computing resources for different frequency bands, which is beneficial to the efficiency optimization of embedded platforms.
[0035] In one embodiment, the target deep neural network model is pruned or quantized at runtime, and the running scale of the target deep neural network model is dynamically adjusted according to the current computing resource status or wind noise level. When the wind noise level is lower than a preset threshold, the scaled-down model is enabled to reduce power consumption.
[0036] It's important to note that runtime pruning refers to dynamically skipping the computation of certain neurons or convolutional channels during model inference based on the characteristics of the input signal or the current running state. For example, for a GRU layer, the computation of neurons with activation values below a threshold, such as 10% of the maximum activation value, can be skipped based on the activation value of the current frame. Quantization converts floating-point model parameters into 8-bit integer representations, reducing model storage space and computational load by approximately 75%. The dynamic adjustment strategy for model size is strongly correlated with wind noise level. A wind noise level threshold is preset, for example, using level N2 as the dividing line. When wind noise levels are detected as N0 or N1, the wind noise itself has relatively little interference with speech, and the benefits of using complex, high-precision neural network models are limited. In this case, a pre-pruned, scaled-down model can be enabled, for example, reducing the number of channels by 30% to 50%, or reducing the GRU hidden layer dimension from 256 to 128, thereby reducing computational load by approximately 40% to 60%. When wind noise levels are N2, N3, or N4, the full model capability is required to ensure noise reduction, and the full model is switched back. The current computing resource status can also be used as a basis for adjustment. For example, when the battery level is below 20% or the device temperature exceeds 50°C, the model running scale can be actively reduced to extend battery life or prevent overheating. Model switching can be done using a hot-switching method, that is, the full model and some parameters of the reduced model are simultaneously retained in memory in advance, and the switching is fast at the frame boundary as needed. The latency introduced by the switching process is less than 20-40ms of a frame duration, and there will be no perceptible audio interruption.
[0037] This implementation achieves the adaptation of the deep neural network noise reduction model to computing resources and wind noise scenarios. It uses a lightweight model to reduce average power consumption in weak wind noise and activates the full model to ensure noise reduction effect in strong wind noise. It is particularly suitable for cycling walkie-talkies worn for a long time, and significantly reduces average power consumption while ensuring noise reduction performance in extreme wind noise.
[0038] like Figure 3 In one embodiment, speech rate information or speech activity density information of the audio signal to be processed is obtained, and the smoothing time constant of the weight parameters when the first noise reduction signal and the second noise reduction signal are fused is dynamically adjusted according to the speech rate information or speech activity density information. The smoothing time constant is smaller when the speech rate is faster or the speech activity density is higher, so as to speed up the weight response speed.
[0039] Understandably, speech rate information can be obtained by estimating the syllable or phoneme rate of the speech signal. A simple and effective method is to calculate the number of zero-crossing rate changes per second or the number of peaks in the short-time energy envelope, classifying speech rates into three levels: slow (less than 2 syllables / second), medium (2-4 syllables / second), and fast (more than 4 syllables / second). Speech activity density refers to the proportion of speech frames per unit time, for example, the percentage of speech frames out of the total number of frames in the most recent second. When applying a first-order low-pass smoothing filter to the fusion weights, a time constant greater than or equal to 0.9 results in lag in weight adjustment when wind noise levels change rapidly; a time constant less than or equal to 0.3 can cause unnatural listening experiences when speech rates are slow or during long silences. By introducing speech rate or speech activity density information, the smoothing time constant can be dynamically adjusted. Specifically, when the speech rate is fast or the speech activity density is higher than 70%, human hearing is less sensitive to changes in acoustic details. In this case, a smaller smoothing time constant, such as 0.3-0.5, is used to allow the fusion weights to respond quickly to changes in wind noise levels, ensuring noise reduction performance. When the speech rate is slow or the speech activity density is lower than 30%, human hearing is more sensitive to noise residue and abrupt changes in timbre. In this case, a larger smoothing time constant, such as 0.85-0.95, is used to ensure a smooth transition of weights and avoid perceptible abrupt changes. When the speech rate is medium or the speech activity density is moderate, a compromise time constant, such as 0.6-0.7, is used. The above time constant range is set based on statistical results from human auditory perception experiments. At fast speech rates, the human ear's temporal resolution for amplitude modulation is approximately 100-200 ms, corresponding to a smoothing time constant of 0.3-0.5 to achieve this response speed. At slow speech rates or in silence, the human ear's sensitivity to steady noise requires a smoothing time constant greater than 0.8 to avoid auditory abruptness.
[0040] This implementation associates the temporal characteristics of speech activity with the smoothing strategy of fusion weights, realizing adaptive optimization of fusion transition behavior to human auditory perception. It speeds up response during continuous fast speech and smoothly transitions during slow speech or silent segments, further improving the subjective auditory quality of the fusion output.
[0041] like Figure 4 In one embodiment, it is detected whether there is non-wind noise transient strong noise in the audio signal to be processed. If so, the confidence of the masking matrix output by the target deep neural network model is reduced for the time segment where the transient strong noise is located, and the processing weight of the adaptive noise reduction path for that time segment is increased.
[0042] It should be noted that non-wind noise transient strong noise includes, but is not limited to, vehicle horns, construction impact sounds, sudden noise from objects hitting microphones, and electrical spark interference. These types of noise are characterized by extremely short durations, typically 20-200 milliseconds, high energy, a signal-to-noise ratio that may be below -10 dB, a wide spectrum, and statistical characteristics different from wind noise. The short-time energy of each frame is calculated, and then the ratio of the energies of adjacent frames is calculated. When the ratio is greater than 3 times and the number of frames is less than 5, it is determined to be transient strong noise. Alternatively, pulse peak detection of spectral flux can be used, triggering when the spectral flux suddenly increases by more than 5 times the normal value. A preset detection threshold is set, for example, when the short-time energy ratio exceeds 4 and the energy of the previous frame is 5 dB lower than the background noise energy, transient strong noise is confirmed to exist. When transient strong noise is detected, the deep neural network model may not be able to correctly estimate the masking matrix, and directly using its output may produce artificial sound or distortion. In this case, this implementation reduces the confidence level of the masking matrix output by the neural network. The weight of the second denoised signal in the fusion weighting is temporarily reduced from the normal value of 0.6 to 0.1-0.2, or the gain of the masking matrix is limited to no more than 6dB. Simultaneously, the processing weight of the adaptive denoising path for this time segment is increased, raising the weight of the first denoised signal from 0.4 to 0.8-0.9. The adjustment duration is matched to the duration of transient noise, generally lasting 100-300 milliseconds before gradually recovering. First-order smoothing is used during weight recovery to avoid audible jumps caused by hard switching.
[0043] This implementation addresses the problem of insufficient generalization ability of neural network noise reduction models for unknown types and extreme transient noise. It utilizes the robustness of traditional methods to temporarily take over processing in the event of sudden strong noise, effectively suppressing distortion or artificial sound caused by model errors, and significantly improving the robustness and reliability of the dual-path noise reduction system in complex acoustic environments.
[0044] In one embodiment, when determining the fusion weights of the first and second noise-reduced signals based on the wind noise level, a piecewise linear mapping function is used instead of a fixed lookup table to further improve the smoothness of weight changes.
[0045] Specifically, the wind noise level N is set to a range of 0 to 4, corresponding to levels N0 to N4. The target weight values W2 of the second noise reduction signal corresponding to the five discrete levels are pre-set. target The weights are as follows: N0 level: 0.9; N1 level: 0.7; N2 level: 0.5; N3 level: 0.3; N4 level: 0.1. The weight W1 of the first denoised signal... target =1-W2 target For the real-time detected wind noise level N, linear interpolation is used to calculate the current weight, where N may be a non-integer, such as being obtained by weighting the confidence scores of the classifier output. For example, when N=1.3, which is between N1 and N2, then W2=W2. target(N1)+(N-1)×(W2 target(N2) -W2 target(N1) =0.7 + 0.3 × (0.5 - 0.7) = 0.64, W1 = 0.36. This piecewise linear mapping avoids abrupt weight changes caused by discrete level jumps, and works synergistically with the subsequent first-order low-pass smoothing filter to make the fused output smoother and more natural. For weak wind noise N0, the neural network path has high fidelity, and the weights should be maximized; for extremely strong wind noise N4, the traditional path is more robust, and the weights should be maximized; intermediate levels have a linear transition. In practical applications, the above target values can be fine-tuned according to the specific device microphone characteristics.
[0046] This implementation further optimizes the method for generating fusion weights, so that even if the wind noise level fluctuates near the boundary, the change in fusion weights is continuous, reducing the burden on smoothing filtering and improving the consistency of system response.
[0047] In one embodiment, the preset classifier for classifying wind noise levels uses a gradient boosting decision tree or a lightweight neural network and is updated through online incremental learning to adapt to individual differences in different cycling helmets and microphone positions.
[0048] In the implementation, the initial model of the classifier is trained offline using a general dataset. The model input features are eight dimensions, including short-time energy, autocorrelation decay slope, spectral flatness, and the energy ratio of low-frequency (0-500Hz) to the full frequency band. The output is a probability distribution from N0 to N4, and the level corresponding to the highest probability is taken as the final wind noise level. During online operation, whenever the user actively triggers the "calibration" mode or the system automatically detects a short window of continuous, stable, and wind-noise-free speech in the background (e.g., a vehicle speed below 10km / h for 10 seconds), the system collects the current audio signal as a "wind-noise baseline" and uses this baseline to update certain statistical parameters of the classifier, such as the mean and variance of the normalized features, or to fine-tune the weights of the leaf nodes of the GBDT model. The update process uses mini-batch gradient descent with a learning rate of 0.01, and each update only iterates 5-10 times, resulting in low computational overhead and no impact on real-time noise reduction performance. Through online incremental learning, the classifier can adapt to the acoustic coupling characteristics of different helmets, such as the difference in wind noise transfer function between full-face and half-face helmets, and the difference in low-frequency energy caused by different microphone installation positions, thereby continuously improving the accuracy of wind noise level detection.
[0049] This implementation solves the problem of insufficient generalization ability of fixed classifiers under different users and equipment, enabling the dual-path noise reduction system to self-optimize according to the usage environment and maintain excellent noise reduction effect over a long period of time.
[0050] In one embodiment, generating the second noise-reduced signal based on the spectral characteristics and the wind noise level includes the following steps: The enhanced feature is obtained by concatenating the spectral features and the wind noise feature vector corresponding to the wind noise level. The enhanced features are input into the global time-frequency encoder and the low-frequency harmonic encoder respectively, and the corresponding global time-frequency features and low-frequency harmonic features are output. The global time-frequency features and low-frequency harmonic features are fused to obtain the fused features; The fused features are input into a gated loop unit, the fused features are expanded in time sequence, and the time sequence enhancement features are output. The second noise-reduced signal is generated by performing a clean spectrum estimation based on the time-series enhancement features.
[0051] Specifically, the concatenation of the spectral features and the wind noise feature vector corresponding to the wind noise level can be achieved by encoding the wind noise level into a one-hot vector and then concatenating it with the spectral features.
[0052] The target deep neural network model adopts a dual encoder structure. The encoder can be composed of multiple two-dimensional convolutional layers (kernel size 3×3, stride 1×2). The global time-frequency encoder focuses on global time-frequency correlation and outputs global time-frequency features; the low-frequency harmonic encoder focuses on low-frequency harmonic structure and outputs low-frequency harmonic features; the intermediate layer uses gated recurrent units (GRU) for temporal modeling, and unfolds the fused features temporally to obtain temporal enhancement features; the decoder uses transposed convolutional layers to realize spectrum reconstruction based on the temporal enhancement features and generate the second noise-reduced signal.
[0053] This embodiment forms an enhanced input by concatenating spectral features with wind noise level coding vectors, and extracts complementary features from a global time-frequency encoder and a low-frequency harmonic encoder respectively before fusing them. Finally, it is reconstructed by GRU time-series modeling and decoder. It can dynamically adjust the internal feature expression according to the current wind noise intensity, effectively preserve the low-frequency speech harmonic structure and suppress broadband wind noise under strong wind noise, and maintain computational efficiency through a parallel dual encoder structure.
[0054] In one embodiment, the step of generating the second denoised signal by performing clean spectrum estimation based on the temporal enhancement features includes the following steps: The temporal enhancement features are input into the decoder, which outputs a masking matrix and a wind noise masking matrix. The second noise reduction signal is determined based on the strength of the wind noise level, the audio signal to be processed, and the masking matrix; or, the wind noise component is estimated based on the spectral characteristics and the wind noise masking matrix, and the second noise reduction signal is determined based on the audio signal to be processed and the wind noise component.
[0055] The target deep neural network model supports two working paradigms, selecting the appropriate paradigm based on the wind noise level. Specifically, for high wind noise levels (i.e., strong wind noise), the wind noise component is first estimated, and then subtracted from the original noisy frequency to obtain the second denoised signal, thus avoiding signal distortion; for low wind noise levels (i.e., weak wind noise), the second denoised signal is directly estimated.
[0056] Optionally, if the wind noise level meets a preset level threshold condition, the wind noise component is estimated based on the spectral characteristics and the wind noise masking matrix, and the second noise reduction signal is determined based on the audio signal to be processed and the wind noise component; otherwise, the second noise reduction signal is determined based on the audio signal to be processed and the masking matrix. Specifically, the wind noise level can be one of five types (N0~N4), with the wind noise intensity increasing sequentially from N0 to N4. The preset level threshold condition is "wind noise level greater than or equal to N3". When the wind noise level is N3 or N4, the method of "estimating the wind noise component based on the spectral characteristics and the wind noise masking matrix, and determining the second noise reduction signal based on the audio signal to be processed and the wind noise component" is used. When the wind noise level is N0, N1, or N2, the method of "determining the second noise reduction signal based on the audio signal to be processed and the masking matrix" is used.
[0057] Specifically, determining the second noise-reduced signal based on the audio signal to be processed and the masking matrix involves multiplying the masking matrix with the audio signal to be processed; estimating the wind noise component based on the spectral features and the wind noise masking matrix involves multiplying the spectral features with the wind noise masking matrix; and determining the second noise-reduced signal based on the audio signal to be processed and the wind noise component involves subtracting the wind noise component from the audio signal to be processed to obtain the second noise-reduced signal.
[0058] This embodiment simultaneously outputs a speech masking matrix and a wind noise masking matrix through a decoder, and dynamically selects the corresponding noise reduction processing method according to the wind noise level, realizing an adaptive paradigm switching of the neural network model's noise reduction path: in weak wind noise scenarios with a high signal-to-noise ratio, direct masking is used to maintain the naturalness of the speech and low latency; in strong / extremely strong wind noise scenarios with an extremely low signal-to-noise ratio, the wind noise extraction paradigm is switched to reduce the risk of distortion by utilizing the learnability of wind noise features, thereby obtaining the optimal noise reduction quality and speech fidelity across the entire wind noise level range.
[0059] In one embodiment, the step of denoising the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal includes: Based on the adaptive noise reduction parameters, noise spectrum estimation is performed on the audio signal to be processed to obtain noise spectrum estimation parameters; The signal-to-noise ratio parameter is determined based on the audio signal to be processed and the noise spectrum estimation parameters; The probability of speech presence is determined based on the signal-to-noise ratio parameter. Based on the probability of speech presence, the spectral amplitude of the audio signal to be processed is corrected, and the time-domain signal is reconstructed to obtain the first denoised signal.
[0060] Specifically, the improved Minimum Controlled Recursive Averaging (IMCRA) method is used to estimate the noise spectrum of the audio signal to be processed based on adaptive denoising parameters, thus obtaining the first denoised signal. The adaptive denoising parameters are determined according to the following formula: α(N) = α_base · (1 + γ·N / N_max) Where α(N) represents the adaptive noise reduction parameter, α_base represents the preset base parameter, γ represents the preset adjustment coefficient, N_max represents the maximum wind noise level, which can be 4, and N represents the current wind noise level.
[0061] Substituting the current wind noise level into the above formula yields the adaptive noise reduction parameters. The noise spectrum estimation parameters are the noise power spectral density of each frequency point and each time frame obtained by recursively updating the adaptive noise reduction parameters. Then, the posterior signal-to-noise ratio is calculated based on the power spectrum and noise spectrum estimation values of the audio signal to be processed. The prior signal-to-noise ratio is then recursively estimated using the decision-guided method. Based on the prior and posterior signal-to-noise ratios, combined with spectral characteristics, the likelihood ratio method under the Bayesian framework is used to estimate the speech presence probability at each time and frequency point. The speech presence probability is used to optimally correct the spectral gain to obtain the gain factor. The gain factor is applied to the complex spectrum of the noisy signal, and the time-domain signal is reconstructed using the inverse short-time Fourier transform (iSTFT) and the overlapping addition method to obtain the first noise-reduced signal.
[0062] Specifically, when the wind noise level meets a set condition (e.g., reaches a set threshold), an over-attenuation factor greater than 1 is used to overestimate the noise estimate based on the spectral gain correction. This results in a more aggressive attenuation of the noise frequencies of the audio signal being processed, enhancing the wind noise suppression effect. This over-attenuation factor increases monotonically with increasing wind noise level.
[0063] In one embodiment, the above method further includes the following steps: The weight parameters corresponding to the first noise reduction signal and the second noise reduction signal are determined according to the wind noise level. The weight parameters are subjected to a first-order low-pass smoothing filter to obtain the filtered weight parameters. The process of fusing the first noise-reduced signal and the second noise-reduced signal includes: The first and second denoised signals are weighted and fused using the filtered weight parameters.
[0064] Specifically, based on the determined wind noise level, and using a preset mapping rule or lookup table, the first and second weights corresponding to the first and second denoised signals are determined respectively, satisfying that the first weight + second weight = 1. The core principle of the mapping rule is: the lower the wind noise level (the higher the signal-to-noise ratio), the greater the weight of the second denoised signal, to fully leverage its advantage in preserving human voice fidelity; the higher the wind noise level (the lower the signal-to-noise ratio), the greater the weight of the first denoised signal, to ensure processing robustness and avoid uncontrollable distortion of the neural network model under extreme unknown noise.
[0065] In one specific embodiment, the value ranges of the first weight w_1 and the second weight w_2 corresponding to levels N0~N4 are as follows: Level N0 (weak wind noise): w_2 = 0.8~1.0, w_1 = 0.0~0.2; Level N1 (moderate to weak wind noise): w_2 = 0.6~0.8, w_1 = 0.2~0.4; Level N2 (medium wind noise): w_2 = 0.4~0.6, w_1 = 0.4~0.6; Level N3 (strong wind noise): w_2 = 0.2~0.4, w_1 = 0.6~0.8; Level N4 (extremely strong wind noise): w_2 = 0.0~0.2, w_1 = 0.8~1.0.
[0066] The first-order low-pass smoothing filter applied to the weight parameters refers to performing a first-order recursive smoothing process on the original weight parameters determined in the current audio signal frame and the filtered weights of the previous audio signal frame. The core function of this operation is to suppress sudden weight changes caused by frequent switching of wind noise levels or jitter in the classifier output, thus avoiding perceptible abrupt changes in the fused output signal, such as sudden volume changes, abrupt changes in timbre, and discontinuities.
[0067] The final audio signal output after weighted fusion can be expressed as: S_fusion = W_2 · S_2 + W_1 · S_1. Where S_1 and W_1 are the first denoised signal and their corresponding filtered weights, respectively, and S_2 and W_2 are the second denoised signal and their corresponding filtered weights, respectively. Since the smoothed weights change continuously, the output signal also exhibits a smooth transition, resulting in a more natural listening experience.
[0068] This embodiment applies a first-order low-pass smoothing filter to the fusion weights obtained by mapping wind noise levels and dynamically adjusts the smoothing time constant according to the rate of change of levels. This effectively suppresses weight abrupt changes caused by frequent switching of wind noise levels or jitter in the classifier output, making the transition of the fusion output signal smoother and more natural. It avoids the abrupt auditory problems such as sudden changes in volume and timbre caused by weight jumps in traditional hard switching schemes. At the same time, the dynamic time constant ensures that the system's response speed is not affected by excessive smoothing when the wind noise level changes in reality, achieving a smooth, jitter-free, and fast transient response fusion control effect.
[0069] In one embodiment, the method further includes the following steps: sequentially or selectively performing dynamic range compression, adaptive gain control, frequency domain equalization, and noise thresholding based on speech activity detection on the denoised audio signal.
[0070] The dynamic range compression is used to limit the amplitude of large signals to avoid saturation distortion; the gain adjustment range of the adaptive gain control is dynamically set according to the wind noise level; the frequency domain equalization is used to compensate for the spectral distortion introduced by the noise reduction process; and the noise thresholding based on voice activity detection is used to set the output signal to zero or perform signal attenuation processing of a preset amplitude at the pure noise frame position of the audio signal to be processed.
[0071] Specifically, in this embodiment, after outputting the denoised audio signal and before entering the final output, a series of post-processing operations are performed on the signal. These operations (dynamic range compression, adaptive gain control, frequency domain equalization, and noise thresholding based on speech activity detection) can be executed sequentially in a fixed order, or some of them can be selectively executed according to the current signal quality or wind noise level (such as skipping dynamic range compression in the case of weak wind noise to reduce processing delay), adapting to different application scenarios and computing resource constraints.
[0072] The dynamic range compression operation is as follows: the instantaneous or short-term amplitude envelope of the noise-reduced audio signal is detected. When the signal amplitude exceeds the preset threshold, the excess part is attenuated according to the set compression ratio (such as 2:1, 3:1) to prevent clipping distortion when the large signal enters the subsequent amplifier or digital-to-analog converter.
[0073] The adaptive gain control operates as follows: It automatically adjusts the gain based on the long-term average amplitude of the noise-reduced audio signal, automatically amplifying low frequencies and preventing pops in high frequencies. The gain adjustment range (i.e., the interval between the maximum and minimum permissible gain) is not fixed but dynamically set according to the current wind noise level—under strong wind noise (N3 / N4), the upper limit of the gain adjustment range is appropriately reduced to prevent residual wind noise from being excessively amplified.
[0074] The frequency domain equalization operation is as follows: A preset or adaptive filtering curve is used to compensate for spectral dips or bulges in the speech band after noise reduction. The noise threshold processing based on speech activity detection is as follows: An independent speech activity detection module or multiplexing the speech presence probability information from the front end is used to perform speech / noise judgment on each frame of the denoised audio signal. When the current frame is determined to be a pure noise frame, the amplitude of the output signal of that frame is set to zero or attenuated to below a preset noise floor threshold (e.g., attenuation of -30dB), thereby avoiding the transmission of residual noise in speech-free segments and improving the user's auditory comfort.
[0075] This embodiment employs multiple post-processing enhancement modules after the fused output, including dynamic range compression, adaptive gain control, frequency domain equalization, and a voice activity detection noise threshold. Dynamic range compression avoids clipping distortion caused by high-energy impacts; adaptive gain control dynamically adjusts the gain range based on wind noise levels to prevent excessive amplification of residual wind noise; frequency domain equalization compensates for spectral distortion introduced by noise reduction processing to restore natural speech; and the voice activity detection noise threshold completely removes residual background noise in speech-free segments. The synergistic effect of these four modules ensures that the final output signal maintains a clear, natural, and comfortable auditory experience even in strong wind noise environments, significantly improving the quality of intercom communication for users in high-intensity noise environments.
[0076] In one embodiment, the noise classification of the audio signal to be processed includes: Extract signal features from the audio signal to be processed, wherein the signal features include at least one of the following: short-time energy features, autocorrelation features, spectral flatness, and low-frequency energy proportion; A preset classifier is used to classify the noise of the audio signal to be processed according to the signal characteristics, so as to obtain the wind noise level corresponding to the audio signal to be processed.
[0077] Among them, short-time energy features refer to the normalized energy value of each frame of audio signal, which reflects the instantaneous intensity of the signal. In high-speed cycling scenarios, wind noise energy is much higher than speech energy and persists. Short-time energy features can be used as an indicator to determine whether wind noise exists. This feature can be obtained by normalizing the sum of squares or absolute values and logarithmic domain energy of each frame of sampling points.
[0078] The autocorrelation feature is used to calculate the autocorrelation function of an audio signal to analyze the similarity between the signal and its delayed replica. Due to the high randomness of wind noise, its autocorrelation coefficient decays rapidly to zero with delay. In contrast, speech signals have quasi-periodicity, and the autocorrelation coefficient shows a significant peak at the delay corresponding to the pitch period. This feature can be used to distinguish between audio signal frames dominated by wind noise and audio signal frames dominated by speech.
[0079] Spectral flatness (SFM) is the ratio of the geometric mean to the arithmetic mean of the spectrum of an audio signal in each frame, used to describe the flatness of the spectrum. Wind noise, with its broadband random characteristics, has a relatively flat spectrum, with a SFM close to 1. Speech signals, due to their harmonic structure, exhibit a peak-and-valley distribution in their spectrum, resulting in a lower SFM (typically less than 0.3). This characteristic can help determine the type and intensity of noise.
[0080] Low-frequency energy proportion refers to the ratio of energy in the low-frequency range (e.g., 0-500Hz) to the total energy across the entire frequency band. Wind noise energy is mainly concentrated in the low-frequency range (below 200Hz), but it extends to the mid-to-high frequencies (up to 2kHz and above) in high-speed cycling scenarios. By calculating the low-frequency energy proportion and considering its changing trend, the wind noise intensity level can be reflected—the stronger the wind noise, the higher the low-frequency proportion. However, as speed increases further, the mid-to-high frequency components increase, and the low-frequency proportion may first rise and then fall, requiring a comprehensive judgment based on other characteristics.
[0081] The pre-defined classifier can be a lightweight machine learning classification model, such as a decision tree, Gaussian mixture model (GMM), support vector machine (SVM), or small neural network. The input to this classifier is one or more of the aforementioned audio features, and the output is a discrete wind noise level index (e.g., N0~N4). The classifier can be pre-trained using labeled multi-level wind noise samples or adaptively updated during operation.
[0082] In one embodiment, before extracting the spectral features of the audio signal to be processed using a target deep neural network model, and generating a second noise-reduced signal based on the spectral features and the wind noise level, the process includes: Obtain the training dataset; The deep neural network model to be processed is trained using the training dataset, and the parameters of the deep neural network model to be processed are updated to obtain the deep neural network model with updated parameters. The updated deep neural network model is pruned to obtain a lightweight target deep neural network model.
[0083] Specifically, the training dataset should cover wind noise intensity across multiple scenarios and levels, including speech samples from different riding speeds (corresponding to different wind noise levels, covering 0~195km / h), different helmet types (affecting acoustic transfer function and wind noise coupling characteristics), different microphone positions (affecting wind noise spectrum distribution), and different vehicle models (affecting the mixing ratio of wind noise and background noise). Each sample is composed of clean speech mixed with real or simulated wind noise at a specific signal-to-noise ratio and labeled with the corresponding wind noise level. During training, the noisy speech dataset is used as the model input, and the temporal spectrum (or masking matrix) of clean speech is used as the supervision target. The model parameters are iteratively updated using the backpropagation algorithm and optimizer. For the trained deep neural network model, redundant connections, neurons, or convolutional channels that contribute little to the model output are removed to reduce the number of model parameters and computational complexity, while maintaining the model's noise reduction performance as much as possible. After pruning, the number of model parameters is significantly reduced, resulting in a lightweight target deep neural network model that can run in real time and with low power consumption on embedded platforms.
[0084] Please see Figure 2 , Figure 2 This is a schematic block diagram of a microphone audio noise reduction device according to an embodiment of the present application. The microphone audio noise reduction device 100 includes: an acquisition module 110, a noise classification module 120, and a noise reduction module 130. The acquisition module 110 is used to acquire an audio signal to be processed; the noise classification module 120 is used to perform wind noise detection on the audio signal to be processed and determine the wind noise level of the audio signal to be processed; the noise reduction module 130 is used to match adaptive noise reduction parameters according to the wind noise level, and perform noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; using a target deep neural network model, extracting the spectral features of the audio signal to be processed, and generating a second noise-reduced signal according to the spectral features and the wind noise level; and fusing the first noise-reduced signal and the second noise-reduced signal to obtain a noise-reduced audio signal.
[0085] For the functions performed by each module of the microphone audio noise reduction device 100, please refer to the description of the various embodiments of the microphone audio noise reduction method described above in this application, which will not be repeated here.
[0086] Please see Figure 5 , Figure 5 This is a schematic block diagram of a computer device according to an embodiment of the present application. The computer device 900 includes a processor 910 and a memory 920 coupled to each other. The memory 920 stores a computer program, and the processor 910 executes the computer program to implement the microphone audio method for high-speed cycling scenarios described in the above embodiments.
[0087] For a description of each step of the processing, please refer to the description of each step in the above embodiment of the microphone audio noise reduction method of this application, which will not be repeated here.
[0088] The memory 920 can be used to store program data and modules. The processor 910 executes various functional applications and data processing by running the program data and modules stored in the memory 920. The memory 920 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function, etc.; the data storage area may store data created based on the use of the computer device 900, etc. In addition, the memory 920 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 920 may also include a memory controller to provide the processor 910 with access to the memory 920.
[0089] In the various embodiments of this application, the disclosed methods, apparatus, and devices can be implemented in other ways. For example, the embodiments of the computer devices described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0091] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium.
[0093] See Figure 6 , Figure 6 This is a schematic block diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 700 stores program data 710, which, when executed, implements the steps of the microphone audio noise reduction method embodiments described above.
[0094] For a description of each step of the processing, please refer to the description of each step in the above embodiment of the microphone audio noise reduction method of this application, which will not be repeated here.
[0095] Any references to memory, storage, database, or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0096] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0097] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0098] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A microphone audio noise reduction method, characterized in that, The method includes: Acquire the audio signal to be processed; Wind noise detection is performed on the audio signal to be processed to determine the wind noise level of the audio signal to be processed; Based on the wind noise level, adaptive noise reduction parameters are matched, and the audio signal to be processed is denoised based on the adaptive noise reduction parameters to obtain a first denoised signal. Using a target deep neural network model, the spectral features of the audio signal to be processed are extracted, and a second noise-reduced signal is generated based on the spectral features and the wind noise level. This includes: concatenating the spectral features and the wind noise feature vectors corresponding to the wind noise level to obtain enhanced features; inputting the enhanced features into a global time-frequency encoder and a low-frequency harmonic encoder respectively, outputting corresponding global time-frequency features and low-frequency harmonic features; fusing the global time-frequency features and low-frequency harmonic features to obtain fused features; inputting the fused features into a gated loop unit, expanding the fused features in a time sequence, and outputting a time-series enhanced feature; and performing clean spectrum estimation based on the time-series enhanced feature to generate the second noise-reduced signal. The first noise-reduced signal and the second noise-reduced signal are fused to obtain a noise-reduced audio signal, including: determining the weight parameters corresponding to the first noise-reduced signal and the second noise-reduced signal respectively according to the wind noise level; performing a first-order low-pass smoothing filter on the weight parameters to obtain filtered weight parameters; and using the filtered weight parameters to perform weighted fusion of the first noise-reduced signal and the second noise-reduced signal.
2. The method according to claim 1, characterized in that, The step of estimating the clean spectrum based on the time-series enhancement features to generate the second noise-reduced signal includes: The temporal enhancement features are input into the decoder, which outputs a masking matrix and a wind noise masking matrix. The second noise reduction signal is determined based on the strength of the wind noise level, the audio signal to be processed, and the masking matrix; or, the wind noise component is estimated based on the spectral characteristics and the wind noise masking matrix, and the second noise reduction signal is determined based on the audio signal to be processed and the wind noise component.
3. The method according to claim 1, characterized in that, The step of denoising the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal includes: Based on the adaptive noise reduction parameters, noise spectrum estimation is performed on the audio signal to be processed to obtain noise spectrum estimation parameters; The signal-to-noise ratio parameter is determined based on the audio signal to be processed and the noise spectrum estimation parameters; The probability of speech presence is determined based on the signal-to-noise ratio parameter. Based on the probability of the speech presence, the spectral amplitude of the audio signal to be processed is corrected, and the time-domain signal is reconstructed to obtain the first noise-reduced signal.
4. The method according to claim 1, characterized in that, The method further includes: Dynamic range compression, adaptive gain control, frequency domain equalization, and noise thresholding based on speech activity detection are sequentially or selectively applied to the denoised audio signal. The dynamic range compression is used to limit the amplitude of large signals to avoid saturation distortion; the gain adjustment range of the adaptive gain control is dynamically set according to the wind noise level; the frequency domain equalization is used to compensate for the spectral distortion introduced by the noise reduction process; and the noise thresholding based on voice activity detection is used to set the output signal to zero or perform signal attenuation processing of a preset amplitude at the position of pure noise frame.
5. The method according to claim 1, characterized in that, The step of detecting wind noise in the audio signal to be processed and determining the wind noise level of the audio signal to be processed includes: Extract signal features from the audio signal to be processed, wherein the signal features include at least one of the following: short-time energy features, autocorrelation features, spectral flatness, and low-frequency energy proportion; A preset classifier is used to classify the noise of the audio signal to be processed according to the signal characteristics, so as to obtain the wind noise level corresponding to the audio signal to be processed.
6. The method according to claim 1, characterized in that, The method further includes: The motion state information or speed information corresponding to the audio signal to be processed is obtained, the wind noise level for the current or future period is predicted based on the motion state information or speed information, and the predicted wind noise level is used to replace the real-time detected wind noise level.
7. The method according to claim 3, characterized in that, The method further includes: The degree of non-stationarity of wind noise in the audio signal to be processed is detected, and the update rate of the noise spectrum estimation is dynamically adjusted according to the degree of non-stationarity, wherein the higher the degree of non-stationarity, the faster the update rate.
8. The method according to claim 1, characterized in that, The method further includes: The target deep neural network model is pruned or quantized at runtime. The running scale of the target deep neural network model is dynamically adjusted according to the current computing resource status or wind noise level. When the wind noise level is lower than a preset threshold, a scaled-down model is enabled to reduce power consumption.
9. The method according to claim 1, characterized in that, The method further includes: The speech rate information or speech activity density information of the audio signal to be processed is obtained, and the smoothing time constant of the weight parameters when fusing the first noise reduction signal and the second noise reduction signal is dynamically adjusted according to the speech rate information or speech activity density information. The smoothing time constant is smaller when the speech rate is faster or the speech activity density is higher, so as to speed up the weight response speed.
10. The method according to claim 2, characterized in that, The method further includes: If any non-wind noise transient strong noise exists in the audio signal to be processed, the confidence of the masking matrix output by the target deep neural network model is reduced for the time segment where the transient strong noise is located, and the processing weight of the adaptive noise reduction path for that time segment is increased.
11. A microphone audio noise reduction device, characterized in that, The device includes: The acquisition module is used to acquire the audio signal to be processed; A noise grading module is used to detect wind noise in the audio signal to be processed and determine the wind noise level of the audio signal to be processed. The noise reduction module is used to match adaptive noise reduction parameters according to the wind noise level, and perform noise reduction on the audio signal to be processed based on the adaptive noise reduction parameters to obtain a first noise-reduced signal; using a target deep neural network model, it extracts the spectral features of the audio signal to be processed, and generates a second noise-reduced signal based on the spectral features and the wind noise level, specifically including: concatenating the spectral features and the wind noise feature vector corresponding to the wind noise level to obtain enhanced features; inputting the enhanced features into a global time-frequency encoder and a low-frequency harmonic encoder respectively, and outputting the corresponding global time-frequency features and low-frequency harmonic features; and combining the global time-frequency features and low-frequency harmonic features... The first and second noise-reduced signals are fused to obtain fused features. These fused features are then input into a gated loop unit, which expands them temporally to output temporal enhancement features. A clean spectrum estimation is performed based on these temporal enhancement features to generate the second noise-reduced signal. The first and second noise-reduced signals are then fused to obtain a noise-reduced audio signal. Specifically, this includes: determining weight parameters corresponding to the first and second noise-reduced signals based on the wind noise level; performing a first-order low-pass smoothing filter on the weight parameters to obtain filtered weight parameters; and using the filtered weight parameters to perform weighted fusion of the first and second noise-reduced signals.
12. A computer device, characterized in that, The device includes a processor and a memory coupled to each other; the memory stores a computer program, and the processor executes the computer program to implement the steps of the method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data that, when executed by a processor, implements the steps of the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Grid denoising method and device based on deep learning and multi-level fusion
CN117390524A
Wind noise suppression method, system and device for audio signal and storage medium
CN118072754A