Audio device wind noise suppression method, apparatus, device, and storage medium
By implementing independent dynamic range control for each frequency band and gain coordination between frequency bands, the problem of damage to the dynamics and listening experience of mid-to-low frequency music when suppressing wind noise in existing audio equipment has been solved, achieving efficient wind noise suppression and sound quality protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINKPLAY TECHNOLOGY INC NANJING
- Filing Date
- 2026-05-13
- Publication Date
- 2026-06-26
AI Technical Summary
Existing audio equipment generally suffers damage to the dynamics and listening experience of mid-to-low frequency music when suppressing wind noise, and lacks the ability to accurately identify and adaptively process high-frequency wind noise, resulting in an overall decline in sound quality.
It adopts independent dynamic range control for each frequency band, and configures low compression threshold, high compression ratio and fast start-up strategy for the high frequency band. It combines spectral flatness detection and drive strength information for adaptive updates, and introduces an inter-band gain coordination mechanism to ensure the dynamic integrity and natural listening experience of the mid and low frequency bands.
While effectively suppressing high-frequency wind noise, it preserves the dynamic integrity and auditory layering of mid- and low-frequency music content to the greatest extent, improves the ability to distinguish wind noise components from normal high-frequency music, and avoids spectral imbalance and abrupt changes in timbre.
Smart Images

Figure CN122294048A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and in particular to methods, apparatus, devices and storage media for suppressing wind noise in audio equipment. Background Technology
[0002] In high-volume operation scenarios of loudspeaker systems, the high-amplitude vibration of the diaphragm drives rapid airflow. When the airflow velocity exceeds a certain threshold, turbulence occurs at the edges of structures such as the sound tube, surround, dust cap, and speaker enclosure, generating broadband aerodynamic noise, commonly known as wind noise. The wind noise spectrum is mainly concentrated in the high-frequency range of 3kHz to 20kHz, dominated by airflow boundary layer separation and the Karman vortex street effect. Its sound pressure level increases nonlinearly with the driving power, reaching -20dBFS to -10dBFS when the power is close to the rated upper limit, severely degrading the signal-to-noise ratio of music signals.
[0003] Existing audio equipment has the following shortcomings in dealing with wind noise: First, existing solutions generally employ full-band amplitude limiters or full-band dynamic range control to apply uniform gain attenuation across the entire frequency band. While suppressing high-frequency wind noise, this approach simultaneously compresses core musical content such as mid- and low-frequency vocals and strings, resulting in a decrease in low-frequency impact, reduced clarity of mid-frequency vocals, and damage to overall dynamic range.
[0004] Secondly, some solutions employ passive hardware filtering protection, such as high-frequency current-limiting resistors in the crossover or fixed high-frequency attenuation structures. These static mechanisms cannot dynamically adjust based on the playback content, device drive status, and wind noise intensity. In non-wind-noise-dominated scenarios, they can introduce unnecessary high-frequency attenuation, affecting sound quality.
[0005] Third, existing technologies generally lack the ability to effectively identify and adaptively intervene in the spectral characteristics of wind noise, making it difficult to distinguish high-frequency wind noise components from normal high-frequency music signals such as cymbals and overtones, and easily leading to problems such as protection lag, miscompression, or overcompression.
[0006] In summary, existing technologies lack the ability to accurately identify the high-frequency concentrated characteristics of wind noise and the ability to process frequency bands independently. There is an urgent need for a new wind noise suppression solution that can effectively suppress high-frequency wind noise while taking into account the dynamic integrity of mid- and low-frequency music and the overall naturalness of the listening experience. Summary of the Invention
[0007] The purpose of this invention is to provide a method for suppressing wind noise in audio devices, in order to solve the technical problems of existing full-band processing solutions that damage the dynamics of mid- and low-frequency music and have low protection efficiency when suppressing wind noise. This method effectively suppresses high-frequency wind noise while preserving the dynamic integrity of mid- and low-frequency music content and the overall naturalness of the listening experience as much as possible.
[0008] The first aspect of this invention provides a method for suppressing wind noise in audio devices, comprising: The input audio signal is acquired and frequency-divided to separate multiple frequency band signals, including at least low-frequency and high-frequency signals. Based on the independent dynamic range control parameters corresponding to each frequency band, independent dynamic range control processing is performed on the signals of each frequency band. Among them, a lower compression threshold, a higher compression ratio, and a shorter start-up time are configured for the high-frequency band signals compared with other frequency band signals, so that the dynamic range control processing of the high-frequency band prioritizes the suppression of the signal peak of the wind noise-dominant frequency band. Extract the wind noise characterization features of the high-frequency band signal, combine them with the device drive strength information to determine the current wind noise scene, and adaptively update the dynamic range control parameters corresponding to the high-frequency band signal based on the determination result. Monitor the gain compression status of signals in each frequency band after dynamic range control processing, and determine the gain difference between frequency bands. When the gain difference exceeds a preset coordination threshold, perform gain coordination processing on the frequency band with smaller compression. The signals from each frequency band, after dynamic range control and gain coordination processing, are recombined into a full-band output signal.
[0009] Preferably, the plurality of frequency band signals include low-frequency band signals, mid-frequency band signals, and high-frequency band signals, and the frequency division processing includes: The full-band audio signal is divided into two levels according to the preset first and second frequency division points; Wherein, the first frequency division point is used to separate the low-frequency band signal and the high-frequency mixed band signal, and the second frequency division point is used to further separate the high-frequency mixed band signal into the mid-frequency band signal and the high-frequency band signal, so that the high-frequency band covers the frequency band that mainly generates wind noise; The low-frequency signal corresponds to a frequency range of 20Hz to 300Hz, the mid-frequency signal corresponds to a frequency range of 300Hz to 3kHz, and the high-frequency signal corresponds to a frequency range of 3kHz to 20kHz.
[0010] Preferably, the frequency division process further includes: The input audio signal is frequency-divided and filtered using a cross-frequency divider filter bank. The cross-frequency divider filter bank matches the amplitude response of adjacent frequency bands at the corresponding frequency division point, and the corresponding high-pass filter path and low-pass filter path have matching phase response, so that the signals of each frequency band can obtain a flat frequency response after recombination, and eliminate the spectral distortion introduced by frequency division.
[0011] Preferably, the frequency division processing further includes: obtaining the group delay introduced by the cross-frequency division filter bank on each frequency band filtering path; and introducing a corresponding delay compensation amount for at least one frequency band signal according to the group delay difference of each frequency band path, so that the signals of each frequency band achieve time domain alignment before reassembly, avoid time domain phase vanishing truth caused by group delay difference, and ensure the time domain integrity of the reassembled signal.
[0012] Preferably, the step of performing independent dynamic range control processing on signals of each frequency band includes: Each frequency band is configured with an independent dynamic range control parameter group. Each parameter group includes at least compression threshold, compression ratio, start-up time, release time, and knee width. The configuration of the parameter groups for each frequency band ensures that the compression start-up speed and compression intensity of the high-frequency band dynamic range control processor for signals exceeding the threshold are higher than those of the corresponding dynamic range control processors for other frequency bands, while the compression start-up speed and compression intensity for the low-frequency band are the lowest. This is to accurately suppress high-frequency wind noise while preserving the dynamic integrity of low-frequency music content to the greatest extent.
[0013] Preferably, the process of performing independent dynamic range control further includes: An independent sidechain detection channel is established for each frequency band. The root mean square level detection method is used to detect the input signal of the corresponding frequency band. A low-pass filter with a preset time constant is used to smooth the square value of the input signal of the corresponding frequency band to generate the instantaneous root mean square level of the corresponding frequency band. The target compression gain is calculated based on the relationship between the instantaneous root mean square level and the corresponding frequency band compression threshold; The target compression gain is smoothed based on the start-up time constant and release time constant of the corresponding frequency band to obtain a smoothed compression gain, and the smoothed compression gain is applied to the signal of the corresponding frequency band.
[0014] Preferably, the step of extracting the wind noise characterization features of the high-frequency signal and determining the current wind noise scene in conjunction with the device drive strength information includes: The high-frequency band signal is frequency domain transformed, and the spectral flatness index is calculated based on the ratio of the geometric mean to the arithmetic mean of the spectral amplitude. The current driving power is estimated based on the rated power, digital volume control value and current root mean square signal level. The wind noise intensity estimate is generated according to the preset mapping relationship between the current driving power and the wind noise intensity. A current wind noise scene determination result is generated based on at least one of the spectrum flatness index and the wind noise intensity estimate; wherein, the higher the spectrum flatness index and the larger the wind noise intensity estimate, the more the wind noise scene determination result tends to be a high wind noise scene.
[0015] Preferably, the adaptive updating of the dynamic range control parameters corresponding to the high-frequency band signal based on the determination result includes: Based on the spectral flatness index and / or the wind noise intensity estimate, corresponding threshold adjustment amounts are generated respectively, and the threshold adjustment amounts of each path are fused and calculated according to preset weights to obtain the comprehensive adjustment amount of the high-frequency band compression threshold. The overall adjustment amount is limited to a preset adjustment range, and the high-frequency band compression threshold is smoothly updated according to a preset update cycle to prevent the threshold from being excessively lowered, causing unnecessary compression to normal high-frequency music content.
[0016] Preferably, the determination of gain differences between frequency bands, when the gain difference exceeds a preset coordination threshold, involves performing gain coordination processing on the frequency band with the smaller compression amount, including: The gain compression amount of each frequency band is acquired in real time, and the gain difference between adjacent frequency bands is calculated. When the absolute value of the gain difference between any adjacent frequency band exceeds the preset coordination threshold, the frequency band with the smaller gain compression amount in the adjacent frequency band is identified as the coordination compensation object. The target coordination gain is determined based on the degree of deviation between the gain difference and the preset coordination threshold, so that the gain difference between adjacent frequency bands converges to the target gain difference range. Within a preset update cycle, the target coordination gain is injected into the coordination compensation object in an exponentially smooth manner, so that the coordination gain smoothly approaches the target coordination gain within a preset transition time, avoiding the introduction of perceptible gain jumps or modulation noise by the coordination action.
[0017] Preferably, the method further includes: when the gain compression of the high-frequency signal exceeds a preset protection threshold, determining that the current audio device is in an extreme wind noise protection scenario; in response to the extreme wind noise protection scenario: adjusting the dynamic range control parameters corresponding to the low-frequency and mid-frequency bands in a coordinated manner to implement full-band protective compression for each frequency band and reduce the overall excitation power of the speaker; sending risk alarm information to the device protection management module and triggering a volume limiting mechanism to limit the upper limit of the device's output volume; and outputting protection prompt information to the device display or associated application to guide the user to actively reduce the volume.
[0018] As a preferred embodiment, real-time quality monitoring of the recombined full-band output signal is also included. The time-domain aligned and linearly superimposed signals of each frequency band after delay compensation are obtained to obtain a full-band output signal. The spectral difference between the full-band output signal and the reference signal in the non-wind noise-dominated frequency band is calculated to obtain the spectral reconstruction error, which characterizes the degree of spectral distortion introduced by the inter-band gain coordination processing in the reconstructed output signal. When the spectrum reconstruction error exceeds a preset error threshold, it is determined that the inter-band coordination gain injection amount exceeds a reasonable range, and feedback triggers the reduction control of the coordination gain; the upper limit of the coordination gain injection amount is dynamically adjusted according to the feedback result of the spectrum reconstruction error to realize closed-loop adaptive control of the coordination gain, so as to prevent the non-wind noise frequency band from being excessively indirectly compressed due to gain coordination processing.
[0019] As a preferred option, dynamic range control parameter self-learning and device-specific calibration are also included: When the audio device has accumulated a preset learning time, it performs device-level parameter self-learning optimization based on historical operating data. The historical operating data includes at least wind noise characterization statistics, wind noise intensity estimation data, and the spectrum reconstruction error data. Based on the self-learning optimization results, for the measured wind noise characteristics of the current equipment instance, at least one of the following parameters is automatically fine-tuned: high-frequency dynamic range control parameters, inter-band coordination parameters, and wind noise intensity estimation parameters. The fine-tuned parameters are written to non-volatile memory for loading when the audio device starts up later, thus achieving device-level wind noise protection that adapts to individual differences.
[0020] A second aspect of the present invention provides an audio device wind noise suppression device, comprising: The frequency division matrix module is used to acquire the input audio signal and perform frequency division processing to separate it into multiple frequency band signals, including at least low-frequency signals and high-frequency signals. A multi-band independent DRC processing module is used to perform independent dynamic range control processing on each frequency band signal based on the independent dynamic range control parameters corresponding to each frequency band. Specifically, for the high-frequency band signal, a lower compression threshold, a higher compression ratio, and a shorter start-up time are configured compared to other frequency band signals, so that the high-frequency band dynamic range control processing prioritizes the suppression of signal peaks in the wind noise-dominant frequency band. The wind noise detection and adaptive threshold update module is used to extract the wind noise characterization features of the high-frequency signal, combine the device drive strength information to determine the current wind noise scene, and adaptively update the dynamic range control parameters corresponding to the high-frequency signal according to the determination result. The inter-band gain coordination module is used to monitor the gain compression status of signals in each frequency band after dynamic range control processing and to determine the gain difference between frequency bands. When the gain difference exceeds a preset coordination threshold, gain coordination processing is performed on the frequency band with smaller compression. The frequency band signal reassembly module is used to reassemble the signals of each frequency band after dynamic range control processing and gain coordination processing into a full-band output signal.
[0021] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the above-described audio device wind noise suppression method.
[0022] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described audio device wind noise suppression method.
[0023] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention employs independent dynamic range control for each frequency band. It employs an aggressive strategy with low compression thresholds, high compression ratios, and rapid oscillation for the high-frequency bands where wind noise is concentrated, while maintaining a more lenient independent processing strategy for the mid- and low-frequency bands. Compared to a uniform compression scheme across the entire frequency band, this invention can accurately suppress high-frequency wind noise in high-volume scenarios, while effectively reducing dynamic compression of mid- and low-frequency music content, thus better preserving the power of low frequencies, the clarity of mid-frequency vocals, and the overall listening experience.
[0024] 2. This invention combines spectral flatness detection and drive intensity tracking to quantify and characterize high-frequency wind noise features, and adaptively updates high-frequency compression parameters accordingly. This mechanism helps improve the ability to distinguish wind noise components from normal high-frequency music content (such as cymbal sounds and string overtones), enabling the suppression strategy to dynamically adjust with changes in playback content and wind noise intensity, thereby reducing the probability of false compression.
[0025] 3. This invention introduces an inter-band gain coordination mechanism to smoothly control the amount of coordinated gain injection, effectively suppressing the spectral imbalance caused by independent frequency division processing, avoiding abrupt changes in timbre or a sense of frequency band fragmentation, and improving the naturalness of the spectral connection and the consistency of the listening experience of the reconstructed output signal.
[0026] 4. In extreme wind noise scenarios, this invention can trigger full-band linkage protection, adjust the dynamic range control parameters of the low-frequency and mid-frequency bands in a linkage manner, and limit the upper limit of the output volume. While effectively suppressing wind noise, it reduces the risk of speaker overload, and balances playback stability and equipment safety. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating the wind noise suppression method for audio devices in this embodiment; Figure 2 This is a schematic diagram of the audio device wind noise suppression device in this embodiment; Figure 3 This is a schematic diagram of one embodiment of the computer device in this example. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] For ease of understanding, the specific process of this invention embodiment is described below. This invention embodiment provides an audio device wind noise suppression method, which can be applied to speakers, smart audio terminals, car audio systems, headphone amplifiers, or other electronic devices with audio playback capabilities. Addressing the wind noise problem that easily occurs in audio devices during high-volume playback, this embodiment utilizes the characteristic that wind noise is mainly concentrated in the high-frequency range. A frequency-band dynamic wind noise suppression system is constructed in the signal output processing link of the audio device to achieve wind noise suppression while minimizing the impact on normal music content and listening experience.
[0030] In some embodiments, the frequency-band wind noise dynamic suppression system may include a frequency division matrix module (CMM), a multi-band independent dynamic range control processing module (MIDM), a frequency band gain coordination module (CGCM), a frequency-band signal reconstruction module (MSRM), and a wind noise detection and adaptive threshold update module (WNDATM). The overall processing flow is as follows: the input PCM signal is separated into three frequency bands by the frequency division matrix module to obtain low-frequency, mid-frequency, and high-frequency signals; subsequently, the multi-band independent dynamic range control processing module performs independent dynamic range control processing on each frequency band; then, the frequency band gain coordination module coordinates the gain differences between frequency bands generated after independent compression; finally, the frequency band signal reconstruction module reconstructs the signals of each frequency band to obtain the output PCM signal; simultaneously, the wind noise detection and adaptive threshold update module can detect wind noise in the high-frequency signal and update the corresponding dynamic range control parameters of the high-frequency band based on the device drive strength information, so that the wind noise suppression strategy matches the current operating state.
[0031] Specifically, such as Figure 1 As shown, the audio device wind noise suppression method of this embodiment may include the following steps: Step S1: Acquire the input audio signal and perform frequency division processing to separate multiple frequency band signals, including at least low-frequency signals and high-frequency signals.
[0032] Step S2: Based on the independent dynamic range control parameters corresponding to each frequency band, perform independent dynamic range control processing on each frequency band signal. Specifically, for the high-frequency band signal, configure a lower compression threshold, a higher compression ratio, and a shorter start-up time compared to other frequency band signals, so that the high-frequency band dynamic range control processing prioritizes suppressing the signal peak of the wind noise-dominant frequency band.
[0033] Step S3: Extract the wind noise characterization features of the high-frequency signal, combine the device drive strength information to determine the current wind noise scenario, and adaptively update the dynamic range control parameters corresponding to the high-frequency signal according to the determination result, so that the dynamic range control parameters match the current wind noise state.
[0034] Step S4: Monitor the gain compression status of each frequency band signal after dynamic range control processing, and determine the gain difference between frequency bands. When the gain difference exceeds the preset coordination threshold, perform gain coordination processing on the frequency band with smaller compression amount to reduce the spectral imbalance caused by independent frequency division compression.
[0035] Step S5: Reassemble the frequency band signals after dynamic range control processing and gain coordination processing into a full-band output signal.
[0036] This embodiment divides the input audio signal into frequencies and performs independent dynamic range control for different frequency bands, enabling wind noise suppression to focus on the high-frequency bands where wind noise tends to concentrate. Simultaneously, by combining the characteristics of high-frequency wind noise with equipment drive strength information to adaptively update the high-frequency control parameters, the wind noise suppression strategy's adaptability to different wind noise conditions and equipment operating states can be enhanced. Furthermore, through inter-band gain coordination and, when necessary, full-band linkage protection, wind noise suppression effectiveness, natural sound quality, and equipment safety can be balanced.
[0037] In some implementations, the frequency division processing of the input audio signal in step S1 may specifically include: performing two-level frequency division on the full-band audio signal (20Hz-20kHz) according to a preset first frequency division point and a second frequency division point; wherein, the first frequency division point is used to separate the low-frequency band signal from the high-frequency mixed band signal, and the second frequency division point is used to further separate the high-frequency mixed band signal into a mid-frequency band signal and a high-frequency band signal, so that the high-frequency band covers the main frequency band of wind noise generation; the first frequency division point is used to distinguish between the low-frequency band and the mid-high frequency band, and the second frequency division point is determined according to the spectral density inflection point in the wind noise spectrum, so that the high-frequency band covers the main frequency band of wind noise generation. Preferably, the low-frequency band signal corresponds to a frequency range of 20Hz to 300Hz, the mid-frequency band signal corresponds to a frequency range of 300Hz to 3kHz, and the high-frequency band signal corresponds to a frequency range of 3kHz to 20kHz.
[0038] Based on the frequency band division method described above, the system achieves precise allocation of processing strategies according to the spectral characteristics analysis of wind noise. Specifically, the low-frequency band (LF) mainly carries low-frequency energy information and auditory impact, and its contribution to wind noise is the lowest. Therefore, the most lenient dynamic range control (DRC) strategy is subsequently configured. The mid-frequency band (MF) mainly carries core musical information such as vocals, strings, and guitars, and its contribution to wind noise is relatively low. Therefore, a more lenient DRC strategy is subsequently configured to prioritize the preservation of the integrity of the mid-frequency information of the music. The high-frequency band (HF) mainly carries high-frequency detail information and is the most important frequency band for generating wind noise. Therefore, the most aggressive DRC strategy is subsequently configured to precisely suppress high-frequency wind noise as the core objective.
[0039] Specifically, the second crossover point (3kHz) was chosen based on the fact that the turbulent noise spectrum of aerodynamic wind noise exhibits a significant spectral density inflection point near 3kHz, and the wind noise pressure level above 3kHz is approximately 3-5 times that below 3kHz (power spectral density ratio). Therefore, using 3kHz as the dividing point between the mid and high frequencies can isolate the dominant wind noise component to the high-frequency band for independent processing to the greatest extent possible. This facilitates the subsequent configuration of differentiated dynamic range control parameters for different frequency bands, achieving the optimal balance between wind noise suppression and maintaining normal listening experience.
[0040] As an example, and not a limitation, a measured analysis of the wind noise spectrum of a WiiM Bar sound stick product (collected at 85dB SPL output power) shows the following: The average sound pressure level of wind noise in the LF band (20-300Hz) is approximately -48dBFS, in the MF band (300Hz-3kHz) approximately -36dBFS, and in the HF band (3kHz-20kHz) approximately -22dBFS. The wind noise in the HF band is approximately 26dB higher than that in the LF band, strongly confirming the rationality of using 3kHz as the crossover point. Based on this analysis, the system confirms that the HF band is the frequency band requiring focused DRC protection.
[0041] In other embodiments, the first and second frequency division points can also be adaptively adjusted according to the audio device type, speaker response characteristics, acoustic structure, or wind noise statistics. This embodiment does not limit this.
[0042] In some implementations, the frequency division process in step S1 of the input audio signal further includes: The input audio signal is frequency-divided and filtered using a cross-frequency filter bank (e.g., a fourth-order Linkwitz-Riley filter bank, i.e., an LR4 filter bank); specifically, a two-stage frequency division structure is used: the first stage uses 300Hz as the division point to divide the full-band signal into a low-frequency component S. LF (n) (LR4 low-pass, cutoff 300Hz) and high-frequency mixed component S HMF(n) (LR4 high-pass filter, cutoff 300Hz); the second stage uses 3kHz as the division point for S. HMF (n) is further divided into intermediate frequency components S MF (n) (LR4 low-pass, cutoff 3kHz) and high-frequency component S HF (n) (LR4 high-pass, cutoff 3kHz).
[0043] The cross-frequency divider filter bank matches the amplitude responses of adjacent frequency bands at the corresponding division points (restoring a 0dB flat response after superposition of the two paths), and the corresponding high-pass and low-pass filter paths have matching phase responses, so that the recombined signals of each frequency band obtain a flat frequency response, eliminating the spectral distortion introduced by the frequency division. As an example, and not a limitation, at a system sampling rate of 48kHz, a fourth-order Linkwitz-Riley low-pass filter can be constructed from two cascaded second-order Butterworth low-pass filters (each stage has a Q value of 0.7071, i.e., 1 / √2), and the coefficient set is converted from the analog prototype to the digital domain (Z domain) through a bilinear transform. Actual measured frequency division accuracy can reach: at the 300Hz division point, the amplitude difference between the two components is less than 0.1dB, and the phase difference is less than 1 degree, meeting the requirements for accurate frequency division.
[0044] In some implementations, the frequency division processing in step S1 further includes: acquiring the group delay introduced by the cross-frequency divider filter bank on each frequency band filtering path; and, based on the group delay difference of each frequency band path, introducing a corresponding delay compensation amount for at least one frequency band signal to ensure that the signals of each frequency band achieve time-domain alignment before reassembly, avoiding time-domain phase vanishing distortion caused by group delay differences, and ensuring the time-domain integrity of the reassembled signal. Since the LR4 filter introduces different group delays in different frequency bands, the signals of the three frequency bands have delay differences in the time domain, and direct addition will produce time-domain phase vanishing distortion; therefore, delay compensation is necessary. As an example, and not a limitation, at a sampling rate of 48kHz, the group delay introduced by the 300Hz LR4 low-pass filter is approximately 3.3ms (159 sampling points), and the group delay introduced by the 3kHz LR4 low-pass filter is approximately 0.33ms (16 sampling points). The system is for high-frequency component S. HF (n) Introducing a delay compensation of 3.63ms (175 sampling points) for the intermediate frequency component S MF (n) A delay compensation of 3.3ms (159 sampling points) is introduced to ensure that the total delay of the low, medium and high signals is uniformly 3.63ms, thus completely eliminating the time-domain phase vanishing truth.
[0045] In one embodiment, step S2, which involves performing independent dynamic range control processing on signals of each frequency band based on independent dynamic range control parameters corresponding to each frequency band, includes: configuring independent dynamic range control parameter groups for each frequency band, wherein each parameter group includes at least a compression threshold, compression ratio, start-up time, release time, and knee width; wherein the configuration of the parameter groups for each frequency band ensures that the compression start-up speed and compression intensity of the high-frequency band dynamic range control processor for signals exceeding the threshold are higher than those of the dynamic range control processors corresponding to other frequency bands, and the compression start-up speed and compression intensity of the low-frequency band are the lowest, so as to preserve the dynamic integrity of low-frequency music content to the greatest extent while accurately suppressing high-frequency wind noise. For example, a first parameter group is configured for the low-frequency signal, a second parameter group is configured for the mid-frequency signal, and a third parameter group is configured for the high-frequency signal; wherein the first parameter group, the second parameter group, and the third parameter group each include at least a compression threshold, a compression ratio, an oscillation start time, a release time, and a knee width; the compression threshold corresponding to the first parameter group is higher than that of the second parameter group and the third parameter group, the compression ratio corresponding to the first parameter group is lower than that of the second parameter group and the third parameter group, and the oscillation start time corresponding to the first parameter group is longer than that of the second parameter group and the third parameter group; the compression threshold, compression ratio, and oscillation start time corresponding to the second parameter group are respectively between those of the first parameter group and the third parameter group; the compression threshold corresponding to the third parameter group is lower than that of the first parameter group and the second parameter group, the compression ratio corresponding to the third parameter group is higher than that of the first parameter group and the second parameter group, and the oscillation start time corresponding to the third parameter group is shorter than that of the first parameter group and the second parameter group, so that the high-frequency signal responds preferentially and wind noise is suppressed.
[0046] As an example, and not a limitation, the compression threshold of the first parameter group (low frequency band) can be -6dBFS, the compression ratio can be 2:1, the start-up time can be 50ms, the release time can be 300ms, and the knee width can be 6dB; the compression threshold of the second parameter group (mid frequency band) can be -9dBFS, the compression ratio can be 3:1, the start-up time can be 20ms, the release time can be 200ms, and the knee width can be 4dB; the compression threshold of the third parameter group (high frequency band) can be -14dBFS, the compression ratio can be 8:1, the start-up time can be 5ms, the release time can be 80ms, and the knee width can be 2dB. Typical DRC parameter configurations for each frequency band are shown in Table 1 below. Table 1 Typical DRC Parameter Configuration Table for Each Frequency Band With the above configuration, the high-frequency band adopts an aggressive configuration of low threshold, high compression ratio, and fast start-up to accurately suppress high-frequency wind noise; the low-frequency band adopts a relaxed configuration of high threshold, low compression ratio, and slow start-up to protect the impact of low-frequency percussion.
[0047] For example, when an audio device (such as a WiiM Bar) plays test content (such as white noise) at 85dB SPL power, the high-frequency signal level is approximately -22dBFS, exceeding the high-frequency compression threshold of -14dBFS by an excess of ΔV. HF The gain is 8dB. The high-frequency band DRC is processed with a compression ratio of 8:1, and the actual gain attenuation is -8×(1-1 / 8)=-7dB. That is, the high-frequency band signal is compressed by about 7dB, and the output level drops to about -29dBFS, effectively suppressing wind noise. At the same time, the low-frequency band level is about -48dBFS, which is far below the low-frequency band threshold of -6dBFS. The low-frequency band DRC does not intervene, and the low-frequency music content is completely unaffected.
[0048] In one embodiment, the independent dynamic range control process in step S2 further includes: An independent sidechain detection channel is established for each frequency band. The root mean square (RMS) level detection method is used to detect the input signal of the corresponding frequency band. A low-pass filter with a preset time constant is used to smooth the square value of the input signal of the corresponding frequency band to generate the instantaneous RMS level. RMS detection reflects the perceived loudness of the signal better than peak detection, avoiding unnecessary compression triggered by transient peaks at a single sampling point. The specific formula for calculating the instantaneous RMS level is as follows: Where n is the current time, L RMS (n) represents the current instantaneous root mean square level, L RMS (n-1) represents the level at the previous moment, and α is the smoothing coefficient. f s Let τ be the sampling rate, x(n) be the sampled value of the input signal in the current frequency band, and τ be the sampling rate. RMS Set to a preset time constant (e.g., 10ms).
[0049] As an example, and not a limitation, at a sampling rate of 48kHz, τ RMS =10ms, If the input signal x(n) at a certain moment in the high-frequency band is 0.25 (corresponding to -12dBFS), and the RMS value L at the previous moment is... RMS (n-1)=0.18 (corresponding to -14.9dBFS), then at the current time L RMS (n)=sqrt(0.99792×0.0324+0.00208×0.0625)=sqrt(0.03246)≈0.1802 (corresponding to -14.9dBFS), which shows that the RMS value slowly tracks the signal changes and avoids erroneous compression triggered by brief transients.
[0050] Based on instantaneous root mean square level L RMSThe relationship between (n) and the corresponding frequency band compression threshold is used to calculate the target compression gain G. target (n); The target compression gain is smoothed based on the start-up time constant and release time constant of the corresponding frequency band to obtain the smoothed compression gain G. smooth (n), to avoid modulation noise caused by sudden gain changes, specifically including: If the target compression gain G target (n) is less than the smoothing gain G from the previous time step. smooth (n-1) (needs to be compressed), then: G smooth (n)=α attack ×G smooth (n-1)+(1-α attack )×G(n); If the target compression gain G target (n) is greater than or equal to the smoothing gain G from the previous time step. smooth (n-1) (needs to be released), then: G smooth (n)=α release ×G smooth (n-1)+(1-α release )×G target (n); where, , T attack and T release These are the start-up time and release time, respectively. The smoothing compression gain is applied to the corresponding frequency band signal to output the dynamic range control result for the corresponding frequency band.
[0051] As an example, and not a limitation, is the high-frequency DRC start-up time T. attack =5ms, release time T release =80ms, sampling rate 48kHz When a high-frequency signal suddenly exceeds the threshold (G) target When G jumps from 0dB to -7dB smooth Approximately 63% of the transition (approximately -4.4dB) is completed within 5ms (240 sampling points), and approximately 98% of the transition (approximately -6.86dB) is completed within 20ms, achieving rapid response oscillation; while when the signal falls back, G... smooth It releases slowly with an 80ms time constant, effectively avoiding the pumping effect.
[0052] In one embodiment, step S3 involves extracting wind noise characteristics from the high-frequency signal and determining the current wind noise scenario by combining this with device drive strength information, including: The high-frequency signal is subjected to frequency domain transformation (e.g., 128-point FFT), and the spectral flatness index SFM is calculated based on the ratio of the geometric mean to the arithmetic mean of the spectral amplitude. The specific calculation formula is as follows: SFM=exp(Σlog|X(k)| / N) / (Σ|X(k)| / N), where X(k) is the amplitude of the k-th frequency component of the spectrum, and N is the number of FFT points. Wind noise has high spectral flatness, while high-frequency music (such as cymbal sounds and violin overtones) has obvious harmonic structure and spectral peaks, resulting in lower spectral flatness. The SFM range is [0, 1]. SFM approaching 1 indicates that the spectrum is close to white noise (characteristic of wind noise), and SFM approaching 0 indicates that the spectrum has obvious peaks (characteristic of music signals).
[0053] To accurately distinguish between real wind noise and white noise effects (such as rain sounds in movies, ocean waves, ASMR, etc.), this embodiment further extracts the logarithmic power spectrum slope of the high-frequency band based on the spectrum flatness (SFM) to construct the Fused Wind Noise Index (FWNI), which serves as the main basis for determining wind noise scenes. At the same time, it can be combined with equipment drive strength information for auxiliary judgment.
[0054] The first step is to estimate the slope of the high-frequency logarithmic power spectrum. This involves performing a frequency domain transformation (e.g., 128-point FFT) on the high-frequency signal to obtain the complex amplitude X(k) at each frequency point, and then calculating the logarithmic power. Where k is the index of the k-th FFT frequency point in the high-frequency band, k∈Ω HF (Ω) HF (This represents the set of FFT frequency points covered by the high-frequency band), where ε is a zero-prevention parameter, taken as 10. -12 X(k) is the complex amplitude of the FFT at the k-th frequency point in the HF band.
[0055] Least squares linear regression was performed on the logarithmic power P(k) at all frequency points within the high-frequency band (3kHz-20kHz) to estimate the slope of the logarithmic power spectrum. : Where N = |Ω HF | represents the total number of frequency points within the high-frequency band. It reflects the average decay rate of the power spectrum in the high-frequency band as the frequency increases, and is the core characteristic quantity that distinguishes wind noise (turbulent broadband noise) from music content (harmonic structure signal).
[0056] The second step involves calculating the wind noise index (FWNI) by weighting and fusing the spectral flatness (SFM) with the slope matching similarity: , where β wind Let σ be the expected value of the theoretical slope for wind noise, taken as -1.7; βThe slope matching window width parameter is set to 0.5 to control the tolerance range of the Gaussian similarity function; SFM is the high-frequency band spectral flatness, the ratio of geometric mean to arithmetic mean; w1 and w2 are the fusion weights of the two features, with w1=0.55 and w2=0.45 respectively; FWNI is the fused wind noise index, ranging from [0, 1]. The Gaussian similarity term exp(-(β) est -β wind ) 2 / (2σ β 2 Using the wind noise theory slope β wind Centered on the slope, the output approaches 1 when the measured slope matches the theoretical slope of the turbulence noise; otherwise, it rapidly decays to 0.
[0057] The third step is wind noise scene determination. After calculating the FWNI, it is compared with the preset wind noise scene determination threshold Th. FWNI =0.60 for comparison: When FWNI>0.60, it is determined to be a high wind noise scene, triggering the lowering of the high-frequency band DRC compression threshold to enhance wind noise suppression; When FWNI≤0.60, it is determined to be a music-dominated scene, and the high-frequency DRC compression threshold is restored to the default value.
[0058] This embodiment can also combine the device drive strength information to assist in correcting the judgment result: based on the rated power, digital volume control value and the current root mean square signal level, the current drive power is estimated, a wind noise intensity estimate is generated, and the estimate is weighted and fused with FWNI (for example, FWNI accounts for 60% and power estimate accounts for 40%) to obtain the final compression threshold adjustment amount.
[0059] Fourthly, to verify the effectiveness of the FWNI algorithm, the following are computational examples for three typical scenarios: Scenario A (Real wind noise, correct protection trigger): The device plays at a high volume of 85dB SPL; the power at each point in the HF band FFT is flat and decreases gradually with frequency (aerodynamic turbulence characteristics). Linear regression yields β. est =-1.65, SFM=0.90 (high flatness). Substituting into FWNI: exp(-(-1.65-(-1.7)) 2 / (2×0.25))=exp(-0.005)≈0.995,FWNI=0.55×0.90+0.45×0.995=0.495+0.448=0.943,Conclusion:FWNI=0.943>0.60,It is determined to be a high wind noise scene, and DRC protection is triggered.
[0060] Scenario B (white noise sound effect, avoiding false positives and not triggering DRC): White noise (rain sound) is played in the movie / TV sound effects. The power spectrum is extremely flat across the entire frequency band, with a slope close to zero. Linear regression yields β. est =+0.02 (nearly horizontal), SFM=0.91 (similarly high flatness). Substituting into FWNI: exp(-(0.02-(-1.7)) 2 / (2×0.25))=exp(-5.92)≈0.003, FWNI=0.55×0.91+0.45×0.003=0.501+0.001=0.502, Conclusion: FWNI=0.502≤0.60, determined to be music-driven, DRC not triggered. Successfully avoided misjudgment (pure SFM solution would misjudge).
[0061] Scenario C (mixture of cymbal sound and slight wind noise, correctly distinguishable): The sound is predominantly cymbal sound with a small amount of wind noise; harmonics cause the slope to be steeper. Linear regression yields β. est =-2.40, SFM=0.62 (higher due to slight wind noise), substituting into FWNI: exp(-(-2.40-(-1.7)) 2 / (2×0.25))=exp(-0.98)≈0.375, FWNI=0.55×0.62+0.45×0.375=0.341+0.169=0.510. Conclusion: FWNI=0.510≤0.60, maintaining the default parameters, the high-frequency details of the cymbals are fully preserved.
[0062] This embodiment has the following significant advantages compared to the solution using only SFM: Elimination of white noise sound effect miscompression: Pure SFM of white noise sound effects (rain, waves, ASMR) completely overlaps with wind noise, causing all original solutions to falsely trigger DRC. After FWNI introduces slope matching, the β of white noise... est Approaching zero (≠-1.7), the fusion index is below the threshold, completely avoiding miscompression and protecting the high-frequency details and dynamic integrity of this type of content; physical mechanisms enhance interpretability: β wind =-1.7 Directly derived from the turbulent power spectral density ∝ f in the Kolmogorov inertial subregion -5 / 3 Based on fluid dynamics theory and experimental acoustic observations of the Karman vortex street effect in loudspeakers, the algorithm has a clear physical basis, and parameter adjustments are theoretically guided, rather than purely empirical thresholds; dual-feature complementarity significantly improves discrimination accuracy: SFM captures whether the spectrum is uniform, β est Capturing whether "spectral attenuation conforms to turbulence laws" involves positive and negative interactions. The pure SFM scheme has a misclassification rate of approximately 12% (mainly from white noise effects), while introducing FWNI is expected to reduce the misclassification rate to below 2%; the computational cost is negligible: β estThe estimated reuse of existing FFT results only requires additional linear regression (approximately 3N multiplications and additions) at approximately N frequency points (N≈50). The additional overhead per frame on the embedded DSP is no more than 0.02ms, which has no impact on the system's real-time performance.
[0063] In one embodiment, drive strength information characterizing the high-volume output state of the device is obtained, and the current drive power is estimated based on the drive strength information. The current drive power is determined according to the rated power, the digital volume control value, and the current root mean square signal level. The estimation formula is as follows: P est =P rated ×(V digital / 100) 2 ×Level rms 2 , where P rated For rated power, V digital For digital volume control values (0-100), Level rms This represents the current RMS signal level.
[0064] According to the current driving power P est A preset mapping relationship between wind noise intensity and frequency flatness is used to estimate the current wind noise intensity, and the result of the current wind noise scene determination is generated by combining the frequency flatness index. As an example, and not a limitation, is the rated power P of a certain device. rated =40W, the user adjusts the volume to V digital =80, current RMS level rms =0.3 (-10.5 dBFS), estimated drive power P est =40×0.64×0.09=2.304W, corresponding to the wind noise coefficient k wind =1.3, predicted wind noise intensity (Wind) dB =20×log10(2.304 / 1)×1.3≈+9.4dB. Based on this, the system further lowers the high-frequency band DRC compression threshold from the default -14dBFS to -16dBFS, enhancing high-frequency protection in advance and preventing wind noise from exceeding the default threshold during peak periods.
[0065] In one embodiment, step S3, which adaptively updates the dynamic range control parameters corresponding to the high-frequency band signal based on the determination result, includes: Based on the spectral flatness index and / or the estimated wind noise intensity, corresponding threshold adjustment amounts are generated respectively. These threshold adjustment amounts are then fused according to preset weights to obtain a comprehensive adjustment amount for the high-frequency compression threshold. For example, a first threshold adjustment amount is generated based on the spectral flatness index; a second threshold adjustment amount is generated based on the current wind noise intensity estimation result; the first and second threshold adjustment amounts are then fused according to preset weights to obtain a comprehensive adjustment amount for the high-frequency compression threshold. The specific fusion calculation formula is as follows: ΔThreshold HF =w sfm ×ΔSFM dB +w power ×ΔPower dB , where w sfm and w power For weighting coefficients; as an example and not a limitation, w sfm =0.6, w power =0.4.
[0066] The overall adjustment amount is limited to a preset adjustment range, and the high-frequency band compression threshold is smoothly updated according to a preset update cycle; by way of example and not limitation, ΔThreshold HF The range is clamped to [-6dB, 0dB] to avoid overly aggressive protection that could cause unnecessary compression of high-frequency music content. The high-frequency DRC compression threshold is updated smoothly every 50ms.
[0067] In one embodiment, step S4, determining the gain difference between frequency bands, and performing gain coordination processing on the frequency band with smaller compression when the gain difference exceeds a preset coordination threshold, includes: acquiring the gain compression amount corresponding to each frequency band in real time, calculating the gain difference between adjacent frequency bands; when the absolute value of the gain difference between any adjacent frequency band exceeds the preset coordination threshold, identifying the frequency band with smaller gain compression among the adjacent frequency bands as the coordination compensation object; determining the target coordination gain based on the deviation between the gain difference and the preset coordination threshold, so that the gain difference between adjacent frequency bands converges to within the target gain difference range; and injecting the target coordination gain into the coordination compensation object in an exponentially smooth manner within a preset update period, so that the coordination gain smoothly approaches the target coordination gain within a preset transition time, avoiding the introduction of perceptible gain jumps or modulation noise by the coordination action. Specific details are as follows: Real-time acquisition of the gain compression status (G) of low-frequency, mid-frequency, and high-frequency bands. LF G MF G HF ); Based on the gain compression states of the low-frequency, mid-frequency, and high-frequency bands, the first gain difference (ΔG) between the high-frequency and mid-frequency bands is determined respectively.HF-MF =G HF -G MF ), and the second gain difference (ΔG) between the mid-frequency and low-frequency bands. MF-LF =G MF -G LF ); The first gain difference and the second gain difference are respectively compared with a preset coordination threshold (ΔG). max The comparison is performed; when the absolute value of the gain difference between any two adjacent frequency bands exceeds the preset coordination threshold, the frequency band with the smaller gain compression among the adjacent frequency bands is identified as the coordination compensation target, and gain coordination processing is performed on the coordination compensation target to suppress the spectral imbalance caused by independent frequency division compression. As an example, and not a limitation, the preset coordination threshold ΔG... max The default value is 6dB; in high-volume, high-wind-noise scenarios, if the high-frequency DRC produces a -9dB gain compression (G HF =-9dB), mid-frequency DRC produces -1dB compression (G MF =-1dB), then the first gain difference ΔG HF-MF =-8dB, exceeding the 6dB threshold, the system detected an excessive gain difference between frequency bands and triggered mid-band compensation: injecting an additional -2dB of coordination gain into the mid-band, so that ΔG HF-MF Narrowed to -6dB to prevent a muffled sound caused by a severe imbalance in the mid-to-high frequency spectrum.
[0068] In one embodiment, step S4, performing gain coordination processing on the coordination compensation object, includes: Based on the degree of deviation between the gain difference and the preset coordination threshold, the corresponding target coordination gain (G) is determined. coord_target This brings the gain difference between the adjacent frequency bands closer to the target gain difference range. The target coordination gain is applied to the coordination compensation object to increase the gain compression of that frequency band; Within a preset update period, the smooth coordination gain G at the current moment is determined based on the weighted result of the smooth coordination gain at the previous moment and the current target coordination gain. coord_smooth (t), specifically injected using an exponential smoothing method, calculated as follows: G coord_smooth (t)=(1-λ)×G coord_smooth (t-1)+λ×G coord_target Where λ is the smoothing coefficient, t is the current update period, and G coord_smooth (t-1) represents the smoothing and coordination gain at the previous time step; The smoothing coordination gain is applied to the coordination compensation object so that the coordination gain smoothly approaches the target coordination gain within a preset transition time, thereby avoiding perceptible gain jumps or modulation noise introduced by the coordination action. As an example, and not a limitation, the smoothing coefficient λ = 0.02 (update cycle every 50ms) ensures that the coordination gain transitions from 0 to the target value within approximately 2 seconds, completely transparent to the user. If the target coordination gain is -2dB, the initial smoothing gain is 0dB. In the first update cycle (after 50ms), the smoothing gain is approximately -0.04dB, in the tenth update cycle (after 500ms), it is approximately -0.366dB, and after approximately 100 update cycles, it approaches the target value of -2dB with a smooth and imperceptible transition.
[0069] In one embodiment, the wind noise suppression method further includes full-band linkage protection: When the gain compression in the high-frequency (HF) band exceeds a preset protection threshold, the audio device is determined to be in an extreme wind noise protection scenario. As an example, and not a limitation, the preset protection threshold can be -12dB; when the high-frequency DRC compression exceeds this value, the speaker is deemed to be at risk of overheating and damage.
[0070] In response to the extreme wind noise protection scenario, the dynamic range control parameters corresponding to the low-frequency (LF) band and the mid-frequency (MF) band are adjusted in a coordinated manner to implement full-band protective compression of the low-frequency, mid-frequency, and high-frequency bands, reducing the overall excitation power of the speaker. For example, the DRC thresholds of the low-frequency and mid-frequency bands are lowered by 3dB and 2dB respectively, implementing mild full-band compression. Risk alarm information is sent to the device protection management module, and a volume limiting mechanism is triggered to limit the upper limit of the device's output volume. For example, the volume limiting mechanism can clamp the maximum digital volume to 90% of the current value. Protection prompt information is output to the device display or associated application to guide the user to actively reduce the volume. For example, a prompt message such as "Speaker protection is activated" can be displayed on the device display and associated application to avoid continuous accumulation of wind noise that could damage the speaker unit.
[0071] In one embodiment, the wind noise suppression method further includes recombining the frequency band signals after dynamic range control processing and gain coordination processing into a full-band output signal, and performing real-time quality monitoring on the recombined output signal, specifically including: The time-domain aligned signals from each frequency band (low-frequency, mid-frequency, and high-frequency) after delay compensation are linearly superimposed to obtain the full-band output signal. The specific reconstruction formula is as follows: S out (n)=S LF_out (n)+S MF_out (n)+S HF_out (n), where S LF_out (n), S MF_out (n), S HF_out(n) represent the processed output signals for each frequency band. Due to the use of a cross-frequency divider filter bank (such as LR4), the three signals meet the condition of precise amplitude and phase alignment when superimposed. The frequency response flatness error after superposition does not exceed ±0.2dB (full band), and the phase error does not exceed 2 degrees, ensuring that the frequency response and phase characteristics of the reconstructed signal are highly close to the ideal full-band signal. As an example, and not a limitation, under a single-frequency sine wave (1kHz) test, the 1kHz signal only enters the mid-frequency band after passing through the frequency divider matrix. The amplitude error of the reconstructed output signal compared to the original 1kHz signal is less than 0.01dB, and the phase error is less than 0.5 degrees. Under a wideband signal (pink noise) test, the frequency response fluctuations after frequency division and reconstruction are within ±0.15dB (20Hz-20kHz), meeting the high-fidelity standard.
[0072] Real-time quality monitoring is performed on the full-band output signal, including at least distortion and noise index monitoring and spectrum reconstruction error monitoring. Specifically, total harmonic distortion plus noise (THD+N) monitoring ensures that the gain modulation noise introduced by the frequency band dynamic range control does not exceed -70dBFS (below 0.03%).
[0073] The spectral difference between the full-band output signal and the reference signal in the non-wind noise-dominated frequency band is calculated to obtain the spectral reconstruction error, which characterizes the degree of spectral distortion introduced by the inter-band gain coordination processing in the reconstructed output signal. When the spectrum reconstruction error exceeds a preset error threshold, it is determined that the inter-band coordination gain injection amount exceeds a reasonable range, and feedback triggers the reduction control of the coordination gain. The upper limit of the coordination gain injection amount is dynamically adjusted based on the feedback result of the spectrum reconstruction error, realizing closed-loop adaptive control of the coordination gain to prevent excessive indirect compression of non-wind noise frequency bands due to gain coordination processing. When the spectrum reconstruction error in the non-wind noise dominant frequency band (such as the low-frequency and mid-frequency main frequency range) exceeds a preset error threshold, it is determined that the coordination gain injection is excessive, and feedback triggers the reduction control of the coordination gain to avoid excessive indirect compression of non-wind noise frequency bands. As an example, and not a limitation, the preset error threshold can be set to -30dB. In the high-volume section of a low-frequency-rich EDM song, the low-frequency band DRC produces approximately -3dB gain compression, and the high-frequency band produces -10dB compression. The inter-band gain coordination module triggers a mid-frequency band coordination gain of -4dB, causing the mid-frequency band spectrum reconstruction error to exceed -28dB (exceeding the -30dB threshold). The system automatically limits the mid-frequency band coordination gain to -2dB, and the spectrum reconstruction error drops back to -32dB (meeting the threshold requirement). While protecting the clarity of human voices, the wind noise suppression effect remains effective.
[0074] In one embodiment, the wind noise suppression method further includes dynamic range control parameter self-learning and device-specific calibration, specifically including: When the audio device has accumulated a preset learning time, it performs device-level parameter self-learning optimization based on historical operating data. The historical operating data includes at least statistical data on wind noise characterization features, wind noise intensity estimation data, and spectrum reconstruction error data. Based on the self-learning optimization results, for the measured wind noise characteristics of the current device instance (e.g., the differences in wind noise intensity among different individual speakers), at least one of the following parameters is automatically fine-tuned: high-frequency dynamic range control parameters, inter-band coordination parameters, and wind noise intensity estimation parameters; as an example and not a limitation, the automatic fine-tuning may include: high-frequency compression threshold adjustment range ±2dB, inter-band coordination gain weight adjustment range ±0.1, and wind noise intensity coefficient adjustment range ±0.2.
[0075] The fine-tuned parameters are written to non-volatile memory for loading upon subsequent startup of the audio device, thus achieving device-level wind noise protection adapted to individual differences. As an example, and not a limitation, after a certain audio device has accumulated 32 hours of operation, the system's self-learning analysis reveals that the high-frequency wind noise intensity of this device instance is approximately 2.1 dB higher than the design value (possibly due to smaller tolerances in the speaker tubes of this batch, resulting in stronger airflow turbulence). The self-learning module lowers the high-frequency compression threshold from the default -14 dBFS to -16.1 dBFS and updates the wind noise intensity coefficient from 1.3 to 1.45, writing the updated parameters to Flash memory. Upon the next device startup, the frequency-band DRC system automatically loads these personalized parameters, improving the device's high-frequency wind noise suppression effect by approximately 2 dB, while maintaining consistent mid-to-low frequency dynamics with the parameters before the update, achieving precise device-level protection.
[0076] like Figure 2 As shown, based on the same inventive concept, this embodiment provides an audio device wind noise suppression device, comprising: The frequency division matrix module is used to acquire the input audio signal and perform frequency division processing to separate it into multiple frequency band signals, including at least low-frequency signals and high-frequency signals. A multi-band independent DRC processing module is used to perform independent dynamic range control processing on each frequency band signal based on the independent dynamic range control parameters corresponding to each frequency band. Specifically, for the high-frequency band signal, a lower compression threshold, a higher compression ratio, and a shorter start-up time are configured compared to other frequency band signals, so that the high-frequency band dynamic range control processing prioritizes the suppression of signal peaks in the wind noise-dominant frequency band. The wind noise detection and adaptive threshold update module is used to extract the wind noise characterization features of the high-frequency signal, combine the device drive strength information to determine the current wind noise scene, and adaptively update the dynamic range control parameters corresponding to the high-frequency signal according to the determination result. The inter-band gain coordination module is used to monitor the gain compression status of signals in each frequency band after dynamic range control processing and to determine the gain difference between frequency bands. When the gain difference exceeds a preset coordination threshold, gain coordination processing is performed on the frequency band with smaller compression to reduce the spectral imbalance caused by independent frequency division compression. The frequency band signal reassembly module is used to reassemble the signals of each frequency band after dynamic range control processing and gain coordination processing into a full-band output signal. When the gain compression of the high-frequency band signal exceeds the preset protection threshold, the full-band linkage protection is triggered. The principle is described in Embodiment 1 and will not be repeated here.
[0077] Based on the same inventive concept, this embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the computer program, when executed by the processor, implements the steps of any of the above-described audio device wind noise suppression methods.
[0078] Based on the same inventive concept, this embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the audio device wind noise suppression method described above.
[0079] The computer device in this embodiment will be described in detail below from the perspective of hardware processing.
[0080] See Figure 3 As shown, the computer device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the above-described audio device wind noise suppression method.
[0081] further, Figure 3 The computer device shown also includes a bus 102 and a communication interface 103, with the processor 100, communication interface 103 and memory 101 connected via the bus 102.
[0082] The memory 101 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0083] The processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 100 or by instructions in software form. The processor 100 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information in memory 101 and, in conjunction with its hardware, completes the method steps of the aforementioned embodiments.
[0084] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the audio device wind noise suppression method.
[0085] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for suppressing wind noise in audio equipment, characterized in that, include: The input audio signal is acquired and frequency-divided to separate multiple frequency band signals, including at least low-frequency and high-frequency signals. Based on the independent dynamic range control parameters corresponding to each frequency band, independent dynamic range control processing is performed on the signals of each frequency band. Among them, a lower compression threshold, a higher compression ratio, and a shorter start-up time are configured for the high-frequency band signals compared with other frequency band signals, so that the dynamic range control processing of the high-frequency band prioritizes the suppression of the signal peak of the wind noise-dominant frequency band. Extract the wind noise characterization features of the high-frequency band signal, combine them with the device drive strength information to determine the current wind noise scene, and adaptively update the dynamic range control parameters corresponding to the high-frequency band signal based on the determination result. Monitor the gain compression status of signals in each frequency band after dynamic range control processing, and determine the gain difference between frequency bands. When the gain difference exceeds a preset coordination threshold, perform gain coordination processing on the frequency band with smaller compression. The signals from each frequency band, after dynamic range control and gain coordination processing, are recombined into a full-band output signal.
2. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, The multiple frequency band signals include low-frequency signals, mid-frequency signals, and high-frequency signals, and the frequency division processing includes: The full-band audio signal is divided into two levels according to the preset first and second frequency division points; Wherein, the first frequency division point is used to separate the low-frequency band signal and the high-frequency mixed band signal, and the second frequency division point is used to further separate the high-frequency mixed band signal into the mid-frequency band signal and the high-frequency band signal, so that the high-frequency band covers the frequency band that mainly generates wind noise; The low-frequency signal corresponds to a frequency range of 20Hz to 300Hz, the mid-frequency signal corresponds to a frequency range of 300Hz to 3kHz, and the high-frequency signal corresponds to a frequency range of 3kHz to 20kHz.
3. The method for suppressing wind noise in audio equipment according to claim 2, characterized in that, The frequency division process also includes: The input audio signal is frequency-divided and filtered using a cross-frequency divider filter bank. The cross-frequency divider filter bank matches the amplitude response of adjacent frequency bands at the corresponding frequency division point, and the corresponding high-pass filter path and low-pass filter path have matching phase response, so that the signals of each frequency band can obtain a flat frequency response after recombination, and eliminate the spectral distortion introduced by frequency division.
4. The method for suppressing wind noise in audio equipment according to claim 3, characterized in that, The frequency division process also includes: The group delay introduced by the cross-frequency divider filter bank on each frequency band filtering path is obtained; according to the group delay difference of each frequency band path, a corresponding delay compensation amount is introduced for at least one frequency band signal to make the signals of each frequency band achieve time domain alignment before reassembly, avoid time domain phase vanishing truth caused by group delay difference, and ensure the time domain integrity of the reassembled signal.
5. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, The process of performing independent dynamic range control on signals of each frequency band includes: Each frequency band is configured with an independent dynamic range control parameter group. Each parameter group includes at least compression threshold, compression ratio, start-up time, release time, and knee width. The configuration of the parameter groups for each frequency band ensures that the compression start-up speed and compression intensity of the high-frequency band dynamic range control processor for signals exceeding the threshold are higher than those of the corresponding dynamic range control processors for other frequency bands, while the compression start-up speed and compression intensity for the low-frequency band are the lowest. This is to accurately suppress high-frequency wind noise while preserving the dynamic integrity of low-frequency music content to the greatest extent.
6. The method for suppressing wind noise in audio equipment according to claim 5, characterized in that, The process of performing independent dynamic range control also includes: An independent sidechain detection channel is established for each frequency band. The root mean square level detection method is used to detect the input signal of the corresponding frequency band. A low-pass filter with a preset time constant is used to smooth the square value of the input signal of the corresponding frequency band to generate the instantaneous root mean square level of the corresponding frequency band. The target compression gain is calculated based on the relationship between the instantaneous root mean square level and the corresponding frequency band compression threshold; The target compression gain is smoothed based on the start-up time constant and release time constant of the corresponding frequency band to obtain a smoothed compression gain, and the smoothed compression gain is applied to the signal of the corresponding frequency band.
7. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, The step of extracting the wind noise characterization features of the high-frequency band signal and combining them with the equipment drive strength information to determine the current wind noise scene includes: The high-frequency band signal is subjected to frequency domain transformation, and the spectral flatness index is calculated based on the ratio of the geometric mean to the arithmetic mean of the spectral amplitude; and / or the current driving power is estimated based on the rated power, digital volume control value and the current root mean square signal level, and a wind noise intensity estimate is generated according to the preset mapping relationship between the current driving power and the wind noise intensity. A current wind noise scene determination result is generated based on at least one of the spectrum flatness index and the wind noise intensity estimate; wherein, the higher the spectrum flatness index and the larger the wind noise intensity estimate, the more the wind noise scene determination result tends to be a high wind noise scene.
8. The method for suppressing wind noise in audio equipment according to claim 7, characterized in that, The adaptive updating of the dynamic range control parameters corresponding to the high-frequency band signal based on the determination result includes: Based on the spectral flatness index and / or the wind noise intensity estimate, corresponding threshold adjustment amounts are generated respectively, and the threshold adjustment amounts of each path are fused and calculated according to preset weights to obtain the comprehensive adjustment amount of the high-frequency band compression threshold. The overall adjustment amount is limited to a preset adjustment range, and the high-frequency band compression threshold is smoothly updated according to a preset update cycle to prevent the threshold from being excessively lowered, causing unnecessary compression to normal high-frequency music content.
9. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, The determination of gain differences between frequency bands, and when the gain difference exceeds a preset coordination threshold, performing gain coordination processing on the frequency band with smaller compression, includes: The gain compression amount of each frequency band is acquired in real time, and the gain difference between adjacent frequency bands is calculated. When the absolute value of the gain difference between any adjacent frequency band exceeds the preset coordination threshold, the frequency band with the smaller gain compression amount in the adjacent frequency band is identified as the coordination compensation object. The target coordination gain is determined based on the degree of deviation between the gain difference and the preset coordination threshold, so that the gain difference between adjacent frequency bands converges to the target gain difference range. Within a preset update cycle, the target coordination gain is injected into the coordination compensation object in an exponentially smooth manner, so that the coordination gain smoothly approaches the target coordination gain within a preset transition time, avoiding the introduction of perceptible gain jumps or modulation noise by the coordination action.
10. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, Also includes: When the gain compression in the high-frequency band exceeds the preset protection threshold, the current audio device is determined to be in an extreme wind noise protection scenario. In response to the extreme wind noise protection scenario, the dynamic range control parameters corresponding to the low-frequency and mid-frequency bands are adjusted in a coordinated manner to implement full-band protective compression of the low-frequency, mid-frequency and high-frequency bands, thereby reducing the overall excitation power of the loudspeaker. Send risk alarm information to the device protection management module and trigger the volume limiting mechanism to limit the upper limit of the device's output volume; Output protection prompts to the device display or associated application to guide users to actively lower the volume.
11. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, This also includes real-time quality monitoring of the recombined output signal: The low-frequency, mid-frequency, and high-frequency signals after delay compensation are time-domain aligned and linearly superimposed to obtain the full-band output signal; The spectral difference between the full-band output signal and the reference signal in the non-wind noise-dominated frequency band is calculated to obtain the spectral reconstruction error, which characterizes the degree of spectral distortion introduced by the inter-band gain coordination processing in the reconstructed output signal. When the spectrum reconstruction error exceeds a preset error threshold, it is determined that the inter-band coordination gain injection amount exceeds a reasonable range, and feedback triggers the reduction control of the coordination gain; the upper limit of the coordination gain injection amount is dynamically adjusted according to the feedback result of the spectrum reconstruction error to realize closed-loop adaptive control of the coordination gain, so as to prevent the non-wind noise frequency band from being excessively indirectly compressed due to gain coordination processing.
12. The method for suppressing wind noise in audio equipment according to claim 1, characterized in that, It also includes dynamic range control parameter self-learning and device-specific calibration, including: When the audio device has accumulated a preset learning time, it performs device-level parameter self-learning optimization based on historical operating data. The historical operating data includes at least statistical data on wind noise characterization features, wind noise intensity estimation data, and spectrum reconstruction error data. Based on the self-learning optimization results, for the measured wind noise characteristics of the current equipment instance, at least one of the following parameters is automatically fine-tuned: high-frequency dynamic range control parameters, inter-band coordination parameters, and wind noise intensity estimation parameters. The fine-tuned parameters are written to non-volatile memory for loading when the audio device starts up later, thus achieving device-level wind noise protection that adapts to individual differences.
13. A wind noise suppression device for audio equipment, characterized in that, include: The frequency division matrix module is used to acquire the input audio signal and perform frequency division processing to separate it into multiple frequency band signals, including at least low-frequency signals and high-frequency signals. A multi-band independent DRC processing module is used to perform independent dynamic range control processing on each frequency band signal based on the independent dynamic range control parameters corresponding to each frequency band. Specifically, for the high-frequency band signal, a lower compression threshold, a higher compression ratio, and a shorter start-up time are configured compared to other frequency band signals, so that the high-frequency band dynamic range control processing prioritizes the suppression of signal peaks in the wind noise-dominant frequency band. The wind noise detection and adaptive threshold update module is used to extract the wind noise characterization features of the high-frequency signal, combine the device drive strength information to determine the current wind noise scene, and adaptively update the dynamic range control parameters corresponding to the high-frequency signal according to the determination result. The inter-band gain coordination module is used to monitor the gain compression status of signals in each frequency band after dynamic range control processing and to determine the gain difference between frequency bands. When the gain difference exceeds a preset coordination threshold, gain coordination processing is performed on the frequency band with smaller compression. The frequency band signal reassembly module is used to reassemble the signals of each frequency band after dynamic range control processing and gain coordination processing into a full-band output signal.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for suppressing wind noise in an audio device as described in any one of claims 1-12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for suppressing wind noise in an audio device as described in any one of claims 1-12.