A wind noise resistant adaptive volume adjustment method and system for cycling earphones

By analyzing the external ambient sound signal and the source audio signal in the cycling headphones, frequency-related compensation gain is generated, which solves the problems of sound quality distortion and unnatural listening experience under wind noise in the existing technology, and realizes a high-quality audio playback experience in wind noise environment.

CN121454963BActive Publication Date: 2026-03-27SHENZHEN ASMAX INFINITE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing cycling headphones lack frequency differentiation and audio content feature analysis in wind noise environments, leading to sound quality distortion, unnatural listening experience, and the risk of hearing damage. Their adjustment mechanisms are not refined enough, affecting user experience.

Method used

By acquiring external ambient sound signals and source audio signals in parallel, analyzing noise spectrum characteristics and audio content features, generating frequency-related target compensation gain, and combining it with a psychoacoustic model for dynamic smoothing and weighting, the final dynamic equalizer gain parameters are generated to achieve fine-tuned volume adjustment.

Benefits of technology

It improves audio intelligibility and listening quality in windy environments, optimizes the listening experience of different types of audio, ensures clear perception of key audio information, avoids signal distortion and auditory fatigue, and enhances naturalness and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121454963B_ABST
    Figure CN121454963B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio signal processing, and discloses an anti-wind-noise adaptive volume adjustment method and system for a riding earphone, the method comprising the following steps: S1, collecting ambient sound and source audio signals in parallel; S2, analyzing the ambient sound to determine a macro adjustment intensity; S3, analyzing the source audio in parallel to identify its content type and transient signals; S4, generating frequency-dependent target method sound quality compensation gain based on a psychoacoustic model; S5, dynamically smoothing and frequency-weighting the target gain according to the content and transient characteristics to generate an application gain; S6, fusing the macro adjustment intensity and the application gain to synthesize a final gain parameter; and S7, applying the final gain to the source audio to output the compensated audio. The present application fuses psychoacoustic model and parallel analysis of audio content, realizes frequency-level accurate compensation, solves the problems of traditional degradation and sudden changes in listening experience, and improves audio clarity and comfort.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, in particular to an anti-wind-noise adaptive volume adjustment method and system for cycling earphones. BACKGROUND

[0002] Cycling is a common way of commuting and fitness, and the demand for users to use earphones to listen to music, podcasts or navigation instructions during the process is growing. However, in the cycling scenario, the high-speed airflow passing through the external microphone of the earphone will produce strong wind noise. This wind noise not only seriously interferes with normal audio playback, reduces the intelligibility of voice content and the appreciation experience of music, but also forces users to manually increase the volume, which is not only inconvenient to operate, but also may damage hearing due to excessive volume. Therefore, how to automatically and intelligently adjust the volume of the earphone to resist wind noise interference has become a technical problem to be solved to improve the user experience of cycling earphones.

[0003] Currently, some audio devices adopt a noise adaptive volume control technology. This technology monitors the energy level of the external environmental noise through the microphone on the device, and according to a preset mapping relationship, the detected noise energy is corresponded to a specific gain value, and then this gain value is applied to the source audio signal to be played. When the monitored environmental noise energy increases, the system correspondingly increases the playback volume of the source audio, and vice versa.

[0004] Although the existing technology can adjust the volume according to the noise size to some extent, there are still some deficiencies: first, this technology adopts a strategy of overall gain adjustment of the source audio signal. The fundamental reason is that this technology only regards noise as a single energy value, ignoring the fact that the masking effect of wind noise on human hearing varies significantly at different frequencies. Therefore, the overall volume increase without considering the frequency will excessively amplify the frequency components that are not severely masked by noise, leading to distorted sound quality and unnatural listening experience, and increasing the risk of hearing damage. Secondly, the existing technology lacks the ability to perceive the attributes of the audio content. Because it does not analyze the intrinsic characteristics of the source audio signal, it cannot distinguish whether the current playback is voice or music. This leads to the system's inability to make differentiated compensation, such as when processing voice, it cannot specifically enhance the key frequency band most relevant to intelligibility; when processing music, it may destroy the impact of transient signals such as drum points due to the sluggish gain response, making the music sound dull and lack of dynamics. Finally, the adjustment mechanism of the existing technology is too direct, lacking a macro control strategy that matches the actual working conditions and user perception. Simply relying on real-time noise energy for direct mapping can easily cause frequent and dramatic changes in volume due to transient fluctuations in airflow, disrupting the smoothness and comfort of the listening experience. SUMMARY

[0005] In view of the deficiencies of the prior art, the present application provides an anti-wind noise adaptive volume adjustment method and system for cycling earphones, which solves the problems of audio quality degradation, abrupt listening experience and loss of audio dynamic details caused by the overall gain adjustment without distinguishing frequency and ignoring audio content characteristics when resisting wind noise in the prior art.

[0006] To achieve the above object, the present application is implemented by the following technical solutions:

[0007] The present application provides an anti-wind noise adaptive volume adjustment method for cycling earphones, which comprises the following steps:

[0008] S1: Collecting external environmental sound signals and source audio signals to be played in parallel;

[0009] S2: Analyzing the external environmental sound signals to determine a macro adjustment intensity level;

[0010] S3: In parallel with analyzing the external environmental sound signals, analyzing the intrinsic characteristics of the source audio signals to obtain content type identification results and transient signal detection results;

[0011] S4: Based on the spectral characteristics of the external environmental sound signals, the spectral characteristics of the source audio signals and the psychoacoustic model, generating a frequency-dependent target compensation gain;

[0012] S5: Based on the transient signal detection results and the content type identification results, dynamically smoothing and frequency weighting the target compensation gain to generate an application gain;

[0013] S6: Fusing the macro adjustment intensity level and the application gain to synthesize the final dynamic equalizer gain parameter;

[0014] S7: Applying the final dynamic equalizer gain parameter to the source audio signals and outputting the compensated audio.

[0015] Preferably, the step S2 comprises:

[0016] Applying a voice activity detection algorithm to the external environmental sound signal frame to identify pure noise frames;

[0017] Calculating the real-time energy value of the pure noise frames;

[0018] Comparing the real-time energy value with a preset noise energy interval and adjustment intensity level mapping table to determine the macro adjustment intensity level.

[0019] Preferably, the preset noise energy interval and adjustment intensity level mapping table is established by an offline calibration method, and the offline calibration method comprises:

[0020] Collecting multiple sets of pure wind noise signal samples under different riding speed working conditions;

[0021] Calculating statistical average energy values of the wind noise signal samples under each working condition;

[0022] According to the statistical average energy values, a mapping relationship between energy intervals and the macro-adjustment intensity levels is established to generate the mapping table.

[0023] Preferably, the step S3 comprises:

[0024] By calculating the spectral centroid and zero-crossing rate features of the source audio signal frame and inputting them into a preset classifier model, the audio frame is classified as speech dominant content or music dominant content to obtain the content type recognition result; and

[0025] By monitoring whether the short-time energy change rate of the source audio signal frame exceeds a preset energy rising threshold, it is determined whether the current frame is a transient signal frame to obtain the transient signal detection result.

[0026] Preferably, the step S4 comprises:

[0027] Calculating the power spectral density of the pure noise frame in the external environmental sound signal;

[0028] Based on the power spectral density of the noise and the absolute threshold of human ear, a real-time masking threshold curve is calculated;

[0029] Calculating the power spectral density of the source audio signal frame;

[0030] At each frequency point, the power spectral density of the source audio signal is compared with the real-time masking threshold curve to calculate the target compensation gain required to raise the masked audio frequency component to an audible state.

[0031] Preferably, the dynamic smoothing process in the step S5 comprises:

[0032] Based on the transient signal detection result, a smoothing coefficient is adaptively selected for the recursive smoothing process of the target compensation gain to generate the application gain; wherein when a transient signal is detected, an attack coefficient and a release coefficient with faster response time are used.

[0033] Preferably, the frequency weighting in the step S5 comprises:

[0034] Based on the content type recognition result, a frequency weighting function matched with the current content type is selected to perform frequency weighting on the application gain; wherein when speech dominant content is identified, the frequency weighting function has a larger weight value in the key frequency band of human voice intelligibility.

[0035] Preferably, the step S6 comprises:

[0036] converting the macro-adjustment strength level into a continuous scaling factor;

[0037] combining the continuous scaling factor with the frequency-weighted application gain to calculate the final dynamic equalizer gain parameter.

[0038] Preferably, the step S7 comprises:

[0039] taking square root of the final dynamic equalizer gain parameter and applying it to the complex spectrum of the source audio signal to adjust the amplitude of the complex spectrum;

[0040] performing inverse short-time Fourier transform on the amplitude-adjusted complex spectrum to reconstruct a time-domain signal.

[0041] The second aspect of the present application provides an anti-wind-noise adaptive volume adjustment system for cycling earphones, configured to perform the above method, and the system comprises:

[0042] an audio input module configured to collect an external environmental sound signal and a source audio signal to be played in parallel;

[0043] an environmental analysis module configured to analyze the external environmental sound signal to determine a macro-adjustment strength level;

[0044] a source audio analysis module configured to analyze the inherent characteristics of the source audio signal to obtain a content type identification result and a transient signal detection result in parallel with the environmental analysis module;

[0045] a compensation generation module configured to generate a frequency-dependent target compensation gain based on the spectral characteristics of the external environmental sound signal, the spectral characteristics of the source audio signal, and a psychoacoustic model;

[0046] a compensation modulation module configured to modulate the target compensation gain based on the content type identification result and the transient signal detection result, and fuse the macro-adjustment strength level to synthesize a final dynamic equalizer gain parameter;

[0047] an audio output module configured to apply the final dynamic equalizer gain parameter to the source audio signal and output the compensated audio.

[0048] The present application provides an anti-wind-noise adaptive volume adjustment method and system for cycling earphones, which has the following beneficial effects:

[0049] 1. This invention improves audio intelligibility and listening quality in strong wind noise environments by introducing a compensation mechanism based on a psychoacoustic model. Instead of uniformly amplifying the audio signal, this method calculates the noise masking threshold in real time and generates a frequency-dependent target compensation gain accordingly, selectively boosting only the audio frequency components perceived by the human ear as masked by noise. This refined processing ensures that key audio information is clearly perceived amidst noise while avoiding signal distortion and auditory fatigue caused by unnecessary overall gain, thus achieving an excellent balance between compensation effectiveness and audio fidelity.

[0050] 2. This invention provides a highly adaptive and refined compensation strategy through in-depth analysis of audio content and dynamic characteristics, thereby optimizing the listening experience of different types of audio. The method analyzes the source audio signal in parallel, identifies whether it is speech-dominated or music-dominated, and detects transient signals. This allows for focused weighting of key intelligibility frequency bands in the speech content and the use of a faster gain response speed to process transient signals. This ensures that the compensation highlights the human voice while maintaining the dynamic impact of the music, avoiding the dullness or distortion that occurs when traditional methods process complex audio.

[0051] 3. This invention establishes a mapping table between noise energy and adjustment intensity through offline calibration, and matches the current noise energy to it in real-time processing to determine a macroscopic adjustment intensity level. This level ultimately acts as a continuous scaling factor on the refined compensation parameters, enabling the overall compensation effect to smoothly transition with changes in riding speed. This achieves seamless adjustment from no compensation to strong compensation, improving the naturalness and comfort of the user experience. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0053] Figure 2 This is a schematic diagram of the system structure of the present invention;

[0054] Figure 3 This is a flowchart illustrating the offline calibration and macro-adjustment strategy establishment method of the present invention;

[0055] Figure 4 This is a flowchart illustrating the real-time parallel signal acquisition and preprocessing method of the present invention;

[0056] Figure 5 This is a flowchart illustrating the method for environmental noise analysis and macroscopic regulation intensity determination of the present invention.

[0057] Figure 6 This is a flowchart illustrating the parallel analysis method for source audio signal features of the present invention.

[0058] Figure 7 A flowchart of the method for generating compensation parameters based on a psychoacoustic model according to the present application is shown in FIG. 1.

[0059] Figure 8 A flowchart of the method for synthesizing and outputting the compensated audio according to the present application is shown in FIG. 6.

[0060] Figure 9 A flowchart of the method for synthesizing and outputting the compensated audio according to the present application is shown in FIG. 6. DETAILED DESCRIPTION

[0061] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the specification of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0062] Please refer to the accompanying drawings in the specification of the present application. Figure 1 , Figure 1 A flowchart of the method according to an embodiment of the present application is shown in FIG. 1. The present application provides an anti-wind-noise adaptive volume adjustment method for a riding earphone, which comprises the following steps:

[0063] S1: Collecting an external environmental sound signal and a source audio signal to be played in parallel. The external environmental sound signal is obtained by at least one microphone arranged on the earphone, and contains wind noise and other background noise. The source audio signal is an audio data stream to be played through a speaker of the earphone.

[0064] S2: Analyzing the external environmental sound signal to determine a macro adjustment intensity level. This step specifically comprises: applying a voice activity detection algorithm to identify pure noise frames in the external environmental sound signal; calculating a real-time energy value of the pure noise frames; and comparing the real-time energy value with a preset noise energy-adjustment intensity mapping table to match and determine the macro adjustment intensity level at the current time.

[0065] S3: In parallel with step S2, analyzing the internal characteristics of the source audio signal. This analysis process comprises two parallel sub-steps:

[0066] One is content type identification, which calculates a series of signal characteristics such as spectral centroid and zero-crossing rate of the source audio signal frames, and classifies the current audio content as voice-dominant content or music-dominant content according to these characteristics through a preset classifier model;

[0067] The other is transient signal detection, which monitors the rising rate of the short-time energy of the source audio signal frames, and marks the current frame as a transient signal frame when the rate exceeds a preset threshold.

[0068] S4: generating frequency-dependent target compensation gains based on psychoacoustic model. This step first performs short-time Fourier transform on the pure-noise frames in the external environmental sound signal to obtain its power spectral density. Then, according to the power spectral density and the human auditory characteristics, the real-time masking threshold curve of the human ear in the current noise environment is calculated.

[0069] Meanwhile, the source audio signal is also subjected to short-time Fourier transform to obtain its power spectral density. By comparing the power spectral density of the source audio signal with the real-time masking threshold at each frequency point, the target compensation gain required to lift the audio frequency components masked by the noise to the audible state is calculated.

[0070] S5: modulating the target compensation gains and synthesizing the final compensation parameters. This step combines the analysis results of steps S2 and S3 to refine the target compensation gains generated in step S4. The processing process includes:

[0071] First, according to the transient signal detection result of step S3, different time smoothing coefficients are adaptively selected for the update process of the target compensation gains, and recursive smoothing processing is performed to ensure that the dynamics of the transient signal are not weakened.

[0072] Second, according to the content type identification result, a frequency-dependent weighting function is applied to the smoothed gains to highlight the key frequency bands of specific content.

[0073] Finally, the macro-adjustment intensity level determined in step S2 is converted into a continuous scaling factor, and the final dynamic equalizer gain parameters are calculated by combining the scaling factor with the weighted gains.

[0074] S6: applying the final compensation parameters and outputting the audio. The final dynamic equalizer gain parameters are applied to the complex spectrum of the source audio signal for amplitude adjustment. Then, the inverse short-time Fourier transform is performed on the adjusted complex spectrum to reconstruct it into a time-domain signal. Finally, the reconstructed time-domain signal is transmitted to the earphone speaker for playback.

[0075] Please refer to the attached Figure 2 , Figure 2 is a schematic diagram of the system structure according to an embodiment of the present application. The present application provides an anti-wind noise adaptive volume adjustment system for cycling earphones, which can be configured in a cycling earphone or its associated computing device for executing the method in the aforementioned embodiments. The system can include:

[0076] The audio input module 100 is configured to collect the external environmental sound signal and the source audio signal to be played in parallel. The external environmental sound signal is derived from at least one external microphone, and the source audio signal is derived from an internal audio data stream.

[0077] The environmental analysis module 200 is electrically connected to the audio input module 100 and is configured to analyze the external ambient sound signal to determine a macro adjustment intensity level.

[0078] The source audio analysis module 300 is electrically connected to the audio input module 100 and is configured to analyze the intrinsic characteristics of the source audio signal in parallel with the environmental analysis module 200 to output a content type identification result and a transient signal detection result.

[0079] The compensation generation module 400 is electrically connected to the audio input module 100 and is configured to generate a frequency-dependent target compensation gain based on the spectral characteristics of the external ambient sound signal and the spectral characteristics of the source audio signal and in accordance with a psychoacoustic model.

[0080] The compensation modulation module 500 is electrically connected to the environmental analysis module 200, the source audio analysis module 300, and the compensation generation module 400 and is configured to fuse the macro adjustment intensity level, the content type identification result, and the transient signal detection result to dynamically smooth, content-weight, and intensity-scale the target compensation gain to synthesize a final dynamic equalizer gain parameter.

[0081] The audio output module 600 is electrically connected to the audio input module 100 and the compensation modulation module 500 and is configured to apply the final dynamic equalizer gain parameter to the source audio signal and reconstruct the processed signal to drive a loudspeaker unit to output.

[0082] In one specific embodiment, the environmental analysis module 200 further comprises:

[0083] The voice activity detection unit 210 is configured to identify pure noise frames in the external ambient sound signal.

[0084] The noise energy calculation unit 220 is electrically connected to the voice activity detection unit 210 and is configured to calculate a real-time energy value of the pure noise frames.

[0085] The intensity matching unit 230 is electrically connected to the noise energy calculation unit 220 and is configured to compare the real-time energy value with a preset noise energy-adjustment intensity mapping table to determine the macro adjustment intensity level.

[0086] In one specific embodiment, the source audio analysis module 300 further comprises:

[0087] The content type identification unit 310 is configured to classify the current audio content as voice-dominant content or music-dominant content by calculating signal characteristics such as spectral centroid and zero-crossing rate of the source audio signal frames and in accordance with a preset classifier model and output a content type identification result.

[0088] The transient signal detection unit 320 is configured to output a transient signal detection result by monitoring a short-time energy rising rate of the source audio signal frame, and outputting the transient signal detection result when the rate exceeds a preset threshold.

[0089] In one specific embodiment, the compensation generation module 400 is configured to perform a short-time Fourier transform on the pure noise frame to obtain a noise power spectral density, calculate a real-time masking threshold curve based on the noise power spectral density, perform a short-time Fourier transform on the source audio signal frame to obtain a source signal power spectral density, and calculate the target compensation gain by comparing the source signal power spectral density with the real-time masking threshold.

[0090] In one specific embodiment, the compensation modulation module 500 is configured to adaptively select a time smoothing coefficient for recursive smoothing processing of the target compensation gain according to the transient signal detection result, apply a frequency-dependent weighting function to the smoothed gain according to the content type identification result, convert the macro adjustment intensity level into a continuous scaling factor, and calculate the final dynamic equalizer gain parameter by combining the scaling factor and the weighted gain.

[0091] Those skilled in the art should understand that the system structure diagram shown is one embodiment of the present application, and the module division in the figure is functional. In actual implementation, the above functions can be allocated to different modules for completion, or integrated by one or more modules. The module can be a hardware circuit configured to perform a specific function, or software instructions stored in a memory and executed by a processor. Figure 2 The system structure diagram shown is one embodiment of the present application, and the module division in the figure is functional. In actual implementation, the above functions can be allocated to different modules for completion, or integrated by one or more modules. The module can be a hardware circuit configured to perform a specific function, or software instructions stored in a memory and executed by a processor.

[0092] Please refer to the accompanying drawings Figure 3 , Figure 3 is a flowchart of an offline calibration and macro adjustment strategy establishment method according to one embodiment of the present application. The method provides a basis for decision-making for online operation of the system, and its execution process is completed before product deployment.

[0093] The offline calibration and macro adjustment strategy establishment method, in one specific embodiment, can include the following steps:

[0094] S11: In a controlled or typical riding environment, using a headset configured with the system of the present application, pure wind noise signal samples are collected from an external microphone under a series of preset, gradient conditions. The conditions can be different riding speeds, for example, from 30 kilometers / hour to 130 kilometers / hour, at intervals of 10 kilometers / hour. The pure wind noise signal sample refers to audio data sampled in digital form without human voice or other non-target environmental sound.

[0095] S12: Signal processing is performed on the wind noise signal samples collected under each working condition to calculate the representative average energy value. This step can specifically include: dividing the collected wind noise signal samples into multiple analysis frames in the time domain; for each analysis frame, calculating its signal energy; finally, calculating the statistical average of the energy values of the multiple analysis frames to obtain a stable and robust energy representation under the working condition. The statistical average energy under one working condition The calculation formula of the statistical average energy under one working condition may be:

[0096] ;

[0097] In the formula, is the statistical average energy value of the wind noise signal samples under a specific working condition ; is the working condition index; is the total number of analysis frames used for statistical averaging under the working condition; is the index of the analysis frame; is the length of the frame (number of sample points); is the index of the sample point within the frame; is the amplitude of the sample point in the frame under the working condition ; is the amplitude of the sample point in the frame under the working condition ;

[0098] is the logarithm operation with base 10. S13: Establish the mapping relationship between the noise energy interval and the macro adjustment strength level, and generate the strategy table. This step can specifically include: setting a set of energy thresholds , which divides the continuous energy value range into non-overlapping energy intervals. At the same time, define a set of discrete macro adjustment strength levels

[0099] corresponding to the energy threshold intervals.

[0100] By comparing the statistical average energy value of each working condition calculated in step S12 with the energy threshold, the corresponding energy interval and macro adjustment strength level are determined. Finally, the noise energy interval and adjustment strength level mapping table is generated and stored in the system non-volatile memory for online real-time processing and query and matching.

[0101] Please refer to the attached Figure 4 , Figure 4is a flow chart of a real-time signal parallel acquisition and preprocessing method according to an embodiment of the present application. The method provides high-quality, formatted digital signal input for subsequent parallel analysis.

[0102] The real-time signal parallel acquisition and preprocessing method, in one specific embodiment, can include the following steps:

[0103] S21: Continuously acquire external environmental sound signals through at least one microphone arranged outside the earphone at a preset sampling rate and bit depth, and convert them into a digital signal stream . The microphone can be an omnidirectional microphone or a microphone array with a specific directivity to capture wind noise and other environmental sounds around the user.

[0104] S22: In parallel, acquire the source audio signal to be played through the internal audio interface to form a digital signal stream . The internal audio interface can be a Bluetooth audio decoder, a USB audio interface, or an audio file reading interface in the memory, to ensure that the acquired source audio signal is synchronized with the signal to be sent to the audio output module.

[0105] S23: Perform synchronized preprocessing on the two digital signal streams and acquired in steps S21 and S22 to generate signal frames suitable for subsequent frequency domain analysis. The preprocessing step can specifically include:

[0106] Frame segmentation: Divide the continuous digital signal stream into analysis frames with a fixed length. The selection of the frame length needs to be balanced between frequency resolution and time resolution. In some embodiments, an overlapping segmentation method can be used, i.e., there is a part of overlapping samples between two adjacent analysis frames to reduce frame boundary effects. The overlap rate can be set to 50% or higher.

[0107] Windowing: Multiply each analysis frame by a window function to smooth the frame boundaries and reduce spectral leakage. The selection of the window function can be a Hamming window, a Hanning window, or a Kaiser window, etc. The windowed signal frame can be represented as:

[0108] ;

[0109] wherein is the amplitude of the th sample in the windowed source audio signal frame; is the amplitude of the amplitude of the sample; value of the window function at the sample position; index of the intra-frame sample, ranging from 0 to ; length of the frame (sample points). The same windowing process is applied to the external environmental sound signal frame .

[0110] After preprocessing, the system will obtain a series of synchronized, windowed external environmental sound signal frames and source audio signal frames, and transmit them to the subsequent environmental noise analysis module and source audio feature analysis module for processing, respectively.

[0111] Please refer to the attached Figure 5 , Figure 5 is a flowchart of an environmental noise analysis and macro-adjustment intensity determination method according to an embodiment of the present application. The method aims to process real-time external environmental sound signals to obtain a macro-adjustment instruction that can represent the current noise intensity.

[0112] The environmental noise analysis and macro-adjustment intensity determination method, in a specific embodiment, can include the following steps:

[0113] S31: Apply a voice activity detection (VAD) algorithm to the preprocessed external environmental sound signal frame to identify the pure noise frame that does not contain human voice. The purpose of this step is to distinguish the target noise such as wind noise from the interfering human voice (e.g., the rider's own speaking voice or the conversation of others) that may exist, to ensure that the subsequent energy calculation is based only on the target noise.

[0114] In an embodiment, the voice activity detection algorithm can be based on the short-time energy and zero-crossing rate features of the signal. In another embodiment, the voice activity detection algorithm can also combine the spectral entropy, spectral flatness and other frequency domain features of the signal for joint decision to improve the accuracy in complex noise environment. The output of this algorithm is a binary flag , when its value is a preset value (e.g., 0), it indicates that the current time frame is determined to be a pure noise frame.

[0115] S32: When step S31 determines that an external environmental sound signal frame is a pure noise frame, calculate the real-time energy value of the pure noise frame and match the macro-adjustment intensity level according to the energy value. This step specifically can include:

[0116] Calculate the real-time energy value of the pure noise frame. Its calculation formula can be:

[0117] ;

[0118] wherein, is the real-time energy value of the time frame ; is the index of the current time frame; is the length (number of sample points) of the frame; is the index of the sample point within the frame; is the amplitude of the pre-processed (framing, windowing) external ambient sound signal at the th sample point of the time frame ; is the operation to convert the power value to the decibel (dB) scale.

[0119] The calculated real-time energy value is compared with the noise energy interval and adjustment intensity level mapping table established in the offline calibration phase.

[0120] According to the energy interval in which the real-time energy value is located, the macro adjustment intensity level corresponding to the current frame is determined and output. The level will be transmitted to the subsequent compensation dynamic characteristic modulation module to control the overall amplitude of the final compensation effect.

[0121] Please refer to the attached Figure 6 , Figure 6 is a flowchart of a source audio signal feature parallel analysis method according to an embodiment of the present application. The method is executed in parallel with the ambient noise analysis, aiming to deeply analyze the intrinsic properties of the audio content to be played, and to provide decision basis for subsequent differentiated and refined compensation.

[0122] The parallel analysis method of the source audio signal feature, in one specific embodiment, can include a content type identification method and a transient signal detection method.

[0123] The content type identification method, in one embodiment, can include the following steps:

[0124] S41: Calculate a series of frequency domain and / or time domain features that can represent the physical properties of the signal for the pre-processed source audio signal frame. The features at least include:

[0125] spectral centroid , which reflects the center of gravity position of the signal spectrum energy, and is usually related to the brightness perception of sound. Its calculation formula is:

[0126] ;

[0127] wherein, is the spectral centroid value of the time frame ; For frequency; For the source audio signal in time frames The results of the short-time Fourier transform (STFT); For the source audio signal at frequency and time frame The power spectral density; This is the lower limit of the calculated frequency range; To calculate the upper limit of the frequency range; This is the index of the current time frame.

[0128] Zero cross rate This characteristic represents the frequency at which the signal waveform crosses zero, and is typically related to the signal's noise characteristics and high-frequency components. Its calculation formula is:

[0129] ;

[0130] In the formula, For time frames The zero crossover rate; For the pre-processed source audio signal in time frames The The amplitude of each sample point; The length of the frame (number of sample points); This is the index of the intra-frame sample; This is a sign function; it outputs 1 when the input is positive, -1 when the input is negative, and 0 when the input is zero. For the pre-processed source audio signal in time frames The The amplitude of each sample point; This is the index of the current time frame.

[0131] In other embodiments, the features may further include well-known audio features in the art, such as spectral entropy, spectral flatness, spectral roll-off point, and Mel-frequency cepstral coefficients (MFCCs), to form a higher-dimensional feature vector, thereby improving the accuracy of classification.

[0132] S42: Input one or more features calculated in step S41 into a pre-defined classifier model to classify the current audio frame as either speech-dominated or music-dominated content. The classifier model can be obtained offline by training on a large number of labeled speech and music samples. This model can be logistic regression, support vector machine (SVM), decision tree, Gaussian mixture model (GMM), or a lightweight neural network. The output of the classifier is a content type flag. , is used to indicate the content attributes of the current frame.

[0133] In another embodiment, the transient signal detection method can be executed in parallel with content type recognition and includes the following steps:

[0134] S43: Real-time monitoring of the short-time energy change rate of the source audio signal frames. This step may specifically include calculating the short-time energy of the current frame. And calculate its energy compared to the previous frame. The difference. Short-time energy. The calculation formula is:

[0135] ;

[0136] In the formula, For time frames The short-term energy value; For the first The square of the amplitude of each sample point represents the instantaneous energy of that sample point; For the pre-processed source audio signal in time frames The The amplitude of each sample point; The length of the frame (number of sample points); This is the index of the current time frame; This is the index of the intra-frame sample.

[0137] S44: Combine the energy difference calculated in step S43 with a preset energy rise threshold. A comparison is made. When the energy difference is greater than a threshold, the condition is met. If the current frame contains a transient signal, then the method determines that the current frame contains a transient signal. The output of this method is a transient signal flag. When its value is a preset value (e.g., 1), it represents the current time frame. The frame was identified as a transient signal. Threshold. It can be a fixed value, or it can be adaptively adjusted based on the results of content type recognition. For example, when the content is identified as speech-dominant, the threshold can be lowered to improve the detection sensitivity of consonant plosives.

[0138] Please see the appendix Figure 7 , Figure 7 This is a flowchart of a method for generating compensation parameters based on a psychoacoustic model according to an embodiment of the present invention. This method is the core computational part of the technical solution of the present invention, aiming to generate a set of ideal compensation parameters that dynamically change over time and are frequency-dependent.

[0139] In a specific embodiment, the method for generating compensation parameters based on a psychoacoustic model may include the following steps:

[0140] S51: Perform short-time Fourier transform (STFT) on the pure noise frame determined by the ambient noise analysis module to obtain its representation in the frequency domain, and calculate the power spectral density of the noise accordingly . Wherein, denotes the frequency, denotes the index of the current time frame.

[0141] S52: Based on the noise power spectral density obtained in step S51 , and in combination with the psychoacoustic model of human ear hearing, calculate the real-time masking threshold curve of the current time frame . The real-time masking threshold curve represents the minimum energy required for the human ear to perceive sound at each frequency point in this noise environment.

[0142] In one embodiment, the calculation of the real-time masking threshold integrates two aspects: one is the absolute hearing threshold inherent to the human ear, which does not change with the noise ; the other is the masking effect caused by the current noise spectrum. The masking effect is modeled by convolving the noise power spectral density with a spread function that describes the spreading effect of masking energy in the frequency domain. The final real-time masking threshold is determined by both, and its calculation formula can be represented as:

[0143] ;

[0144] In the formula, is the real-time masking threshold at frequency and time frame ; is the frequency; is the index of the time frame; is the absolute hearing threshold of the human ear at frequency ; is the integral symbol, representing the accumulation of all contributions at frequency ; is the noise power spectral density at frequency and time frame ; is the integral variable, representing different frequency components of the noise spectrum; is the spread function that describes the spreading effect of masking energy in the frequency domain, representing the contribution of energy at frequency to the masking threshold at frequency .

[0145] S53: Perform short-time Fourier transform on the source audio signal frame synchronized with the current pure noise frame to obtain its power spectral density The power spectral density of the source audio signal is then compared with the real-time masking threshold calculated in step S52 The comparison is made at each corresponding frequency point to calculate the target compensation gain needed to make the masked audio frequency component audible .

[0146] In one specific embodiment, for any frequency , if is lower than , it indicates that the audio component at this frequency is masked by noise and gain needs to be applied. Otherwise, no extra gain needs to be applied. The target compensation gain can be calculated as:

[0147] ;

[0148] wherein is the target compensation gain at frequency and time frame ; is the real-time masking threshold at frequency and time frame ; is the power spectral density of the source audio signal at frequency and time frame ; is a preset masking margin constant greater than zero, which ensures that the compensated signal energy is stably higher than the masking threshold, thus obtaining better perceptual effect.

[0149] The output of this step is a set of refined target gain parameters varying with frequency and time.

[0150] Please refer to the accompanying Figure 8 , Figure 8 is a flow chart of a compensation dynamic characteristic modulation and final parameter synthesis method according to one embodiment of the present application. The method fuses multi-dimensional information obtained through parallel analysis and performs refined processing on the target compensation gain generated in the foregoing steps to synthesize the compensation parameters finally applied to the audio signal.

[0151] The compensation dynamic characteristic modulation and final parameter synthesis method, in one specific embodiment, can include the following steps:

[0152] S61: Based on the transient signal detection result, dynamically smooth the target compensation gain to generate an application gain that varies continuously in time domain and can maintain the audio dynamic characteristics. This step specifically can include:

[0153] Pre-set at least two groups of time smoothing coefficients: one group is attack coefficient for steady-state signal and release coefficient ; the other group is attack coefficient for transient signal and release coefficient . Generally, is greater than , to correspond to faster gain response time.

[0154] According to the value of transient signal flag , and compare the target compensation gain of the current frame with the size relationship of the gain applied in the last frame, to adaptively select the smoothing coefficient .

[0155] The applied gain is updated in a first-order recursive smoothing manner:

[0156] ;

[0157] In the formula, is the applied gain after dynamic smoothing in frequency and time frame ; is the adaptively selected time smoothing coefficient; is the target compensation gain in frequency and time frame ; is the applied gain in frequency and the previous time frame ; is the frequency; is the index of the current time frame.

[0158] The selection logic of smoothing coefficient is:

[0159] If , the gain is in the attack stage, when indicates a transient signal, , otherwise ;

[0160] If , the gain is in the release stage, when indicates a transient signal, , otherwise .

[0161] S62: Based on the content type identification result, adaptively frequency-weight the applied gain generated in step S61. This step can specifically include:

[0162] According to the content type flag a frequency weighting function that matches the current content type is selected from a pre-defined library of functions .

[0163] For example, when the content is indicated as speech dominant, the selected will have large weight values in the critical bands for human voice intelligibility (e.g. 300 Hz to 4 kHz) and small or 1 in other bands. When the content is indicated as music dominant, the selected may be a function with weight values of 1 in all bands to preserve the spectral balance of music.

[0164] The weighting function is applied to the application gain to obtain a weighted application gain.

[0165] S63: Combine the macro adjustment strength level with the dynamic equalizer gain parameter to obtain the final dynamic equalizer gain parameter . This step can include:

[0166] The macro adjustment strength level output by the ambient noise analysis module is converted to a continuous scaling factor ranging between 0 and 1 by a pre-defined mapping function . The function maps the lowest strength level to 0 or a value close to 0 and the highest strength level to 1.

[0167] The scaling factor, the weighted application gain from step S62, and a reference gain (value 1 representing no change) are combined to calculate the final dynamic equalizer gain parameter. The combination formula can be:

[0168] ;

[0169] wherein is the final dynamic equalizer gain parameter at frequency and time frame ; is the function that maps the macro adjustment strength level to the continuous scaling factor; is the macro adjustment strength level at time frame ; is the application gain that has been dynamically smoothed at frequency and time frame ; is the frequency weighting function selected according to the current content type; is the frequency; is the index of the current time frame.

[0170] The formula ensures that when the macro-adjustment strength is the lowest, close to 0, the final close to 1, i.e. almost no compensation. When the macro-adjustment strength increases, and the fine compensation effect determined by

[0171] Please refer to the attached Figure 9 , Figure 9 is a flowchart of a compensation method for synthesizing and outputting audio according to an embodiment of the present application. The method is responsible for applying the final compensation parameters generated in the foregoing steps to the source audio signal, and generating the final audio output for the user to listen to.

[0172] The compensation method for synthesizing and outputting audio, in a specific embodiment, can include the following steps:

[0173] S71: Apply the final dynamic equalizer gain parameter set to the complex spectrum of the source audio signal to adjust its amplitude. Since the gain parameters are calculated for the power or energy of the signal, and the adjustment of the complex spectrum is for its amplitude, the gain parameters need to be square-rooted before being multiplied by the complex spectrum. The calculation formula of this process can be:

[0174] ;

[0175] In the formula, is the amplitude-adjusted complex spectrum; is the original complex spectrum of the source audio signal in the time frame ; is the final dynamic equalizer gain parameter calculated in the frame; is the frequency; is the index of the time frame. This step accurately adjusts the energy of each frequency component while not changing the phase of the original signal.

[0176] S72: Perform inverse short-time Fourier transform (ISTFT) on the series of amplitude-adjusted complex spectrum frames in step S71 to reconstruct them from the frequency domain to the time domain signal. In an embodiment, when the signal preprocessing stage adopts overlapping segmentation, this step correspondingly adopts the overlap-add method to smoothly splice the processed time domain frames to generate a continuous digital time domain output signal without frame boundary mutation.

[0177] S73: Output the digital time domain signal reconstructed in step S72 to the audio playback unit. Inside the audio playback unit, Digital-to-Analog Conversion and power amplification are performed in sequence.

[0178] Finally, the amplified analog audio signal is sent to the speaker transducer of the earphone to produce the sound that the user finally hears, which has been adjusted in volume by the wind noise resistance adaptive volume adjustment.

[0179] While the embodiments of the application have been illustrated and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the spirit and scope of the application, which is defined by the appended claims and their equivalents.

Claims

1. A method for wind noise-resistant adaptive volume adjustment for cycling headphones, characterized in that, The method includes the following steps: S1. Parallel acquisition of external ambient sound signals and source audio signals to be played; S2. Analyze the external ambient sound signal to determine a macroscopic adjustment intensity level; S3. In parallel with the analysis of external environmental sound signals, analyze the intrinsic characteristics of the source audio signal to obtain content type identification results and transient signal detection results; S4. Based on the spectral characteristics of the external ambient sound signal, the spectral characteristics of the source audio signal, and the psychoacoustic model, generate frequency-related target compensation gain. S5. Based on the transient signal detection result and the content type identification result, the target compensation gain is dynamically smoothed and frequency-weighted to generate an application gain; S6. Integrate the macroscopic adjustment intensity level with the application gain to synthesize the final dynamic equalizer gain parameters; S7. Apply the final dynamic equalizer gain parameters to the source audio signal and output the compensated audio. Step S4 includes: Calculate the power spectral density of pure noise frames in the external ambient sound signal; Based on the power spectral density of the noise and the absolute hearing threshold of the human ear, the real-time masking threshold curve is calculated. Calculate the power spectral density of the source audio signal frame; At each frequency point, the power spectral density of the source audio signal is compared with the real-time masking threshold curve to calculate the target compensation gain required to raise the masked audio frequency components to an audible state.

2. The wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 1, characterized in that, Step S2 includes: A speech activity detection algorithm is applied to external ambient sound signal frames to identify pure noise frames; Calculate the real-time energy value of the pure noise frame; The real-time energy value is compared with a preset noise energy range and regulation intensity level mapping table to determine the macro-regulation intensity level.

3. The wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 2, characterized in that, The preset noise energy range and adjustment intensity level mapping table is established through an offline calibration method, which includes: Multiple sets of pure wind noise signal samples were collected under different cycling speed conditions; Calculate the statistical average energy value of the wind noise signal samples under each operating condition; Based on the statistical average energy value, a mapping relationship is established between the energy range and the macro-regulation intensity level to generate the mapping table.

4. The wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 1, characterized in that, Step S3 includes: By calculating the spectral centroid and zero crossover rate features of the source audio signal frame and inputting them into a preset classifier model, the audio frame is classified into speech-dominated content or music-dominated content to obtain the content type recognition result; and By monitoring whether the short-term energy change rate of the source audio signal frame exceeds a preset energy rise threshold, it is determined whether the current frame is a transient signal frame, so as to obtain the transient signal detection result.

5. The wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 1, characterized in that, The dynamic smoothing process in step S5 includes: Based on the transient signal detection results, a smoothing coefficient is adaptively selected for recursive smoothing processing in the update process of the target compensation gain to generate the application gain; wherein, when a transient signal is detected, an attack coefficient and a release coefficient with a faster response time are used.

6. A wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 5, characterized in that, The frequency weighting in step S5 includes: Based on the content type identification result, a frequency weighting function that matches the current content type is selected, and the application gain is frequency-weighted.

7. A wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 6, characterized in that, Step S6 includes: Convert the macro-adjustment intensity level into a continuous scaling factor; The final dynamic equalizer gain parameter is calculated by combining the continuous scaling factor with the frequency-weighted application gain.

8. A wind noise-resistant adaptive volume adjustment method for cycling headphones according to claim 1, characterized in that, Step S7 includes: The square root of the final dynamic equalizer gain parameter is taken and applied to the complex spectrum of the source audio signal to adjust the amplitude of the complex spectrum. Perform an inverse short-time Fourier transform on the complex spectrum after amplitude adjustment to reconstruct it into a time-domain signal.

9. A wind noise-resistant adaptive volume adjustment system for cycling headphones, characterized in that, The system is used to perform the wind noise-resistant adaptive volume adjustment method for cycling headphones according to any one of claims 1-8, the system comprising: The audio input module is configured to acquire external ambient sound signals and the source audio signal to be played in parallel; The environmental analysis module is configured to analyze the external ambient sound signal to determine a macroscopic modulation intensity level; The source audio analysis module is configured to analyze the intrinsic characteristics of the source audio signal in parallel with the environment analysis module to obtain content type identification results and transient signal detection results; The compensation generation module is configured to generate a frequency-dependent target compensation gain based on the spectral characteristics of the external ambient sound signal, the spectral characteristics of the source audio signal, and a psychoacoustic model. The compensation modulation module is configured to modulate the target compensation gain based on the content type recognition result and the transient signal detection result, and to fuse the macroscopic adjustment intensity level to synthesize the final dynamic equalizer gain parameters. An audio output module is configured to apply the final dynamic equalizer gain parameters to the source audio signal and output compensated audio.

Citation Information

Patent Citations

  • DRC control method for improving indoor riding reducibility of motorcycle riding earphone

    CN121037739A

  • KR1017303860000B1