A DRC control method for improving indoor riding restoration of motorcycle riding earphones

By using a hierarchical adaptive control architecture, differentiated audio processing parameters are generated in real time, which solves the contradiction between voice clarity and sound fidelity in motorcycle riding headphones under high and low noise environments. This achieves stable audio output in different environments and improves the robustness and naturalness of the system.

CN121037739BActive Publication Date: 2026-02-03SHENZHEN ASMAX INFINITE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511568597.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-03
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

The existing DRC control method of motorcycle riding headphones cannot simultaneously meet the requirements of speech clarity and sound fidelity in both high-noise and low-noise environments, resulting in excessive sound compression in quiet environments and an unnatural listening experience.

Method used

A hierarchical adaptive control architecture is adopted. Audio signals are acquired through a microphone and divided into main signal paths and control signal paths. By using a scene judgment module, a joint acoustic state encoder, a parameter synthesizer, and an artifact monitor, differentiated audio processing parameters are generated in real time to achieve dynamic range and equalization processing, thus forming a closed-loop control.

Benefits of technology

A balance between speech clarity and sound fidelity was achieved under different acoustic environments, avoiding hard switching and improving the robustness of the system and the coherence and naturalness of the output audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037739B_ABST
    Figure CN121037739B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio signal processing, and discloses a DRC control method for improving indoor riding restoration of a motorcycle riding earphone, which comprises the following steps: judging a macro acoustic scene; analyzing a micro acoustic state; generating audio processing parameters according to the scene, the state and a correction vector of feedback, and applying the audio processing parameters to an audio signal. The method further comprises the following steps: analyzing an output signal by using an artifact monitor to detect processing artifacts, and generating the correction vector feedback to a parameter synthesizer to form a closed-loop control. The system comprises a scene judgment module, a joint acoustic state encoder, a parameter synthesizer, a parameterized audio processing engine and an artifact monitor. The application combines macro scene judgment with micro closed-loop self-adaption, solves the contradiction between the noise reduction performance under high noise and the sound high restoration under low noise, and improves the audio quality under different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, specifically to a DRC control method for improving the fidelity of indoor riding in motorcycle headphones. Background Technology

[0002] A core function of motorcycle riding headphones is to ensure clear voice communication between users in noisy environments such as high-speed riding. To achieve this, existing audio processing systems typically employ Dynamic Range Control (DRC) technology to process the noise-reduced signal, enhancing human voices and further suppressing residual background noise.

[0003] However, the usage scenarios for motorcycle riding headphones are not static. Besides using them in noisy environments like high-speed riding, users may also make calls in quiet environments such as indoors. These two scenarios present significantly different, even contradictory, audio processing needs. At high speeds, environmental noise is extremely high, requiring users to speak louder; in this case, the primary goal of audio processing is to ensure speech intelligibility. But in quiet indoor environments, where environmental noise is low, users speak at a normal volume; in this case, the user's requirements for natural sound and fidelity become the primary considerations.

[0004] Existing DRC control strategies typically employ a fixed set of parameters designed to handle the most demanding cycling scenarios. This powerful compression strategy, optimized for high-noise environments, can unnecessarily overprocess normal speech signals when directly applied to quiet environments such as indoors. This results in excessive compression of sound dynamics, a flat and unnatural sound, and severely compromises sound fidelity. Therefore, existing technologies generally suffer from the inability to adaptively adjust processing strategies according to changes in the acoustic environment, making it difficult to simultaneously meet users' dual demands for speech clarity and sound fidelity in different scenarios.

[0005] Therefore, this invention proposes a DRC control method to improve the indoor riding fidelity of motorcycle riding headphones, in order to overcome the shortcomings of the prior art. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a DRC control method for improving the sound fidelity of motorcycle riding headphones in indoor riding environments, solving the problem that motorcycle riding headphone audio processing systems cannot simultaneously achieve both high speech clarity in high-noise environments and high sound fidelity in low-noise environments.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a DRC control method for improving the indoor riding fidelity of motorcycle riding headphones, the method comprising the following steps:

[0008] S1. Acquire the audio signal collected by the microphone and distribute the audio signal to the main signal path and the control signal path;

[0009] S2. Analyze the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generate a scene flag;

[0010] S3. In the control signal path, the audio signal is analyzed by a joint acoustic state encoder to generate an acoustic state vector describing the current acoustic environment and human voice state.

[0011] S4. A set of audio processing parameters is generated in real time using a parametric synthesizer based on the scene flags, acoustic state vector, and a feedback correction vector.

[0012] S5. In the main signal path, a parameterized audio processing engine is used to perform dynamic range and equalization processing on the audio signal using the audio processing parameters to generate the processed audio signal as the final output.

[0013] S6. The processed audio signal is analyzed by an artifact monitor to detect audio processing artifacts introduced by the dynamic range and equalization processing, and the correction vector is generated and fed back to the parametric synthesizer to form a closed-loop control.

[0014] Preferably, in step S1, the step of acquiring the audio signal collected by the microphone and distributing the audio signal to the main signal path and the control signal path includes:

[0015] The audio signal is input into a weak noise reduction module for preprocessing to generate a preprocessed audio signal, thereby improving the robustness of subsequent acoustic feature extraction.

[0016] The control signal path receives the preprocessed audio signal.

[0017] Preferably, in step S2, the step of analyzing the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generating a scene marker includes:

[0018] The energy of the non-human voice portion of the audio signal is statistically analyzed to obtain the environmental noise level.

[0019] The ambient noise level is compared with a preset scene judgment threshold to generate a scene flag, which is used to indicate whether the current scene is a high noise scene or a low noise scene.

[0020] Preferably, in step S3, the step of analyzing the audio signal through a joint acoustic state encoder in the control signal path includes:

[0021] The audio signal in the control signal path is segmented into short audio frames, and the acoustic features of each short audio frame are extracted to obtain a feature vector.

[0022] The feature vector is input into the joint acoustic state encoder, and the joint acoustic state encoder outputs the acoustic state vector. The acoustic state vector includes: a speech probability characterizing the possibility that the current frame is human speech, a voice intensity level quantifying the degree of effort required for the user's voice, an estimated signal-to-noise ratio characterizing the relative intensity of the voice signal and the background noise, and an environmental noise feature subvector describing the acoustic characteristics of the background noise.

[0023] Preferably, step S4, which involves generating a set of audio processing parameters in real time using a parametric synthesizer based on the acoustic state vector and a feedback correction vector, includes:

[0024] The parameter synthesizer fuses the acoustic state vector and the correction vector to generate the audio processing parameters, and performs temporal smoothing filtering on the generated audio processing parameters to ensure that the audio processing parameters change continuously and smoothly between different audio frames.

[0025] Preferably, the audio processing parameters include:

[0026] DRC parameters are used to control dynamic range compression, and the DRC parameters include compression threshold, compression ratio, start time, and release time;

[0027] Equalizer gain vector used to adjust specific frequency bands.

[0028] Preferably, the temporal smoothing filtering is implemented using a first-order low-pass filter, and the final audio processing parameters applied to the current frame are calculated according to the following formula:

[0029] ;

[0030] In the formula, These are the audio processing parameters that will be applied to the current frame. These are the raw audio processing parameters generated by the parametric synthesizer for the current frame. These are the audio processing parameters that were finally applied in the previous frame. This is the smoothing coefficient.

[0031] Preferably, in step S5, the step of performing dynamic range and equalization processing on the audio signal using the audio processing parameters through a parameterized audio processing engine in the main signal path includes:

[0032] The audio signal is input into a main noise reduction module for processing to obtain a noise-reduced audio signal.

[0033] The dynamic range and equalization processing are applied to the noise-reduced audio signal.

[0034] Preferably, in step S6, the step of analyzing the processed audio signal using an artifact monitor to detect audio processing artifacts introduced by the dynamic range and equalization processing includes:

[0035] The artifact monitor analyzes the processed audio signal to detect noise sucking or over-compression distortion introduced by the dynamic range and equalization processing, and generates the correction vector based on the detection results.

[0036] The present invention also provides a DRC control system for improving the indoor riding fidelity of motorcycle riding headphones, the system comprising:

[0037] The audio acquisition module is used to acquire audio signals captured by the microphone and distribute the audio signals to the main signal path and the control signal path;

[0038] The scene determination module is used to analyze the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generate a scene flag.

[0039] A joint acoustic state encoder, connected to the control signal path, is used to analyze the audio signals in the control signal path to generate an acoustic state vector describing the current acoustic environment and the state of human voice.

[0040] A parameter synthesizer, connected to the scene judgment module, the joint acoustic state encoder, and an artifact monitor, is used to generate a set of audio processing parameters in real time based on the scene flag, the acoustic state vector, and a feedback correction vector.

[0041] A parametric audio processing engine connects the main signal path and the parametric synthesizer, and is used to perform dynamic range and equalization processing on the audio signal in the main signal path using the audio processing parameters to generate a processed audio signal as the final output.

[0042] An artifact monitor, connected to the output of the parametric audio processing engine, is used to analyze the processed audio signal to detect audio processing artifacts introduced by the dynamic range and equalization processing, and generate the correction vector to feed back to the parametric synthesizer, thereby forming a closed-loop control.

[0043] This invention provides a DRC control method for improving the indoor riding fidelity of motorcycle riding headphones. It has the following beneficial effects:

[0044] 1. This invention, through pre-judgment of the macroscopic acoustic scene, enables the dynamic range control system to select differentiated basic processing strategies between high-noise cycling environments and low-noise indoor environments. This design avoids adopting a single, compromise control scheme, ensuring that effective noise suppression and voice enhancement are prioritized in cycling scenarios, while focusing on maintaining the original dynamics and details of the sound in indoor scenarios. This effectively addresses the conflicting needs for noise reduction performance and sound reproduction in different scenarios.

[0045] 2. This invention generates an acoustic state vector that includes multiple dimensions such as speech probability, vocal intensity, and noise characteristics, providing the control system with a refined and comprehensive perception capability of the current audio signal. Compared to relying solely on simple speech activity detection, this multi-dimensional state description enables subsequent dynamic range control to act more precisely on the target voice, while making more reasonable adaptive adjustments to the background sound environment, significantly improving the clarity and intelligibility of the voice and enhancing the overall listening experience.

[0046] 3. This invention employs a parametric synthesizer to generate audio processing parameters in real time and continuously based on the macroscopic scene, microscopic acoustic state, and feedback correction information. This generative control method endows the system with a high degree of adaptability, enabling it to smoothly respond to continuous changes in the acoustic environment, rather than making hard switches between a few preset modes. Therefore, even in complex scenarios where ambient noise and user speaking volume fluctuate constantly, the system can provide stable and appropriate dynamic range control, ensuring the consistency of the output audio.

[0047] 4. This invention introduces a closed-loop control mechanism including an artifact monitor. Through real-time analysis of the output signal, it can actively identify and suppress audio processing artifacts such as noise sucking and over-compression caused by improper dynamic range processing. This self-correcting capability ensures that even with significant dynamic adjustments, the output audio quality remains pure and natural, fundamentally improving the system's robustness and the fidelity of the final output audio. Attached Figure Description

[0048] Figure 1 This is a functional module structure block diagram of the DRC control system of the present invention;

[0049] Figure 2 This is a flowchart of the DRC control method of the present invention;

[0050] Figure 3 This is a schematic diagram showing the relationship between the input and output levels of the dynamic range compression function of the present invention;

[0051] Figure 4 This is a schematic diagram of the hierarchical control logic of the parameter synthesizer of the present invention.

[0052] Among them, 10 is the audio acquisition module; 20 is the weak noise reduction module; 30 is the scene judgment module; 40 is the joint acoustic state encoder; 50 is the parametric synthesizer; 60 is the parametric audio processing engine; 70 is the main noise reduction module; and 80 is the artifact monitor. Detailed Implementation

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] Reference Figure 1 and Figure 2 This invention provides a DRC control method for improving the indoor riding fidelity of motorcycle riding headphones. In one specific embodiment, the core idea of ​​this method is to construct a hierarchical adaptive control architecture, which divides the control logic of audio signal processing into two levels: macroscopic acoustic scene judgment and microscopic acoustic state closed-loop adaptation. The method may include the following steps:

[0055] In step S1, an audio acquisition module 10 acquires the audio signal acquired by the microphone and copies and distributes the audio signal to two parallel processing paths: a main signal path P1 and a control signal path P2.

[0056] In step S2, a scene determination module 30 analyzes the audio signal to determine the current macroscopic acoustic scene. This module analyzes the energy in non-speech segments of the signal to distinguish between a high-noise cycling scene and a low-noise indoor scene, and generates a scene flag accordingly. This scene flag serves as a macroscopic control command to guide the selection of subsequent processing strategies.

[0057] In step S3, within the control signal path P2, a joint acoustic state encoder 40 performs a refined acoustic state analysis of the audio signal. To improve the accuracy of the analysis, in some embodiments, the audio signal may be preprocessed by a weak noise reduction module 20 before entering the joint acoustic state encoder 40. The joint acoustic state encoder 40 outputs an acoustic state vector describing the current acoustic environment and the state of the human voice, which provides detailed microscopic information for subsequent parameter generation.

[0058] In step S4, a parametric synthesizer 50 generates a set of audio processing parameters in real time based on the scene markers, acoustic state vectors, and a correction vector fed back from subsequent steps, all generated in the preceding steps. This module is the core of the control system, and its internal hierarchical control logic is as follows: Figure 4 As shown, a basic parameter strategy is first selected based on macroscopic scene indicators, and then continuous and fine adjustments are made based on microscopic acoustic state vectors and correction vectors.

[0059] In step S5, in the main signal path P1, the audio signal is first processed by a main noise reduction module 70. Subsequently, a parametric audio processing engine 60 loads the set of audio processing parameters generated by the parametric synthesizer 50, and performs dynamic range and equalization processing on the noise-reduced audio signal to generate the processed audio signal as the final output.

[0060] In step S6, an artifact monitor 80 analyzes the processed audio signal output by the parametric audio processing engine 60 to detect audio processing artifacts introduced during dynamic range and equalization processing. Based on the detection results, this module generates the correction vector and feeds it back to the input of the parametric synthesizer 50, thus forming a closed-loop control circuit to monitor and correct the processing effect of the entire system in real time.

[0061] The following will refer to the appendix. Figure 1 To be continued Figure 4 The specific steps and modules of the methods and systems disclosed in the embodiments of the present invention will be described in detail.

[0062] In step S1, refer to Figure 1 The audio acquisition module 10 is responsible for acquiring raw audio signals from one or more microphones. After being digitized, the audio signal is copied into two identical signal streams within the system. The first signal stream is directed to the main signal path P1, which is then used for subsequent main noise reduction and dynamic range processing. The second signal stream is directed to the control signal path P2, which is used to analyze the acoustic environment and generate control parameters.

[0063] In a preferred embodiment, to improve the accuracy of subsequent control parameter generation, especially in environments with strong noise such as wind noise, a weak noise reduction module 20 is provided at the beginning of the control signal path P2. This module receives the signal from the audio acquisition module 10 and performs preliminary noise suppression processing on it. The purpose of this preprocessing is not to completely eliminate noise, but to moderately reduce the interference of noise on subsequent acoustic feature extraction, thereby enhancing the robustness of the scene judgment module 30 and the joint acoustic state encoder 40 during analysis.

[0064] The function of the weak noise reduction module 20 can be implemented through frequency domain processing. Specifically, the input signal is segmented into short-time audio frames and transformed to the frequency domain through short-time Fourier transform (STFT). For each frame, the system obtains an estimate of the current background noise power spectrum based on statistical information from historical non-speech frames. Based on this noise estimate, a spectral gain function is calculated and applied to the spectrum of the current frame.

[0065] A feasible method for calculating spectral gain is based on the power spectrum subtraction method. First, based on the input signal power spectrum and the estimated noise power spectrum, a gain value is calculated. This calculation process can be expressed by the following formula:

[0066] ;

[0067] In the formula: For application to the Frame, First Spectral gain at each frequency point; For the input audio signal at the 1st Frame, First Power spectrum values ​​at each frequency point; To estimate the background noise in the first... Frame, First Power spectrum values ​​at each frequency point; This is a preset over-reduction factor, whose value is usually greater than or equal to 1, used to control the intensity of noise removal; It is a preset spectral floor parameter, a small positive value, used to limit the maximum attenuation and avoid creating unnatural-sounding silent areas; For a function, return and The larger value in; This is for square root operations.

[0068] Calculate the gain Then, it is multiplied by the spectral amplitude of the input signal to obtain the noise-suppressed spectral amplitude. The processed spectral amplitude and the original spectral phase are recombine to form a complex spectrum, which is then transformed back to the time domain by the inverse short-time Fourier transform (ISTFT) to obtain the preprocessed audio signal. This preprocessed signal is then transmitted to the scene determination module 30 and the joint acoustic state encoder 40.

[0069] In step S2, refer to Figure 1 The scene determination module 30 analyzes the audio signal and generates a scene flag to characterize the current macroscopic acoustic scene. This scene flag is then transmitted to the parametric synthesizer 50 as the basis for selecting the basic processing strategy.

[0070] To accurately quantify the level of ambient noise, the user's voice needs to be excluded from the energy statistics. Therefore, the operation of the scene determination module 30 relies on the result of a speech activity detection (VAD). This VAD result can be generated by a separate VAD unit or directly taken from the speech probability in the acoustic state vector output by the joint acoustic state encoder 40 in subsequent steps. For each audio frame, the VAD result can be represented as a binary flag indicating whether the current frame is speech.

[0071] The scene determination module 30 receives audio signals from the audio acquisition module 10 or the weak noise reduction module 20 and processes them frame by frame. First, it calculates the short-time energy of each audio frame. Let the first frame be... The input signal of the frame is ,in The sampling point index represents the short-time energy of that frame. It can be calculated as:

[0072] ;

[0073] In the formula: For the first The short-time energy of a frame; For the first The first frame The signal amplitude at each sampling point; This represents the total number of sampling points in each frame.

[0074] Obtain the short-time energy of the current frame. and the VAD flag of the current frame. (in, Indicates a non-speech frame. After representing a speech frame, the scene determination module 30 updates the long-term estimate of the ambient noise level using a recursive averaging method. This update process is only performed when the current frame is determined to be a non-speech frame to ensure that the statistical results reflect the true background noise level. Ambient noise level The update can be performed according to the following formula:

[0075] ;

[0076] In the formula: The updated result is the first The ambient noise level of the frame; For the first The ambient noise level of the frame; For the current (number) The short-time energy of a frame; For the current (number) Voice activity detection flags for frames; It is a smoothing coefficient, with a value between 0 and 1, used to control the update rate of noise level estimation.

[0077] After calculating the current environmental noise level Then, the system compares it with a preset scene judgment threshold. Compare. If Greater than If so, a scene flag indicating a high-noise scene is generated; if Less than or equal to This generates a scene flag indicating a low-noise scene. This scene flag is then output to the parametric synthesizer 50.

[0078] In step S3, refer to Figure 1 The audio signal in the control signal path P2 is sent to the joint acoustic state encoder 40 for refined acoustic state analysis. The purpose of this module is to generate a multi-dimensional acoustic state vector that can comprehensively and accurately describe the detailed characteristics of the current audio frame and provide a basis for the decision of the parameter synthesizer 50 in step S4.

[0079] This analysis process first segments the continuous audio signal (from the audio acquisition module 10 or the weak noise reduction module 20) into a series of partially overlapping short audio frames. For each frame, the system extracts a set of features that characterize its acoustic properties, forming a high-dimensional feature vector. Extractable features include, but are not limited to: Mel-frequency cepstral coefficients (MFCC), linear predictive cepstral coefficients (LPCC), spectral entropy, spectral centroid, zero-crossing rate, etc.

[0080] Subsequently, the feature vector of each frame is input to the joint acoustic state encoder 40. In one specific embodiment, this encoder can be implemented by a pre-trained deep neural network model, such as a recurrent neural network (RNN) or an attention-based transformer network. This network receives the feature vector as input and outputs a structured acoustic state vector. This acoustic state vector includes the following key components:

[0081] Speech probability: A scalar value, typically between 0 and 1, representing the likelihood that the current frame contains human speech. This value can be directly used by other modules for speech activity detection.

[0082] Vocal intensity level: A scalar value used to quantify the effort exerted by the speaker. For example, this value can distinguish between normal conversation, loud talking, or shouting, as these different vocal signals have different dynamic characteristics.

[0083] Estimated Signal-to-Noise Ratio (SNR): A scalar value used to quantify the relative intensity of speech signal components and background noise components. This estimation process is performed in the frequency domain. First, for the spectrum of the current frame, the posterior SNR needs to be obtained. It is defined as the ratio of the power spectrum of the noisy speech to the estimated power spectrum of the noise.

[0084] Ambient noise feature vector: A vector used to describe the acoustic characteristics of the background noise itself, such as the shape, stationarity, or ripple characteristics of its spectrum. This enables the system to respond differently to different types of noise, such as steady-state wind noise and varying traffic noise.

[0085] Among these steps, estimating the signal-to-noise ratio (SNR) is a crucial one. It relies on the prior SNR... Accurate estimation of the power spectrum of clean speech to the estimated power spectrum of noise. A widely adopted estimation method is a decision-directed recursive update approach, with the following update formula:

[0086] ;

[0087] In the formula: For the first Frame, First The estimated prior signal-to-noise ratio at each frequency point; It is a smoothing factor, taking a value between 0 and 1, used to balance the weight of historical estimates and current observations; For the previous frame (the first frame) (frame) in the The power spectrum of the pure speech signal estimated at each frequency point; For the current frame (the first frame) (frame) in the The power spectrum of the noise signal estimated at each frequency point; For the current frame (the first frame) (frame) in the The a posteriori signal-to-noise ratio calculated at each frequency point; For a function, return and The larger value in the range is used to ensure the non-negativity of the signal-to-noise ratio estimate.

[0088] By all frequency points Averaging or other forms of integration yield a scalar value representing the overall signal-to-noise ratio of the current frame. Finally, the complete acoustic state vector containing all the above components is output to the parametric synthesizer 50.

[0089] In step S4, refer to Figure 1 and Figure 4The parameter synthesizer 50, as the control core of the system, is responsible for generating a set of audio processing parameters in real time. This module receives and fuses three input signals from different sources: a macroscopic scene marker generated by the scene judgment module 30, a microscopic acoustic state vector generated by the joint acoustic state encoder 40, and a correction vector fed back by the artifact monitor 80.

[0090] like Figure 4 As shown, the internal working mechanism of the parametric synthesizer 50 follows a hierarchical logic. At the macro level, the received scene flag (e.g., high-noise scene or low-noise scene) is used as a selector to choose the parameter set that best matches the current scene from multiple preset base parameter sets. For example, when the scene flag is a high-noise scene, the system selects a parameter set optimized for cycling environments, which has stronger noise suppression and voice enhancement characteristics. Conversely, it selects a parameter set optimized for indoor environments, focusing on sound reproduction.

[0091] At the micro level, the system performs real-time, fine-tuned dynamic adjustments based on a selected set of fundamental parameters using acoustic state vectors and correction vectors. Individual components of the acoustic state vector are used to fine-tune specific parameter values. For example, vocal intensity levels can adjust the threshold for dynamic range compression, and the estimated signal-to-noise ratio can affect the gain distribution of the equalizer. Simultaneously, correction vectors from the artifact monitor 80 are used to compensate for parameters, offsetting or reducing detected audio processing artifacts.

[0092] The audio processing parameters generated by the synthesizer 50 primarily include DRC parameters for dynamic range control and equalizer gain vectors for frequency domain adjustment. The DRC parameters determine how the system adjusts its gain based on the level of the input signal. (See reference...) Figure 3 The basic function of a dynamic range compressor is to compress compressors that exceed a preset threshold. The signal level is compressed at a certain ratio Compression is performed. Its output level... With input level The relationship can be described by the following formula:

[0093] ;

[0094] In the formula: The level of the output signal, measured in decibels; The input signal level is expressed in decibels (dB). The compression threshold is expressed in decibels. The compression ratio is a dimensionless positive number, representing the compression ratio for every increase in input level. Decibels, the output level increases by 1 decibel.

[0095] Besides compression threshold and compression ratio The DRC parameters also include start time and release time, which are used to control the rate of gain change when the signal level crosses the threshold.

[0096] Because the acoustic environment and input signal are continuously changing, the parametric synthesizer 50 generates a set of raw audio processing parameters frame by frame. There will be abrupt changes. To avoid these abrupt changes directly affecting the audio signal and introducing sudden changes in sound, the generated parameters need to be smoothed over time. This process is implemented using a first-order low-pass filter to calculate the final audio processing parameters applied to the current frame. The calculation process is as follows:

[0097] ;

[0098] In the formula: For the first The final audio processing parameter vector applied to the frame; For the first The frame is a vector of raw audio processing parameters directly generated by the parametric synthesizer; For the previous frame (the first frame) The final applied audio processing parameter vector (frame); It is a smoothing coefficient, with a value between 0 and 1, used to control the smoothness of parameter updates.

[0099] The final audio processing parameter vector obtained after smoothing It is output to the parametric audio processing engine 60.

[0100] In step S5, refer to Figure 1 The audio signal located on the main signal path P1 undergoes a series of processes to ultimately generate the audio output as the system output. The core of this processing path is to apply parameters generated by the control signal path P2 to adjust the dynamic range and equalization of the audio signal.

[0101] In a preferred embodiment, the processing flow on the main signal path P1 first includes a main noise reduction module 70. This module receives the raw audio signal from the audio acquisition module 10 and performs primary noise suppression operations on it. Compared to the weak noise reduction module 20 in the control signal path, the main noise reduction module 70 typically employs a more in-depth algorithm, aiming to eliminate background noise to the greatest extent possible, thereby improving the clarity and intelligibility of human voices in the final output audio. The signal output by this module is the noise-reduced audio signal.

[0102] Subsequently, the denoised audio signal is transmitted to the parametric audio processing engine 60. This engine simultaneously receives the smoothed final audio processing parameter vector from the parametric synthesizer 50. The core function of the parametric audio processing engine 60 is to act as an execution unit, processing... The parameters defined in the code are applied precisely to the audio signal.

[0103] The application process mainly includes two parts: dynamic range control and equalization processing. Dynamic range control is based on... This is achieved through DRC parameters (i.e., compression threshold, compression ratio, start-up time, and release time). An internal level detector continuously monitors the level of the noise-reduced audio signal and compares it to the compression threshold. When the signal level changes, the engine smoothly adjusts the gain applied to the signal according to the start-up and release time requirements, ensuring its dynamic characteristics conform to the input-output relationship defined by the compression threshold and compression ratio.

[0104] Balanced processing is based on The equalizer gain vector is used to implement this. This process is usually done in the frequency domain. The denoised audio signal (or the signal after DRC processing) is transformed to the frequency domain to obtain its complex spectrum. Then, its spectral amplitude is multiplied by the equalizer gain vector at each frequency point. This process can be expressed by the following formula:

[0105] ;

[0106] In the formula: In the first Frame, First The amplitude of the signal spectrum after equalization at each frequency point; In the first Frame, First The spectral amplitude of the input noise-reduced audio signal at each frequency point; In the first Frame, First The equalizer gain applied at each frequency point, this value is determined by the parameter vector. supply.

[0107] Processed spectral amplitude Combined with the original spectral phase, it is transformed back to the time domain through an inverse Fourier transform. Finally, the signal processed by the parametric audio processing engine 60, i.e., the processed audio signal, serves as the final output of the entire system.

[0108] In step S6, refer to Figure 1The entire system forms a closed-loop control system through an artifact monitor 80. The input of this module is connected to the output of the parametric audio processing engine 60 to receive and analyze the final processed audio signal. Its output is connected to the input of the parametric synthesizer 50, feeding back the analysis results as a correction vector. This feedback loop allows the system to monitor its processing performance and proactively correct audio processing artifacts introduced by improper parameter settings.

[0109] The artifact monitor 80 performs frame-by-frame analysis of the processed audio signal to detect two specific audio processing artifacts: noise-pumping and over-compression.

[0110] The noise sucking effect refers to the perceptible, unnatural fluctuations in the level of background noise between sentences with and without speech. The artifact monitor 80 detects this effect by tracking the background noise energy of the processed audio signal only in frames determined to be non-speech, using speech probability information from the joint acoustic state encoder 40. First, the output signal frame is calculated... short-term energy Then, update the noise level estimate of the output signal. If the rate of change of the noise level exceeds a preset suction effect threshold... If the rate of change is 1, then a noise suction effect is considered to have been detected. This rate of change can be calculated using the following formula:

[0111] ;

[0112] In the formula: For the first A measure of the noise sucking effect of a frame; For the first Noise level estimation of the output signal after frame update; For the first Estimation of the noise level of the output signal of a frame.

[0113] Overcompression distortion refers to the excessive compression of the dynamic range, resulting in a flattened, lacking, and even harmonic distortion in the sound. The artifact monitor 80 detects this distortion by analyzing the dynamic characteristics of the output signal; a feasible metric is the signal's crest factor. First, the output signal frame... peak factor :

[0114] ;

[0115] In the formula: For the first The crest factor of the frame; For the first The first frame The amplitude of the output signal at each sampling point; This represents the peak value of the signal amplitude in that frame. It is a square root operation, and its internal value is the root mean square (RMS) value of the signal in that frame; This represents the total number of sampling points in each frame.

[0116] Calculated crest factor With a preset overcompression threshold Compare. If Below If so, it is considered that excessive compression distortion has been detected.

[0117] Based on the above detection results, the artifact monitor 80 generates a correction vector. This vector is a multi-dimensional vector containing suggestions for adjusting specific audio processing parameters. For example, if a noise sucking effect is detected, then... The element corresponding to the DRC release time will be set to a positive value to suggest that the synthesizer 50 extend the release time. If excessive compression distortion is detected, the element corresponding to the DRC compression ratio or compression threshold will be adjusted to suggest that the synthesizer 50 reduce the compression level. This correction vector The feedback is then sent to the parameter synthesizer 50, which is used to correct the parameter generation process of the next frame in step S4, thereby achieving closed-loop adaptive control of the output audio quality.

[0118] To further illustrate the collaborative workflow of each module in the embodiments of the present invention, the following will be illustrated using a typical usage scenario: The user first uses motorcycle riding headphones in a quiet indoor environment, and then enters a noisy environment of high-speed riding.

[0119] Initially, the user is in a quiet indoor environment. The background noise level in the signal acquired by the audio acquisition module 10 is very low. At this time, the scene judgment module 30 calculates the ambient noise level during non-speech periods. The scene judgment threshold is lower than the preset threshold. Therefore, a low-noise scene marker is generated. Simultaneously, the acoustic state encoder 40 analyzes the signal; when the user speaks normally, the output acoustic state vector indicates a high speech probability, normal vocal intensity, and a high estimated signal-to-noise ratio. The parametric synthesizer 50 receives the low-noise scene marker and selects a set of basic parameters focused on sound fidelity. Based on this parameter set, and combined with real-time information from the acoustic state vector, the parametric synthesizer 50 generates a set of DRC parameters with a low compression ratio, a high compression threshold, and a relatively flat equalizer gain, aiming for minimal processing of the audio signal. The parametric audio processing engine 60 applies these parameters, ensuring that the dynamic range and frequency domain characteristics of the output audio closely approximate the original input, guaranteeing high fidelity. At this stage, due to the weak processing intensity, the artifact detector 80 typically does not detect any processing artifacts, and its feedback correction vector is empty or zero.

[0120] As the user begins riding and gradually accelerates, wind noise and engine noise increase sharply. Scene judgment module 30 detects energy levels during non-voice periods. The rapid increase has led to a rise in environmental noise levels. Exceeded the threshold in a short period of time The scene indicator then switches to a high-noise scene. At the same time, the acoustic state vector output by the joint acoustic state encoder 40 also changes, the estimated signal-to-noise ratio decreases significantly, and the vocal intensity level increases accordingly when the user increases the volume to drown out the noise.

[0121] Upon receiving a high-noise scene indicator, the parametric synthesizer 50 immediately switches to a base parameter set designed for cycling environments with enhanced noise suppression capabilities. Based on this, and taking into account the low signal-to-noise ratio and high vocal intensity level in the acoustic state vector, it further generates DRC parameters with a lower compression threshold and higher compression ratio, as well as equalizer gain to boost the vocal frequency band. Notably, the switch from low-noise to high-noise parameters is not instantaneous but rather a smooth transition via a parameter smoothing formula, ensuring auditory continuity. The main noise reduction module 70 and the parametric audio processing engine 60 apply this new, more sophisticated set of parameters to powerfully reduce noise and compress the dynamic range of the signal, effectively extracting and enhancing the vocal signal in a high-noise background.

[0122] During this high-intensity processing, the closed-loop feedback of the artifact monitor 80 becomes crucial. If an excessively high compression ratio leads to over-compression distortion (i.e., crest factor...), it can cause significant damage. Below the threshold Or, an excessively fast release time can trigger a noise sucking effect (i.e., the rate of change of noise level). Above the threshold The artifact detector 80 will generate a correction vector containing correction suggestions. This feedback is then sent to the parametric synthesizer 50. In the next processing cycle, the parametric synthesizer 50 will take this feedback into account and adjust the parameters it generates appropriately (for example, by appropriately reducing the compression ratio or extending the release time). This will allow it to actively suppress processing artifacts while ensuring the clarity of vocals, ultimately achieving a stable, clear, and natural-sounding output in the dynamic environment of high-speed cycling.

[0123] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A DRC control method for improving the indoor riding fidelity of motorcycle riding headphones, characterized in that, The method includes the following steps: S1. Acquire the audio signal collected by the microphone and distribute the audio signal to the main signal path and the control signal path; S2. Analyze the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generate a scene flag; S3. In the control signal path, the audio signal is analyzed by a joint acoustic state encoder to generate an acoustic state vector describing the current acoustic environment and human voice state. S4. A set of audio processing parameters is generated in real time using a parametric synthesizer based on the scene flags, acoustic state vector, and a feedback correction vector. S5. In the main signal path, a parameterized audio processing engine is used to perform dynamic range control and equalization processing on the audio signal using the audio processing parameters to generate the processed audio signal as the final output. S6. The processed audio signal is analyzed by an artifact monitor to detect audio processing artifacts introduced by the dynamic range control and equalization processing, and the correction vector is generated and fed back to the parameter synthesizer to form a closed-loop control.

2. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, Step S1, which involves acquiring the audio signal collected by the microphone and distributing the audio signal to the main signal path and the control signal path, includes: The audio signal is input into a weak noise reduction module for preprocessing to generate a preprocessed audio signal, thereby improving the robustness of subsequent acoustic feature extraction. The control signal path receives the preprocessed audio signal.

3. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, Step S2, which involves analyzing the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generating a scene marker, includes: The energy of the non-human voice portion of the audio signal is statistically analyzed to obtain the environmental noise level. The ambient noise level is compared with a preset scene judgment threshold to generate a scene flag, which is used to indicate whether the current scene is a high noise scene or a low noise scene.

4. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, In step S3, the step of analyzing the audio signal through a joint acoustic state encoder in the control signal path includes: The audio signal in the control signal path is segmented into short audio frames, and the acoustic features of each short audio frame are extracted to obtain a feature vector. The feature vector is input into the joint acoustic state encoder, and the joint acoustic state encoder outputs the acoustic state vector. The acoustic state vector includes: a speech probability characterizing the possibility that the current frame is human speech, a voice intensity level quantifying the degree of effort required for the user's voice, an estimated signal-to-noise ratio characterizing the relative intensity of the voice signal and the background noise, and an environmental noise feature subvector describing the acoustic characteristics of the background noise.

5. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, Step S4, which involves generating a set of audio processing parameters in real time using a parametric synthesizer based on the acoustic state vector and a feedback correction vector, includes the following steps: The parameter synthesizer fuses the acoustic state vector and the correction vector to generate the audio processing parameters, and performs temporal smoothing filtering on the generated audio processing parameters to ensure that the audio processing parameters change continuously and smoothly between different audio frames.

6. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 5, characterized in that, The audio processing parameters include: DRC parameters used to perform dynamic range control include compression threshold, compression ratio, start time, and release time; Equalizer gain vector used to adjust specific frequency bands.

7. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 5, characterized in that, The temporal smoothing filtering is achieved through a first-order low-pass filter, and the final audio processing parameters applied to the current frame are calculated according to the following formula: ; In the formula, These are the audio processing parameters that will be applied to the current frame. These are the raw audio processing parameters generated by the parametric synthesizer for the current frame. These are the audio processing parameters that were finally applied in the previous frame. This is the smoothing coefficient.

8. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, In step S5, the step of performing dynamic range control and equalization processing on the audio signal using the audio processing parameters through a parameterized audio processing engine in the main signal path includes: The audio signal is input into a main noise reduction module for processing to obtain a noise-reduced audio signal. The dynamic range control and equalization processing are applied to the noise-reduced audio signal.

9. The DRC control method for improving the indoor riding fidelity of motorcycle riding headphones according to claim 1, characterized in that, Step S6, which involves analyzing the processed audio signal using an artifact monitor to detect audio processing artifacts introduced by the dynamic range control and equalization processing, includes: The artifact monitor analyzes the processed audio signal to detect noise sucking or over-compression distortion introduced by the dynamic range control and equalization processing, and generates the correction vector based on the detection results.

10. A DRC control system for improving the indoor riding fidelity of motorcycle riding headphones, applied to the method described in any one of claims 1-9, characterized in that, The system includes: The audio acquisition module is used to acquire audio signals captured by the microphone and distribute the audio signals to the main signal path and the control signal path; The scene determination module is used to analyze the acoustic characteristics of the audio signal to determine the current macroscopic acoustic scene and generate a scene flag. A joint acoustic state encoder, connected to the control signal path, is used to analyze the audio signals in the control signal path to generate an acoustic state vector describing the current acoustic environment and the state of human voice. A parameter synthesizer, connected to the scene judgment module, the joint acoustic state encoder, and an artifact monitor, is used to generate a set of audio processing parameters in real time based on the scene flag, the acoustic state vector, and a feedback correction vector. A parametric audio processing engine connects the main signal path and the parametric synthesizer, and is used to perform dynamic range control and equalization processing on the audio signal in the main signal path using the audio processing parameters to generate a processed audio signal as the final output. An artifact monitor, connected to the output of the parametric audio processing engine, is used to analyze the processed audio signal to detect audio processing artifacts introduced by the dynamic range control and equalization processing, and generate the correction vector to feed back to the parametric synthesizer, thereby forming a closed-loop control.

Citation Information

Patent Citations

  • Adaptive sound effect management method and system with scene analysis module

    CN120547494A

  • Interactive virtual reality generation method

    CN120560498A