Audio signal self-adaptive compensation adjusting system and method applied to earphone

By integrating a multimodal sensing module and a dynamic compensation state machine into the headphones, the cause of headphone leakage can be identified and the compensation gain can be dynamically adjusted, solving the problem of distinguishing between passive and physiological leakage in headphones and improving audio quality and user experience.

CN121194099APending Publication Date: 2025-12-23BESING TECH SHENZHEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511706321.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Current headphone technology cannot accurately distinguish between passive leakage caused by loose headphone fit and physiological leakage caused by user activity, resulting in a single compensation strategy that affects audio quality and user comfort.

Method used

A multimodal sensing module is used to simultaneously acquire acoustic signals and physiological vibration signals. Through cross-modal fusion analysis, the cause of leakage is identified, and the compensation gain is dynamically adjusted to distinguish between passive and physiological leakage, thereby achieving intelligent audio compensation.

Benefits of technology

It improves the accuracy of leakage source classification, ensures reliable system operation under complex dynamic conditions, and enhances the intelligence level of audio compensation and user auditory comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121194099A_ABST
    Figure CN121194099A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio signal processing, and discloses an audio signal adaptive compensation adjustment system and method applied to an earphone, and the system comprises a multi-mode sensing module which is configured to be used for collecting an actual acoustic signal, a physiological vibration signal and a reference audio signal; a theoretical compensation calculation module configured to calculate a theoretical compensation gain based on the reference audio and the actual acoustic signal; the leakage source identification module is configured to be used for extracting acoustic and vibration characteristics and executing cross-modal fusion analysis so as to output a leakage source classification state; and the dynamic compensation state machine and application module is configured for selecting different smoothing parameters and smoothing the theoretical gain according to the leakage source state so as to calculate the final compensation gain and applying the final compensation gain to the reference audio signal. According to the invention, through cross-modal fusion analysis of acoustic and physiological vibration, accurate distinguishing and dynamic smooth compensation of leakage sources are realized, and hearing comfort is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, specifically to an audio signal adaptive compensation and adjustment system and method for headphones. Background Technology

[0002] In the use of headphones (such as in-ear headphones), the acoustic seal between the headphone unit and the user's ear canal is crucial for achieving accurate low-frequency response. However, due to loose fit or user activity, acoustic leakage is inevitable, resulting in significant loss of low-frequency energy and ruining the original audio experience. Therefore, developing technologies that can monitor leakage in real time and perform dynamic audio compensation has become key to improving headphone sound quality.

[0003] To address the aforementioned issues, existing technologies typically employ a compensation scheme based on acoustic feedback. This scheme uses a feedback microphone deployed inside the earphone to acquire the actual acoustic signal within the ear canal in real time and compares this signal with a reference audio signal played by the system. By analyzing the energy difference between the two in the low-frequency band, the system can calculate a theoretical compensation gain and apply this gain back to the audio signal, hoping to offset the sound loss caused by leakage.

[0004] While existing acoustic feedback schemes can detect acoustic leakage to some extent, they have several shortcomings. These schemes heavily rely on single acoustic sensor information, and their fundamental limitation is that while the system can detect leakage, it cannot understand its cause. Specifically, whether it's stable passive leakage due to loose headphones or transient physiological leakage caused by jaw movements such as speaking or chewing, both manifest as similar low-frequency energy loss on the feedback microphone. The system cannot distinguish between these two completely different leakage sources based solely on acoustic data. Furthermore, due to the inability to accurately identify the cause of leakage, existing technologies typically employ a one-size-fits-all fixed smoothing strategy when applying compensation. The drawback of this design is that when the system uses a slow smoothing coefficient for stability, it cannot respond promptly to transient physiological leakage; conversely, if the system blindly uses a fast coefficient, it will cause frequent gain fluctuations when facing stable leakage, producing a twitching effect that severely impacts the accuracy of audio compensation and the user's auditory comfort. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an adaptive compensation and adjustment system and method for audio signals applied to headphones, which solves the problem that existing technologies rely on only a single acoustic signal and cannot distinguish between passive leakage and physiological leakage, thus resulting in a single compensation strategy and poor listening experience.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of the present invention provides an audio signal adaptive compensation and adjustment system for headphones, the system comprising: A multimodal sensing module is configured to acquire actual acoustic signals and physiological vibration signals within the ear canal, and to receive reference audio signals; The theoretical compensation calculation module is connected to the multimodal sensing module and is configured to calculate the theoretical compensation gain based on the reference audio signal and the actual acoustic signal. The leakage source identification module, connected to the multimodal sensing module and the theoretical compensation calculation module, is configured to: determine acoustic leakage characteristics based on the theoretical compensation gain, determine physiological vibration characteristics based on the physiological vibration signal, and perform cross-modal fusion analysis on the acoustic leakage characteristics and the physiological vibration characteristics to output a leakage source classification status characterizing the cause of leakage; The dynamic compensation state machine and application module, connected to the theoretical compensation calculation module and the leakage source identification module, are configured to: select different smoothing parameters to smooth the theoretical compensation gain according to the leakage source classification state, calculate the final compensation gain, and apply it to the reference audio signal.

[0007] Preferably, the theoretical compensation calculation module is specifically configured for: Perform time-frequency transformation on the reference audio signal and the actual acoustic signal; Calculate the low-frequency energy of the reference signal and the low-frequency energy of the actual acoustic signal within the preset target low-frequency band; The difference between the low-frequency energy of the reference signal and the low-frequency energy of the actual acoustic signal is calculated in the logarithmic domain to obtain the theoretical compensation gain.

[0008] Preferably, the leakage source identification module is configured to: The sequence of the theoretically compensated gain is used as the acoustic leakage characteristic; The physiological vibration signal is bandpass filtered, and the short-time energy of the filtered signal is calculated to obtain the physiological vibration characteristics.

[0009] Preferably, the leakage source identification module performs the cross-modal fusion analysis specifically including: Within the analysis window, the cross-modal coherence between the acoustic leakage characteristics and the physiological vibration characteristics is calculated; Calculate the causal time-domain delay between the acoustic leakage characteristics and the physiological vibration characteristics.

[0010] Preferably, the leakage source identification module is further configured to: The cross-modal coherence is compared with the coherence threshold, and the causal time-domain delay is compared with the physiological causal time window; When the theoretical compensation gain exceeds the amplitude threshold, the cross-modal coherence is greater than the coherence threshold, and the causal time-domain delay is within the physiological causal time window, the leakage source classification state is determined to be a physiological leakage state. When the theoretical compensation gain exceeds the amplitude threshold, but the cross-modal coherence is not greater than the coherence threshold or the causal time-domain delay is not within the physiological causal time window, the leakage source classification state is determined to be a passive leakage state.

[0011] Preferably, the dynamic compensation state machine and application module are specifically configured for: When the leakage source classification status is the passive leakage status, the slow coefficient is selected as the smoothing parameter; When the leakage source classification status is the physiological leakage status, the rapid coefficient is selected as the smoothing parameter; Wherein, the slow coefficient is greater than the fast coefficient.

[0012] Preferably, the dynamic compensation state machine and application module use a first-order infinite impulse response filter to smooth the theoretical compensation gain, wherein the smoothing parameter is the smoothing coefficient of the first-order infinite impulse response filter.

[0013] Preferably, the dynamic compensation state machine and application module are further configured to: The final compensation gain in the logarithmic field is converted into a linear amplitude multiplier; In the frequency domain, the linear amplitude multiplier is applied to the frequency domain representation of the reference audio signal only within a preset target low-frequency band; The frequency domain signal after gain is applied is subjected to inverse time-frequency transformation and superposition addition to reconstruct the output audio signal.

[0014] Preferably, the multimodal sensing module includes: The acoustic sensing unit is a feedback microphone deployed inside the earphone, used to collect the actual acoustic signal; The vibration sensing unit, which may be an inertial measurement unit, an accelerometer, or a bone conduction sensor, is used to acquire the physiological vibration signals.

[0015] A second aspect of the present invention provides an adaptive compensation adjustment method for audio signals applied to headphones, the method comprising the following steps: Acquire actual acoustic and physiological vibration signals within the ear canal, and receive reference audio signals; The theoretical compensation gain is calculated based on the reference audio signal and the actual acoustic signal; Acoustic leakage characteristics are determined based on the theoretical compensation gain, and physiological vibration characteristics are determined based on the physiological vibration signal; Perform cross-modal fusion analysis on the acoustic leakage characteristics and the physiological vibration characteristics to output a leakage source classification status characterizing the cause of leakage; Based on the classification status of the leakage source, different smoothing parameters are selected to smooth the theoretical compensation gain in order to calculate the final compensation gain; The final compensation gain is applied to the reference audio signal to obtain the compensated output audio signal.

[0016] This invention provides an adaptive compensation and adjustment system and method for audio signals applied to headphones. It has the following beneficial effects: 1. This invention, by setting up a multimodal sensing module to simultaneously acquire acoustic signals and physiological vibration signals, and utilizing a leakage source identification module to perform cross-modal fusion analysis on the characteristics of the two signals, can accurately distinguish between passive leakage caused by loose wearing and physiological leakage caused by activities such as speaking and chewing. This solves the deficiency of traditional solutions that cannot identify the cause of leakage based solely on acoustic signals, and improves the accuracy of leakage source classification.

[0017] 2. This invention, by setting up a dynamic compensation state machine and application module, can intelligently select different smoothing parameters to handle the theoretical compensation gain based on the classification status output by the leakage source identification module. This dynamic adjustment mechanism enables the system to respond quickly when facing physiological leakage and to transition smoothly when facing passive leakage, avoiding incorrect or over-compensation, and improving the intelligence level of audio compensation and the user's auditory comfort.

[0018] 3. This invention introduces the calculation of the causal time-domain delay of acoustic leakage characteristics and physiological vibration characteristics into cross-modal fusion analysis and compares it with the physiological causal time window. This technical feature provides stronger physical constraints for determining the correlation between the two signals, ensuring that the system can filter out occasional, unrelated vibration interference, further enhancing the robustness of physiological leakage event identification, and ensuring the reliable operation of the compensation system under complex dynamic conditions. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the audio signal adaptive compensation and adjustment system of the present invention; Figure 2 This is a flowchart illustrating the adaptive compensation and adjustment method for audio signals according to the present invention. Figure 3 This is a schematic diagram of the internal structure of the multimodal sensing module of the present invention; Figure 4 This is a schematic diagram of the internal structure of the theoretical compensation calculation module of the present invention; Figure 5 This is a schematic diagram of the internal structure of the leakage source identification module of the present invention; Figure 6 This is a schematic diagram of the internal structure of the dynamic compensation state machine and application module of the present invention. Detailed Implementation

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Please see the appendix Figure 1 , Figure 1 This is a schematic diagram of the structure of an audio signal adaptive compensation adjustment system for headphones according to the present invention. The system may include: A multimodal sensing module is configured to simultaneously acquire at least two signals: actual acoustic signals within the ear canal acquired by an acoustic sensing unit, and physiological vibration signals acquired by a vibration sensing unit. This multimodal sensing module can also receive a reference audio signal to be played.

[0022] The theoretical compensation calculation module is connected to the multimodal sensing module and is configured to compare the reference audio signal and the actual acoustic signal to calculate the theoretical compensation gain required to compensate for the energy difference between the two in the low-frequency band.

[0023] The leakage source identification module is connected to both the multimodal sensing module and the theoretical compensation calculation module. This module is configured to perform cross-modal fusion analysis on the theoretical compensation gain (as a characteristic of acoustic leakage) and physiological vibration signals. By calculating the cross-modal coherence and causal time-domain delay between the two signals, the module determines whether the leakage is caused by passive wearing or active physiological activity, and outputs the leakage source classification status accordingly.

[0024] The dynamic compensation state machine and application module connect the theoretical compensation calculation module and the leakage source identification module. Based on the state output by the leakage source identification module, this dynamic compensation state machine and application module automatically switches its internal algorithm state machine. Based on this state, the state machine selects different smoothing parameters (e.g., slow or fast coefficients) to process the theoretical compensation gain, calculates the final compensation gain, applies it to the low-frequency band of the reference audio signal, and finally outputs the compensated audio signal.

[0025] Please see the appendix Figure 2 , Figure 2This is a schematic flowchart of an audio signal adaptive compensation adjustment method for headphones according to the present invention. The method may include the following steps: S100: Multimodal signal synchronous acquisition. The system acquires three signal streams: Reference audio signal, i.e., the original playback signal; The actual acoustic signal comes from the feedback microphone inside the headphones; Physiological vibration signals originate from an IMU or bone conduction sensor. The signals are divided into frames for processing.

[0026] S200: Calculation of theoretical compensation gain. First, time-frequency transformation is performed on the reference signal and the actual acoustic signal. Next, the signal energy within the preset low-frequency band is calculated. Finally, the difference between the low-frequency energy of the reference signal and the low-frequency energy of the actual acoustic signal is calculated in the logarithmic domain to obtain the original theoretical compensation gain characterizing the acoustic leakage amplitude.

[0027] S300: Leakage source identification. First, two feature sequences are extracted: acoustic leakage features (e.g., which can be represented by theoretically compensated gain) and physiological vibration features (e.g., obtained by filtering the physiological vibration signal and calculating its short-time energy).

[0028] Secondly, within the analysis window, the cross-modal coherence of acoustic leakage characteristics and physiological vibration characteristics is calculated (e.g., through normalized cross-correlation) to quantify the degree to which they occur synchronously.

[0029] Simultaneously, causal time-domain delays were calculated to verify whether physiological vibration characteristics (in accordance with physical laws) precede acoustic leakage characteristics in the time domain.

[0030] Finally, based on the preset leakage amplitude threshold, coherence threshold, and expected delay range, the leakage sources are classified and the leakage source classification status is output (e.g., Active, Passive, None).

[0031] S400: Dynamic Compensation and State Switching. This step dynamically adjusts the smoothing method for the theoretical compensation gain based on the leakage source classification state to calculate the final smoothed compensation gain.

[0032] This processing employs a smoothing filter (e.g., a first-order IIR filter), whose smoothing coefficients are dynamically determined by the state machine: If the state is passive leakage, the state machine selects a slow coefficient so that the final gain slowly and steadily approaches the theoretical gain, achieving a strong correction.

[0033] If the state is physiological leakage, the state machine selects a fast coefficient so that the final gain only follows the approximate envelope of the theoretical gain, quickly smoothing out its severe instantaneous peaks and valleys, and achieving artifact suppression.

[0034] If the state is leak-free, the gain eventually decays smoothly to 0.

[0035] S500: Signal compensation and reconstruction. The resulting logarithmic domain smoothed compensation gain is then converted back to linear amplitude multipliers.

[0036] In the frequency domain, the amplitude multiplier is applied to the preset target low-frequency band of the reference signal.

[0037] Finally, a short-time inverse Fourier transform is performed on the frequency domain signal after the gain is applied, and the final output audio signal with adaptive low-frequency compensation is reconstructed by the overlapping addition method.

[0038] This invention identifies the true physical source of low-frequency leakage through the fusion analysis of acoustic and vibration sensors, and dynamically switches the state machine of the low-frequency compensation algorithm according to the classification of the leakage source. When the leakage is determined to be passive, the algorithm performs stable low-frequency enhancement correction; when the leakage is determined to be physiological activity leakage, the algorithm switches to an artifact suppression state, thereby solving the problem of algorithm instability in dynamic scenarios while ensuring the low-frequency compensation effect.

[0039] Please see the appendix Figure 3 , Figure 3 This is a schematic diagram of the internal structure of a multimodal sensing module according to an embodiment of the present invention. This multimodal sensing module is the real-time data input terminal of the system of the present invention, configured to acquire and preprocess multiple data streams required by the subsequent low-frequency compensation algorithm module.

[0040] The multimodal sensing module may specifically include an acoustic sensing unit. This acoustic sensing unit is a microphone deployed near the headphone's sound outlet, with its acoustic port facing the user's ear canal; it is also commonly referred to in the art as a feedback microphone. The function of this acoustic sensing unit is to capture the actual acoustic pressure signal within the user's ear canal in real time, and this actual acoustic pressure signal is digitized and represented as... This reflects the actual audio content received by the user's eardrum after passing through the acoustic characteristics of the user's ear canal and due to the seal of the earpiece. This is an index for discrete-time samples.

[0041] This multimodal sensing module also includes a vibration sensing unit. This vibration sensing unit is configured to detect physical vibrations transmitted to the earphone shell via the skull or body tissues, caused by the user's physiological activities (such as jaw movements, speaking, and chewing). The output signal of this vibration sensing unit is digitized and represented as a physiological vibration signal. This physiological vibration signal is a non-acoustic basis for subsequent modules to determine the cause of low-frequency leakage.

[0042] In one specific embodiment, the vibration sensing unit can be a triaxial or uniaxial accelerometer in an inertial measurement unit (IMU). This accelerometer captures bone conduction vibrations with specific patterns caused by actions such as speaking or chewing by detecting instantaneous acceleration changes in the earphone shell in a specific direction. In another embodiment, the vibration sensing unit can also be a piezoelectric ceramic sensor or a bone conduction pickup known in the art; such sensors can efficiently convert the aforementioned physical vibrations into electrical signals.

[0043] The multimodal sensing module may also include a signal preprocessing unit. This signal preprocessing unit receives the outputs from the acoustic sensing unit and the vibration sensing unit, as well as a digital reference audio signal from the audio playback core. (i.e., the original signal to be played), and process these signals to make them suitable for subsequent digital signal processing modules. The operation of this signal preprocessing unit may include the following steps: S110: Signal Acquisition and Conversion. This signal preprocessing unit uses an analog-to-digital converter to convert the analog signals output from the acoustic and vibration sensing units into signals with a preset sampling rate. digital signals and .

[0044] S111: Signal synchronization. This signal preprocessing unit processes digital signals... , and reference signal Time alignment is performed to establish a common time reference. This synchronization operation ensures... It reflects The actual result in the ear canal at the corresponding time of playback, and and They share a common time reference, which forms the basis for subsequent cross-modal causal time-domain analysis.

[0045] S112: Windowing and Framing. This signal preprocessing unit applies a window function to the three synchronized continuous signal streams and segments them, cutting the signal into segments with fixed lengths. and fixed frame shift The processing frame. This step converts the continuous signal stream into a series of discrete data blocks for subsequent modules to process frame by frame.

[0046] The signal preprocessing unit outputs the framed data stream to the theoretical compensation calculation module and the leakage source identification module, respectively, for subsequent processing.

[0047] Please see the appendix Figure 4 , Figure 4This is a schematic diagram of the internal structure of a theoretical compensation calculation module according to an embodiment of the present invention. This theoretical compensation calculation module is connected to a multimodal sensing module, and its core function is to calculate the theoretical low-frequency compensation gain in real time based on acoustic sensing signals.

[0048] The theoretical low-frequency compensation gain characterizes the dynamic low-frequency enhancement required to compensate for low-frequency loss caused by wear leakage (regardless of its cause). The theoretical compensation gain is calculated by this theoretical compensation calculation module. It will be used as one of the analysis objects of the subsequent leakage source identification module (i.e., as acoustic leakage characteristics), and also as the target of algorithm smoothing for the dynamic compensation state machine and application module.

[0049] The specific workflow of this theoretical compensation calculation module may include the following steps: S210: Signal time-frequency transformation. This theoretical compensation calculation module receives the reference audio signal preprocessed and framed by the multimodal sensing module. and actual acoustic signal This theoretical compensation calculation module transforms the two time-domain signal frames from the time domain to the frequency domain by performing a time-frequency transformation (e.g., Short-Time Fourier Transform, STFT), obtaining their values ​​in the 1st... Frame, First Frequency domain representation of frequency band and .

[0050] S220: Target low-frequency band energy calculation. The system pre-sets the target low-frequency compensation band. (For example, the frequency range of 20Hz to 250Hz), the target low-frequency compensation band defines the frequency range in which dynamic low-frequency enhancement is required.

[0051] This step S220 is in Within the scope, through the and The energy is calculated by performing energy calculations on the frequency domain representation (e.g., summing the squares of its amplitudes) to obtain the reference signal at the th . Frame target low-frequency energy And the actual acoustic signal at the 1st Frame target low-frequency energy When acoustic leakage exists, The value will be lower than .

[0052] S230: Theoretical compensation gain calculation. This step S230 is based on the calculation obtained in S220. and The gain value needed to compensate for the difference between the two is calculated. In one embodiment, this gain is calculated in the logarithmic domain, i.e., by calculating... and The logarithm of the ratio between them, and applying a nonnegativity constraint (e.g., by taking 0 and calculating the maximum value), to obtain the first... Theoretical compensation gain of frames The calculation may also include operations to maintain numerical stability (such as adding a minimum value to the denominator).

[0053] The theoretical compensation calculation module in each processing frame All output the theoretical compensation gain Sequence. This The sequence objectively reflects the instantaneous degree of acoustic leakage in the ear canal. This occurs when the user speaks or chews. It can change drastically due to the instantaneous deformation of the ear canal, leading to The sequence exhibits large, rapid fluctuations.

[0054] Should The sequence is transmitted to the leak source identification module for use as acoustic leak characteristics; simultaneously, the The sequence is also transmitted to the dynamic compensation state machine and application module as the original target gain for smoothing.

[0055] Please see the appendix Figure 5 , Figure 5 This is a schematic diagram of the internal structure of a leakage source identification module according to an embodiment of the present invention. This leakage source identification module is the core decision-making unit of the system of the present invention, and its function is to intelligently determine the low-frequency leakage (i.e., theoretical compensation gain) calculated by the theoretical compensation calculation module. The physical causes of ).

[0056] The leakage source identification module is connected to both the theoretical compensation calculation module and the multimodal sensing module. It receives... Sequence (as acoustic leakage characteristics) and original physiological vibration signals (As a criterion for physiological activity), by performing cross-modal fusion analysis, discrete leakage source classification status is output. This status signal It will directly control the algorithmic behavior of the dynamic compensation state machine and application module, enabling it to intelligently switch between dynamic low-frequency compensation and artifact suppression.

[0057] The leakage source identification module may specifically include a feature extraction unit, a cross-modal analysis unit, and a leakage source classification unit.

[0058] The feature extraction unit is responsible for extracting key feature sequences from the two input signal streams for subsequent analysis. The operation of this feature extraction unit may include the following steps: S311: Acoustic Leakage Feature Extraction. This feature extraction unit receives the theoretical compensation gain sequence from the theoretical compensation calculation module. The wave pattern (i.e., its envelope) of this sequence in the time domain directly reflects the amplitude variation of acoustic leakage. Therefore, this feature extraction unit uses this theoretically compensated gain sequence. Directly used as an acoustic leakage characteristic : ; In the formula, For the first Acoustic leakage characteristics of the frame; The first one calculated by the theoretical compensation calculation module Theoretical compensation gain for frames; This is the index of the current processing frame.

[0059] S312: Physiological vibration feature extraction. This feature extraction unit receives raw physiological vibration signals from the multimodal sensing module. (or already framed) To extract the signal components most relevant to jaw movements (such as speaking and chewing), this feature extraction unit first... Perform bandpass filtering (e.g., using an IIR or FIR bandpass filter to extract the frequency band from 80Hz to 300Hz) to obtain the filtered signal. This frequency band corresponds to the main energy range of bone conduction speech and chewing vibrations.

[0060] Subsequently, this feature extraction unit calculates the filtered signal. In the The short-time energy of a frame as a physiological vibration characteristic : ; In the formula, For the first Physiological vibration characteristics of frames; For discrete-time sample indexes; This is the index of the current processing frame; The number of samples for frame shift; The number of samples is the frame length; The physiological vibration signal after bandpass filtering at time point Sample values; This indicates calculating the square of the signal sample value.

[0061] The cross-modal analysis unit is responsible for processing the two feature sequences output by the feature extraction unit. and A comparative analysis of time and causality was conducted. This comparative analysis was conducted in length [length missing]. The analysis is performed within a sliding window of the frame. The operation of this cross-modal analysis unit may include the following steps: S321: Cross-modal coherence analysis. This step, S321, is used to quantify acoustic leakage. and physiological vibration The degree of synchronization fluctuation in the time domain. In one embodiment, this coherence analysis is performed by calculating the correlation between two sequences within an analysis window. Normalized cross-correlation or Pearson correlation coefficient within To achieve: ; In the formula, In the first Cross-modal coherence (correlation coefficient) within the analysis window at the end of the frame; This is the index of the current processing frame; To calculate the summation index of frames within the analysis window; The length of the analysis window; for Acoustic leakage characteristics prior to the frame; for Physiological vibration characteristic values ​​prior to the frame; Acoustic leakage characteristic sequence In the present The mean within the window; Physiological vibration characteristic sequence In the present The mean within the window.

[0062] S322: Causal Time-Domain Delay Analysis. Step S322 utilizes a physical prior: the vibration of the mandibular movement must physically precede or be synchronous with the change in ear canal volume caused by that movement. Step S322 finds the peak time-domain delay by calculating the cross-correlation function of the two signals. : ; In the formula, In the first The optimal time delay that maximizes the cross-correlation function within the analysis window at the end of the frame; To find the parameter that will maximize the subsequent expression. The operation; The number of latency frames tested within the preset search range; This is the index of the current processing frame; To calculate the summation index of frames within the analysis window; The length of the analysis window; for Acoustic leakage characteristics prior to the frame; for Before the frame and with an additional delay Physiological vibration characteristic values ​​of the frame.

[0063] The leakage source classification unit is a decision logic unit. It summarizes the feature amplitudes from the feature extraction unit. and coherence from cross-modal analysis units. and causal delay .

[0064] S331: State Decision. This leakage source classification unit compares the above analysis results with a set of preset thresholds in each frame. Output discrete leakage source classification status The decision-making logic is as follows: Define threshold: Acoustic leakage amplitude threshold. This indicates that a significant low-frequency leakage has occurred.

[0065] Cross-modal coherence threshold (e.g., 0.7).

[0066] Physiological causal time window (e.g., [-10,5] frames).

[0067] Decision-making rules: ; In the formula, For the first Classification status of the leakage source of the frame; This is a state of physiological leakage; It is in a passive leakage state; It is in a leak-free state; For the first Acoustic leakage characteristics of the frame; The preset acoustic leakage amplitude threshold is used to determine whether the leakage is significant; For the first Cross-modal coherence of frames; The preset cross-modal coherence threshold is used to determine whether the two are strongly correlated; For the first Optimal frame latency; A preset physiological causal time window is defined as an acceptable delay range (e.g., [-10, 5] frames) to determine whether the delay conforms to the physical causal law. Indicates belonging to; This indicates that it does not belong to this category.

[0068] Status definition: (Physiological leakage): This indicates that a significant low-frequency leakage has been detected, and that the leakage is highly correlated with and causally related to physiological vibrations.

[0069] (Passive leakage): This indicates that a significant low-frequency leakage has been detected, but the leakage is not related to physiological vibrations (such as loosening of glasses or glasses temples) or is not causally consistent in time (such as external impact).

[0070] (No leakage): This indicates that the seal is good and no compensation is required.

[0071] The leak source identification module will ultimately [implement this] The status signal is output to the dynamic compensation state machine and the application module as a control command for switching the algorithm state.

[0072] Please see the appendix Figure 6 , Figure 6 This is a schematic diagram of the internal structure of a dynamic compensation state machine and application module according to an embodiment of the present invention. The dynamic compensation state machine and application module are the final execution unit of the present invention, configured to receive the original target gain from the theoretical compensation calculation module. And the classification status of the leakage source identification module. The core function of this dynamic compensation state machine and application module is: based on... The behavior of the dynamically switched low-frequency compensation algorithm is calculated to determine the final applied smoothing gain. And apply it to audio signals.

[0073] The dynamic compensation state machine and application module may specifically include an algorithm state machine unit and a gain application and reconstruction unit.

[0074] The algorithm state machine unit elevates traditional gain smoothing processing into a multi-state algorithm based on leakage source classification. The operation of this algorithm state machine unit may include the following steps: S410: Definition and initialization of the gain smoother. The state machine unit of this algorithm uses a gain smoother to adjust the original theoretical gain. The smoother is processed to eliminate its transient fluctuations. In one embodiment, the smoother can be a first-order infinite impulse response filter (i.e., an exponential smoother). The general recursive formula for the smoother is as follows: ; In the formula, The time frame index of the currently processed frame; This is the time frame index of the previous processed frame; Calculate the final smoothing compensation gain for the current frame. The previous frame (i.e., the first frame) The final smoothing compensation gain of the output (frame) is calculated, which is the historical state value of the IIR filter; It is the theoretical compensation gain for the current frame. It is the key dynamic smoothing coefficient, whose value is between 0 and 1, and determines the response speed of the smoother.

[0075] This algorithm's state machine unit is pre-defined with at least two sets of parameters. Coefficients: A set of slow coefficients for passive leakage. and a set of rapid coefficients for physiological leakage .in This determines that the smoother has different time constants under different states.

[0076] S411: State transition decision and dynamic smoothing coefficient selection. The state machine unit of this algorithm receives the classification state output by the leakage source identification module. And select the dynamic smoothing coefficient for the current frame based on the classification status. .

[0077] like (Passive Leakage): The state machine switches to the strong correction state. At this time... Set to slow coefficient This slow factor is typically close to 1 (e.g., 0.999 ≤ 1). ≤0.9999). Due to 1- Very small, making right The response to changes is extremely slow, thus achieving stable correction of persistent low-frequency losses.

[0078] like (Physiological leakage): The state machine switches to the artifact suppression state. At this time... Set as fast coefficient The rapidity factor is relatively small (e.g., 0.7 ≤ 0.7). ≤0.9). Due to 1- Larger, smoother The instantaneous peak value has a rapid suppression or smoothing effect, making It does not fluctuate drastically, thus suppressing pump sensation artifacts caused by the algorithm over-responding to physiological activities.

[0079] like (No leakage): At this time Approaching 0 It can be set to Or another attenuation factor. The smoother will It decays smoothly to 0.

[0080] S412: Final gain calculation. The state machine unit of this algorithm is based on the dynamics determined in S411. The final smoothing compensation gain is calculated in real time using a recursive formula. sequence.

[0081] The gain application and reconstruction unit is responsible for applying the calculated gain. The low-frequency portion of the reference audio signal is applied, and the output signal is reconstructed. The operation of this gain application and reconstruction unit may include the following steps: S421: Linear transformation of gain. This gain application and reconstruction unit will provide final smooth compensation gain in the logarithmic domain. Convert back to linear amplitude multipliers : ; In the formula, For the first The linear amplitude multiplier (or linear gain) of the frame. For the algorithm state machine unit in the th The final smoothing compensation gain calculated for each frame.

[0082] S422: Frequency domain compensation application. This gain application relates to the frequency domain representation of the reference audio signal received by the reconstruction unit. and linear multipliers This gain application and reconstruction unit only operates within the preset target low-frequency band. (corresponding to the frequency band index) to Apply this gain within: ; In the formula, For the first Frame, First The complex spectrum value of the frequency domain signal after bandwidth compensation; For the first Frame, First The complex spectrum value of the original reference signal in the frequency band; For the first The linear amplitude multiplier of the frame; This is the index of the currently processed frame; For frequency band index; The preset target low-frequency compensation band The starting frequency band index; The preset target low-frequency compensation band The end frequency band index. For the target low frequency band. Outside of frequency bands, the signal spectrum remains unchanged.

[0083] S423: Time-domain signal reconstruction. This gain application and reconstruction unit reconstructs the compensated frequency-domain signal. Perform an inverse time-frequency transform operation (e.g., inverse short-time Fourier transform), and resynthesize the framed signals using an overlap-add method to obtain the final continuous time-domain output audio signal. .Should It is then sent to the headphone's digital-to-analog converter and speaker for playback.

[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An adaptive compensation adjustment system for audio signals applied to a headphone, characterized in that, The system comprises: a multi-modal sensing module configured to collect an actual acoustic signal and a physiological vibration signal in an ear canal, and receive a reference audio signal; a theoretical compensation calculation module connected to the multi-modal sensing module, configured to calculate a theoretical compensation gain based on the reference audio signal and the actual acoustic signal; a leakage source identification module connected to the multi-modal sensing module and the theoretical compensation calculation module, configured to determine an acoustic leakage feature based on the theoretical compensation gain, determine a physiological vibration feature based on the physiological vibration signal, and perform a cross-modal fusion analysis on the acoustic leakage feature and the physiological vibration feature to output a leakage source classification state characterizing a leakage cause; a dynamic compensation state machine and application module connected to the theoretical compensation calculation module and the leakage source identification module, configured to select different smoothing parameters to smooth the theoretical compensation gain according to the leakage source classification state, to calculate a final compensation gain, and to apply the final compensation gain to the reference audio signal.

2. The adaptive compensation system for audio signals applied to earphones according to claim 1, wherein, The theoretical compensation calculation module is specifically configured to: perform time-frequency transformation on the reference audio signal and the actual acoustic signal; calculate a reference signal low-frequency energy and an actual acoustic signal low-frequency energy in a preset target low-frequency band; calculate a difference between the reference signal low-frequency energy and the actual acoustic signal low-frequency energy in a logarithmic domain to obtain the theoretical compensation gain.

3. The adaptive compensation system for audio signals applied to earphones according to claim 1, wherein, The leakage source identification module is configured to: use a sequence of the theoretical compensation gain as the acoustic leakage feature; perform band-pass filtering on the physiological vibration signal, and calculate a short-time energy of a filtered signal to obtain the physiological vibration feature.

4. The adaptive compensation system for audio signals applied to earphones according to claim 1, wherein, The leakage source identification module performing the cross-modal fusion analysis specifically includes: calculating a cross-modal coherence between the acoustic leakage feature and the physiological vibration feature in an analysis window; calculating a causal time-domain delay between the acoustic leakage feature and the physiological vibration feature.

5. The adaptive compensation system for audio signals applied to earphones according to claim 4, characterized in that, The leakage source identification module is further configured to: compare the cross-modal coherence with a coherence threshold, and compare the causal time-domain delay with a physiological causal time window; when the theoretical compensation gain exceeds an amplitude threshold, and the cross-modal coherence is greater than the coherence threshold, and the causal time-domain delay is within the physiological causal time window, determine the leakage source classification state as a physiological leakage state; when the theoretical compensation gain exceeds the amplitude threshold, but the cross-modal coherence is not greater than the coherence threshold or the causal time-domain delay is not within the physiological causal time window, determine the leakage source classification state as a passive leakage state.

6. The adaptive compensation system for audio signals applied to earphones according to claim 5, wherein, The dynamic compensation state machine and application module is specifically configured to: when the leakage source classification state is the passive leakage state, select a slow coefficient as the smoothing parameter; when the leakage source classification state is the physiological leakage state, select a fast coefficient as the smoothing parameter; wherein the slow coefficient is greater than the fast coefficient.

7. The adaptive compensation system for audio signals applied to earphones according to claim 6, characterized in that, The dynamic compensation state machine and application module smoothes the theoretical compensation gain using a first-order infinite impulse response filter, wherein the smoothing parameter is a smoothing coefficient of the first-order infinite impulse response filter.

8. The adaptive compensation system for audio signals applied to earphones according to claim 1, wherein, The dynamic compensation state machine and application module is further configured to: convert the final compensation gain in the logarithmic domain to a linear amplitude multiplier; in the frequency domain, apply the linear amplitude multiplier to a frequency domain representation of the reference audio signal only within a preset target low frequency band; perform inverse time-frequency transform and overlap-add on the frequency domain signal after applying the gain to reconstruct an output audio signal.

9. The adaptive compensation system for audio signals applied to earphones according to claim 1, wherein, The multi-modal sensing module includes: an acoustic sensing unit, which is a feedback microphone deployed in the earphone, for collecting the actual acoustic signal; a vibration sensing unit, which is an inertial measurement unit, an accelerometer, or a bone conduction sensor, for collecting the physiological vibration signal.

10. An adaptive compensation adjustment method for audio signals applied to a headphone, characterized in that, An application of the adaptive compensation adjustment system for audio signals of an earphone according to any one of claims 1-9, comprising the following steps: collecting the actual acoustic signal and the physiological vibration signal in the ear canal, and receiving a reference audio signal; calculating a theoretical compensation gain based on the reference audio signal and the actual acoustic signal; determining an acoustic leakage characteristic based on the theoretical compensation gain, and determining a physiological vibration characteristic based on the physiological vibration signal; performing cross-modal fusion analysis on the acoustic leakage characteristic and the physiological vibration characteristic to output a leakage source classification state characterizing the leakage cause; selecting different smoothing parameters to smooth the theoretical compensation gain according to the leakage source classification state to calculate a final compensation gain; applying the final compensation gain to the reference audio signal to obtain a compensated output audio signal.