Bone conduction hearing aids and bone conduction hearing aids
By injecting calibration signals into the wearer's skull and collecting response signals, estimating the skull's transmission characteristics and generating calibration parameters, the problem of unstable sound quality in bone conduction hearing devices under different users and wearing conditions is solved, achieving consistency and stability of output sound quality.
Patent Information
- Application Number
- CN202511543877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-28
AI Technical Summary
Existing bone conduction hearing devices cannot accurately reflect the transmission characteristics of the bone conduction path under different users and wearing conditions, resulting in unstable output sound quality.
By injecting calibration signals into the wearer's skull, using reference sensors to collect response signals, estimating the skull's transmission characteristics, and generating calibration parameters to adjust the amplitude and phase of the audio signal, individualized compensation is achieved.
It ensures the consistency and stability of output sound quality under different users and different wearing conditions, avoiding sound quality drift caused by changes in wearing conditions and individual differences.
Smart Images

Figure CN121037758B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hearing aid technology, and more particularly to a bone conduction hearing aid method and a bone conduction hearing aid. Background Technology
[0002] Currently, there are a few technologies related to device self-calibration in the field of bone conduction hearing devices. For example, US Patent No. 10658995B1 (hereinafter referred to as Document 1) discloses a calibration method for bone conduction headphones. It plays a reference tone through an air conduction transducer, then uses a bone conduction transducer to replicate the same frequency, adjusts the volume of the bone conduction output vibration to match the air path output, and simultaneously measures the sound pressure in the ear canal and records the transducer drive voltage to generate an equalization filter for subsequent filtering of bone conduction audio.
[0003] However, the scheme in Reference 1 still has the following shortcomings: the calibration process relies on airborne sound as a reference, the path is indirect, and it cannot truly reflect the transmission characteristics of the bone conduction path itself; the measurement results are easily affected by environmental noise, ear canal shape, and sealing conditions, resulting in poor stability; when the wearing position or contact state changes, the airborne sound reference is difficult to reflect the differences in time, and the sound quality is prone to inconsistency after repeated wearing. Summary of the Invention
[0004] The purpose of this invention is to provide a bone conduction hearing aid method and a bone conduction hearing aid that addresses the shortcomings of existing technologies. The aim is to obtain transmission characteristic information that can truly reflect the bone conduction path under different users and different wearing conditions, and thereby ensure that the output sound quality of the hearing aid device is stable and consistent.
[0005] This invention achieves the above-mentioned objective through the following technical solution: a bone conduction hearing aid method, comprising the following steps: a signal injection step: injecting a preset calibration signal into the skull of the wearer through at least one bone conduction transducer; a response acquisition step: acquiring the response signal formed in the skull by the calibration signal through a reference sensor disposed on the surface of the wearer's skull; a transmission characteristic estimation step: comparing the calibration signal with the response signal to estimate the transmission characteristics of the wearer's skull; a calibration parameter generation step: generating calibration parameters based on the skull transmission characteristics to compensate for individual differences; and a parameter application step: during subsequent audio playback, adjusting the amplitude and / or phase of the audio signal to be played according to the calibration parameters, and outputting it through the bone conduction transducer. In detail:
[0006] Signal injection and response acquisition steps: The processor generates a preset calibration signal. The signal can be a linear chirped signal, a polyphonic signal, or a pseudo-random sequence. The calibration signal is injected into the skull via a bone conduction transducer. A reference sensor acquires the response signal after transmission through the skull. It also performs preprocessing such as noise suppression and DC component removal.
[0007] The linear chirped signal is generated according to the following formula:
[0008] ;
[0009] in, For the first The calibration signal at each sampling point For amplitude coefficient, The starting frequency, For discrete-time sampling index, Sampling rate, For the termination frequency, For the frequency sweep rate and satisfy , For frequency sweep duration, A windowing function for starting and stopping.
[0010] Linear chirping continuously covers a wide frequency band within a short time period, effectively exciting the bone path throughout the target frequency band, facilitating subsequent frequency domain alignment and transfer function estimation. The start and end frequencies cover the main speech frequency band (e.g., 100–4000Hz), and the amplitude uses a cosine window for gradual opening and stopping, avoiding transient pops and improving coherence. Furthermore, compared to schemes using white noise or airborne sound references, chirping provides controllable and highly repeatable injection into the bone path, and is more robust to external noise and subtle wearer movements, improving the estimated SNR and consistency.
[0011] The polyphonic signal is generated according to the following formula:
[0012] ;
[0013] in, It is a continuous-time multi-tone signal. For continuous time variables, The number of polyphonic components, For component index, For the first Each component amplitude, For the first Each component frequency, For the first Each component has an initial phase.
[0014] Multi-tone signals provide high-energy, resolvable excitation peaks at several key frequencies, facilitating accurate amplitude and phase estimation and subband feature construction at these frequencies. Frequency points can be arranged with logarithmic intervals or densely packed into speech-sensitive bands; the amplitude and phase of each component can be optimized to reduce intermodulation and masking. Furthermore, compared to continuous frequency sweeps, multi-tone signals are easier to repeat averaging and self-test, enabling rapid locking of channel variations within specific frequency bands, making them suitable for micro-calibration and rapid updates.
[0015] Frequency domain transformation steps: For subsequent analysis, short-time Fourier transforms are performed on both the calibration signal and the response signal. The calculation formulas are as follows:
[0016] ;
[0017] in, Indicates signal At angular frequency and time shift The time-frequency distribution below, This is used for analysis of the window function. After this processing, the spectral components at discrete frequency points are obtained. and Short-Time Fourier Transform (STFT) decomposes the non-stationary response into locally stable segments in time and frequency, improving robustness over time averaging. Frame length and overlap are set according to sampling rate and real-time requirements (e.g., N=256 or 512, 50% overlap), and the window function uses Hann / Hamming to suppress sidelobe leakage. Furthermore, superior to the coarse-grained statistics of pure whole-segment FFT, STFT can maintain coherence under slight wearing jitter, improving the temporal stability of bone path transmission estimation.
[0018] Power spectrum estimation steps: Based on the framed signal, calculate the power spectrum for the calibration signal and the response signal respectively.
[0019] The calibration power spectrum of the calibration signal is:
[0020] ;
[0021] in, For the first Input calibration power spectrum at each frequency point The number of segments used for spectral averaging. For segmented indexes, For the first The segment calibration signal is in the first The complex spectrum at a frequency point It is a complex conjugate operator. This is a frequency point index. A stable calibration power spectrum is obtained by performing Welch averaging on the injected signal, which serves as a normalization reference and eliminates the influence of random amplitude fluctuations on the transfer function. The number of segments for piecewise averaging is... The settings are configured based on a trade-off between duration and computing power; abnormal segments (such as saturation / abrupt changes) should be removed. Furthermore, this approach is more robust than single-segment energy metrics, significantly reducing variance and thus improving the accuracy of H estimation.
[0022] The cross-power spectrum of the response signal and the calibration signal is as follows:
[0023] ;
[0024] in, For the first The cross-power spectrum of the response at each frequency point and the input. The number of segments used for spectral averaging. For segmented indexes, For the first The segment response signal at the first The complex spectrum at a frequency point For the first The segment calibration signal is in the first Complex conjugate of the frequency spectrum, This is a frequency point index. Cross-spectral analysis extracts the coherent components from the input to the response, suppressing noise and interference components unrelated to the input. It also includes calibration of the power spectrum. Synchronous segmentation and window function; superimposed spectral leakage correction and bias compensation. Furthermore, it outperforms the direct ratio method and exhibits better robustness to uncorrelated noise and mild nonlinearity.
[0025] Transmission characteristic estimation steps: Based on the power spectrum above, the skull transmission characteristics are estimated using the following formula:
[0026] ;
[0027] in, Indicates at frequency point The skull transfer function (discrete transfer function) is as follows. For the first Cross power spectrum at each frequency point For the first Calibration self-power spectrum at each frequency point This is a frequency point index. Directly characterizing the bone path complex frequency response for the current user and wearing method is the sole basis for subsequent compensation. The amplitude and phase curves are smoothed and phase unwrapped; low-confidence frequency points are eliminated. Furthermore, unlike static equalization based on databases or factory averages, this scheme measures the dynamic skull transfer function per person per measurement. High individualization accuracy.
[0028] To improve reliability, the spectral coherence coefficient can be further calculated:
[0029] ;
[0030] in, For the first The spectral coherence coefficients at each frequency point, and their values range from 1 to 2. , For the complex modulus of the cross power spectrum, To input the self-calibrated power spectrum, In response to the self-power spectrum, For frequency point index. Used to determine the reliability of frequency point estimates. Frequency points with insufficient reliability are discarded or weighted. Spectral coherence coefficient. This measure assesses the linear correlation between input and response, used to filter out low-confidence frequencies affected by noise / motion, preventing miscompensation. The threshold can be set between 0.6 and 0.8; it can be combined with the SNR dual criterion. Unlike approaches based solely on amplitude thresholds, the coherence criterion can identify seemingly energetic but uncorrelated interference, improving compensation stability.
[0031] The transmission characteristic estimation step includes a sub-band averaging step: to suppress single-frequency jitter and reduce computational complexity, the frequency is divided into several sub-bands, and the function of the skull transmission characteristics within each sub-band is used. The subband transmission characteristics are obtained by performing a weighted average. This result serves as input for generating subsequent calibration parameters. Subband transmission characteristics. We obtain it from the following formula:
[0032] ;in, For the first The representative value of the complex transmission characteristics of each subband (subband characteristic value). For the first The set of frequency points contained in each subband For the first Weighting factors for each frequency point For the first Skull transfer function at a frequency point For sub-band index, This involves indexing frequency points. Specifically, dense frequency points are mapped to several sub-bands available for engineering use, yielding representative multiple transmission functions for each band, facilitating real-time implementation and storage. Low-frequency sub-bands are denser (emphasizing speech intelligibility ranges); weights can be calculated using the spectral coherence coefficient γ. 2 Or, the importance of speech. Furthermore, unlike the sawtooth spectrum of point-by-point compensation, subbanding brings smooth and stable compensation, which is more suitable for edge implementations with power consumption and latency constraints.
[0033] Calibration parameter generation steps: After obtaining the skull transmission characteristics or sub-band characteristic value Subsequently, this case utilizes optimization methods to generate compensation parameters. As calibration parameters, the compensated joint response satisfies the following equation:
[0034] ;
[0035] in, The target frequency response curve is preset. These are the compensation parameters for the frequency points. This represents the skull transfer function in the continuous frequency domain. It should also be noted that the calibration compensation function in the continuous frequency domain is... During implementation, it is discretized into a table of compensation coefficients at frequency points / subbands. This is for use in subsequent audio compensation. This represents the skull transfer function in the continuous frequency domain. This is a continuous frequency variable. This function is used for theoretical modeling and objective function derivation. Indicates the sampling rate Transform Length Below, at the frequency point The discrete cranial transfer function at a given location is obtained by calculating the power spectrum and solving for it after performing Fourier transforms on the calibration signal and the response signal. The relationship between the two is as follows: .
[0036] Compensation parameters It is obtained by any step of the following weighted least squares optimization method, regularized least squares method, and regularized inverse filtering method:
[0037] The calculation steps of the weighted least squares optimization method are as follows: The objective function is established using the following formula, and the calibration parameters are obtained:
[0038] ;
[0039] The optimal solution of this equation This refers to the compensation parameters at the frequency point. Among them, This is a regularization factor used to suppress excessive gain. For calibration compensation function In the frequency domain, the square of the second norm, In frequency The calibration compensation function is as follows. In frequency The transmission characteristics of the skull below, In frequency The target frequency response curve below, In frequency The weighting factor below, For continuous frequency variables. Specifically, by minimizing and The difference in weights approximates the target response across the entire frequency domain. Reflecting the importance and credibility of the frequency band; target It can be configured as a flat or voice-optimized curve. Furthermore, unlike empirical parameter tuning or fixed EQ, this method takes minimizing the target error as its guiding principle, and has clear optimizability and repeatability.
[0040] Regularized least squares calculation steps: To prevent excessive gain in weak frequency bands, a regularization term is introduced to obtain the calibration parameters:
[0041] ;
[0042] in, This is a regularization factor used to suppress excessive gain. For compensation function In the frequency domain, the square of the second norm, In frequency The compensation function below, In frequency The transmission characteristics of the skull below, In frequency The target frequency response curve below, In frequency The weighting factor below, It is a continuous frequency variable;
[0043] The final solution is:
[0044] ;
[0045] In frequency The calibration compensation function is as follows. for The complex conjugate, The square of the amplitude of the transmission characteristic, As a regularization factor, For continuous frequency variables. Specifically, a regularization term is added to the least squares approach to limit the compensation strength and avoid excessive gain and distortion at path traps. Regularization coefficients It can adapt to either coherence or noise levels; it supports band splitting settings. Furthermore, compared to unconstrained optimization, regularization is more robust in the weak response frequency band, resulting in a more natural sound and higher safety.
[0046] Calculation steps for regularized inverse filtering: Calculate according to the following formula:
[0047] ;
[0048] in, In frequency The compensation function below, for The complex conjugate, The square of the amplitude of the transmission characteristic, As a regularization factor, It is a continuous frequency variable. Specifically, with The generalized inverse is used as an approximate solution for compensation, and... It features a stable small denominator, low computational complexity, and rapid deployment. It can be coupled with band noise / coherence; phase compensation can be limited to control total latency when necessary. Furthermore, compared with global optimization methods, inverse filtering is simpler to implement, consumes fewer resources, and is suitable for low-power devices or rapid micro-calibration.
[0049] Subband closed-form solution calculation steps: Divide the discrete frequency points obtained through Fourier transform into multiple subbands. Within each subband, perform a weighted average of the transmission characteristics of each frequency point to obtain the representative transmission characteristics of that subband, and generate calibration parameters accordingly. Within the subband, the calibration parameters are generated according to the following formula:
[0050] ;
[0051] in, For the first The complex compensation parameters of each sub-band For the first The complex conjugate of the transmission characteristics at each frequency point For the first Complex value of the target frequency response curve at each frequency point For the first The square of the amplitude of the transmission characteristic at each frequency point, For the first Weighting factors for each frequency point For the first The regularization factor of each subband, For the first The set of frequency points of each sub-band For sub-band index, This serves as the frequency point index. Specifically, an explicit closed-form solution is obtained in the subband domain, facilitating interpretation and debugging, and can be directly written into the lookup table parameters. The numerator represents the relevant terms. The denominator is the energy term. It can be used in conjunction with gain and slope constraints for secondary adjustments. Furthermore, compared to numerical iteration, closed-loop solution convergence is more stable and computation is more predictable, which is beneficial for large-scale mass production and rapid on-line calibration.
[0052] The calibration parameter generation process includes engineering constraint steps: to ensure natural sound quality and hardware safety, amplitude range and adjacent sub-band slope constraints are applied to the calibration parameters.
[0053] ;
[0054] ;
[0055] in, This represents the minimum allowable gain (in dB) for a single subband. This represents the maximum allowable gain (in dB) for a single subband. It is a logarithmic function with base 10. For the first The complex modulus of the subband calibration parameters (compensation parameters), For sub-band index, This is the maximum gain difference threshold (in dB) between adjacent sub-bands. Specifically, by limiting the upper and lower limits and the slope of adjacent bands, overcompensation and comb-like spectra are suppressed, ensuring safety and naturalness. The upper limit is, for example, +6dB, the lower limit is, for example, −12dB, and the slope threshold is, for example, 2dB / band, which can be adjusted according to regulations and amplifier capabilities. Furthermore, unlike pure algorithm optimization, engineering constraints conform to auditory and hardware boundaries, ensuring usability and batch consistency. Additionally, limiting the gain change rate of adjacent sub-bands prevents spectral jaggedness and reduces resonance and metallic tones. The threshold is combined with the target curve and hardware response settings; it works in conjunction with the overall gain limit. Therefore, unlike solutions that adjust independently band by band, slope constraints result in a smoother spectrum and more stable cross-scene timbre.
[0056] Parameter application steps: After generating and storing the calibration parameter table, enter normal playback mode. This includes audio input steps, frame segmentation and frequency domain transformation steps, frequency domain compensation steps, inverse transformation steps, and bone conduction output steps.
[0057] Audio input steps: The audio the user wants to listen to is a time-domain signal. .
[0058] Framing and Frequency Domain Transformation Steps: The input audio is framed and transformed to the frequency domain. The system then converts the time-domain signal... Divide the data into multiple frames (each frame lasting 10–20 ms), and perform a Fast Fourier Transform (FFT) on each frame to obtain the result. Frame spectrum .
[0059] Frequency domain compensation step: Parameter table generated based on calibration parameters. (Frequency-by-frequency or sub-band compensation coefficients), parameter table The index corresponds to the frequency point of each frame of the spectrum in the application phase, and the amplitude and / or phase are adjusted for each frequency point:
[0060] ;
[0061] in, For the first The compensated spectrum of the frame For the first The spectrum of a frame; The compensation parameters obtained in the calibration parameter generation step During implementation, it is discretized into a compensation coefficient table at frequency points / sub-bands for audio compensation. (Compensation coefficient table) The compensation parameter table obtained from the calibration parameter generation step has a discrete form derived from the continuous calibration compensation function. The sampling or subband mapping is performed. Specifically, the compensation parameters are applied to the input spectrum at the frame level, and then inversely transformed and synthesized to achieve low-latency real-time pre-equalization. The analysis / synthesis link adopts a unified window and overlap rate to ensure consistency with the frequency division of S3 / S4. Unlike the high-latency scheme of long time-domain FIR, frequency-domain weighting is more computationally efficient and has lower latency, making it easier to implement in wearable devices.
[0062] It should also be noted that: This represents the compensation function generated in the continuous frequency domain, that is, the function generated in the ideal frequency domain. The calibration parameters are obtained through optimization methods, with the goal of making... ,in This is the target frequency response curve. The discrete form of the compensation parameters, i.e., the parameter table, is composed of... At discrete frequency points Upsampling, or mapping from sub-band averaging, is used for compensation of the actual audio signal. The relationship between the two is: .
[0063] Inverse transform step: convert the spectrum of each frame By transforming back to the time domain using the inverse Fourier transform (IFFT), we obtain the first... Frame compensation signal Multiple frames are spliced together using OLA (Optical Overlay) or a window function to form a continuous compensated time-domain signal, resulting in the final compensated signal. :
[0064] ;
[0065] Bone conduction output steps:
[0066] Final compensation signal The sound is converted into mechanical vibrations by a bone conduction transducer driven by a digital-to-analog converter (DAC) and transmitted to the skull, thus achieving a stable and consistent sound quality.
[0067] It also includes a parameter smoothing update step: when the calibration parameters are updated, an interpolation smoothing method is used to avoid abrupt changes in sound quality. The interpolation smoothing uses the following formula:
[0068] ;
[0069] in, For discrete time intervals, For the effective date Subband calibration parameters The first time that took effect at the previous discrete time. Subband calibration parameters For the newly calculated first Subband calibration parameters For time interpolation coefficients ( ), For sub-band index, It uses discrete-time indexing. Specifically, the old and new parameters are interpolated and transitioned to avoid skipping and abrupt changes in timbre caused by instantaneous switching. The transition time of 100–300ms is perceptibly seamless; it can be linked to the listening threshold (mute or low-energy band switching). Therefore, compared to direct replacement, smooth updates significantly reduce the risk of abrupt changes in listening experience and improve wearing comfort.
[0070] Furthermore, the phase is unwrapped and calculated according to the following formula between the transmission characteristic estimation step and the calibration parameter generation step: ,in, For the first The unwrapping phase at each frequency point For phase unwrapping function, For transfer function The phase operator results, For frequency point index.
[0071] Furthermore, the group delay is estimated and calculated between the transmission characteristic estimation step and the calibration parameter generation step using the following formula: ,in, For the first Group delay at each frequency point For discrete difference operators, For the first Phase at each frequency point For the first angular frequency at each frequency point For frequency point index.
[0072] Furthermore, the frequency weights in the transmission characteristic estimation step are set according to the following formula: ,in, For the first The weight of each frequency point The normalization constant is For the first Spectral coherence coefficient at each frequency point For the first The speech importance function value at each frequency point For frequency point index.
[0073] A bone conduction hearing aid, comprising:
[0074] Bone conduction transducers are used to inject preset calibration signals into the wearer's skull and play adjusted audio signals.
[0075] A reference sensor is used to acquire the response signal of the calibration signal in the skull; and
[0076] The processor, electrically connected to the bone conduction transducer and the reference sensor, is configured to:
[0077] The calibration signal is compared with the response signal to estimate the skull transmission characteristics of the wearer;
[0078] Calibration parameters are generated based on the skull transmission characteristics to compensate for individual differences;
[0079] During audio playback, the amplitude and / or phase of the audio signal to be played are adjusted according to the calibration parameters and output to the bone conduction transducer.
[0080] The beneficial effects of this invention are:
[0081] In the signal injection and response acquisition steps, a preset calibration signal is injected through a bone conduction transducer, and the response in the skull is directly acquired by a reference sensor, avoiding reliance on airborne acoustic reference and obtaining real feedback that is strongly correlated with the bone conduction path.
[0082] In the transmission characteristic estimation step, the calibration signal is compared with the response signal to obtain the skull transmission characteristics. This can accurately capture the influence of changes in wearing position or contact state on signal transmission, avoid the instability of the sound curve due to changes in wearing method and individual differences, and reduce the influence of sound quality drift.
[0083] In the calibration parameter generation and parameter application steps, individualized parameters are generated based on transmission characteristics and applied to the subsequent audio playback path, ensuring that the output sound quality remains consistent between different users and when the same user wears the device repeatedly.
[0084] In detail, to address the issue of inconsistent sound quality among different users and under different wearing conditions in existing technologies, the specific implementation process is as follows:
[0085] First, when the device is powered on or re-worn, the bone conduction transducer injects a preset calibration signal into the wearer's skull. This signal can be a chirped signal, a polyphonic signal, or a pulse sequence, used to cover the dominant speech frequency band required for hearing aids. By calibrating the signal directly on the bone conduction path, the uncertainty of the airborne sound path can be avoided, providing a stable input reference for subsequent calculations.
[0086] Secondly, the device acquires the response signals generated in the skull through a reference sensor. This reference sensor can be a bone surface accelerometer, a structural vibration pickup, or a shell vibration microphone. By directly acquiring the skull vibration response, the device accurately reflects the bone conduction transmission characteristics of the current user and the current wearing state, avoiding deviations caused by ear canal structure, seal, or external noise interference.
[0087] The processor then compares the injected calibration signal with the acquired response signal to estimate the transmission characteristics of the skull. These transmission characteristics can be the amplitude and phase information of a frequency band, or simplified into a set of filtering parameters using mathematical algorithms. This process accurately reflects individual differences between users, as well as variations in transmission characteristics caused by different wearing positions and contact pressures for the same user.
[0088] After obtaining the transmission characteristics, the processor generates corresponding calibration parameters based on the preset target frequency response curve. These calibration parameters can exist in the form of bandwidth gain or phase compensation tables, and are used to pre-correct subsequent audio signals. Through this step, automatic compensation can be performed for individual differences and wearing variations, thereby ensuring the consistency of output sound quality.
[0089] Finally, during audio playback, the processor adjusts the input audio signal in real time according to the calibration parameters and outputs the corrected signal to the bone conduction transducer. This ensures that the sound heard by the user is consistent with the target frequency response curve, without significant differences due to variations in wearing conditions or individual differences.
[0090] Through the above technical solution, the present invention can automatically acquire transmission characteristic information that truly reflects the bone conduction path under different users and different wearing conditions, and convert this information into applicable calibration parameters, thereby maintaining stable and consistent sound quality during multiple wears. Attached Figure Description
[0091] Figure 1 This is a schematic flowchart of the bone conduction hearing aid method described in this invention.
[0092] Figure 2 This is a schematic block diagram of the hardware structure of the bone conduction hearing aid described in this invention. Detailed Implementation
[0093] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It is understood that the accompanying drawings are provided for reference and illustration only, and are not intended to limit the present invention. The connection relationships shown in the accompanying drawings are only for clear description and do not limit the connection method.
[0094] like Figures 1-2 As shown, in one embodiment of the present invention, the bone conduction hearing aid method includes a complete self-calibration process, the process of which is as follows:
[0095] First, upon powering on the device or detecting re-wearing, the bone conduction transducer injects a preset calibration signal into the wearer's skull. This calibration signal can be a chirped signal, a multi-tone combination signal, or a pseudo-random sequence covering the main speech frequency band, typically ranging from 100Hz to 4000Hz, and lasting approximately 2–5 seconds. By injecting this known signal directly into the bone conduction path, subsequent estimations can be ensured to be referenced to a reliable input, thus avoiding uncertainties introduced by airborne sound references.
[0096] Secondly, the reference sensor configured in the device collects the response signal formed in the bone tissue by the calibration signal on the skull surface. The reference sensor can be a miniature accelerometer, a vibration pickup, or a high-sensitivity microphone attached to the outer shell, preferably with a bandwidth of at least 3kHz, and data is collected synchronously via a high-sampling-rate analog-to-digital converter. By directly acquiring the bone conduction response, the skull transmission status between different users or under different wearing conditions for the same user can be accurately reflected.
[0097] The processor then compares the calibration signal with the response signal to estimate the skull transmission characteristics under the current wearing condition. Specifically, the system can use Fast Fourier Transform to perform frequency domain analysis on the two signals and obtain the transfer function by the ratio of the cross power spectrum to the calibration autopower spectrum. The obtained transmission characteristics are typically represented as amplitude and phase curves at different frequencies. To improve stability, the processor also performs smoothing, outlier removal, and banding processing on the estimation results, ultimately obtaining parameter data for 16 to 24 frequency bands.
[0098] After obtaining the transmission characteristics, the processor generates corresponding calibration parameters based on the target frequency response curve. The target frequency response curve can be a flat response, a speech enhancement curve, or a personalized curve incorporating hearing loss compensation. The calibration parameters can be generated using a weighted least squares optimization method or a regularized inverse filtering method. For example, in the regularized inverse filtering implementation, the calibration compensation function G(f) can be given by the following equation:
[0099] ;
[0100] Where G(f) is the calibration compensation function at frequency f, H(f) is the estimated cranial transfer function, H*(f) is its complex conjugate, and |H(f)| 2 The amplitude is squared, and λ is the regularization factor. This method effectively avoids excessive gain in weak response bands, thereby improving the robustness of the algorithm. The generated calibration parameters are stored in the device in the form of a band gain and / or phase compensation table for subsequent audio processing.
[0101] Finally, during audio playback, the processor adjusts the input audio signal in real time according to the calibration parameters. Specifically, the audio signal is segmented into frames, subjected to frequency domain transformation, and the amplitude and phase of each frequency band are adjusted using a lookup table. Then, an inverse transformation is performed, and the signals are superimposed for output. To ensure a smooth user experience, the system employs a smooth transition strategy when switching between old and new parameters, avoiding abrupt changes in sound quality. After this processing, the audio signal is transmitted to the user's skull via a bone conduction transducer, ensuring consistent and stable output sound quality for different users or under different wearing conditions.
[0102] A bone conduction hearing aid method includes: a signal injection step, a response acquisition step, a transmission characteristic estimation step, a transmission characteristic estimation step, and a calibration parameter generation step. Wherein:
[0103] Signal injection procedure: A preset calibration signal is injected into the wearer's skull via at least one bone conduction transducer. Specifically, when the device is powered on or a change in wearing status is detected, the bone conduction transducer injects the preset calibration signal into the wearer's skull. Regarding the signal type, it can be a linear chirp (continuous frequency sweep from low to high frequencies), a multi-tone combination (simultaneous superposition of several fixed frequencies), or a pseudo-random binary sequence. Regarding the frequency range, it is recommended to cover a speech bandwidth of 100Hz to 4000Hz. Regarding the signal length and amplitude, it should last 2–5 seconds, with the amplitude controlled below a comfort threshold, such as an RMS not exceeding 80dB bone conduction sound pressure level. Furthermore, directly calibrating the signal in the bone pathway reduces airborne noise interference and ensures controllability at the input end.
[0104] Response Acquisition Steps: The response signal formed within the skull by the calibration signal is acquired using a reference sensor placed on the surface of the wearer's skull. Specifically, during signal injection, the reference sensor (e.g., a miniature accelerometer attached to the skull surface) acquires the vibration response within the bone. A sampling rate of 24 kHz and a resolution of 24 bits are recommended. For preprocessing, a high-pass filter is used to remove DC bias before signal acquisition, and a low-pass filter is used to remove noise above 5 kHz. For quality assessment, the signal-to-noise ratio (SNR) is calculated to ensure that the SNR is ≥ 15 dB within the effective signal bandwidth; otherwise, acquisition is repeated. This ensures that the acquired signal reflects the true bone path, avoiding uncertainties introduced by airborne acoustic paths.
[0105] Transmission characteristic estimation step: The calibration signal is compared with the response signal to estimate the skull transmission characteristics of the wearer. Specifically, the processor compares the known calibration signal with the reference sensor response to estimate the skull transmission characteristics. This is specifically achieved using cross-power spectroscopy. With calibration self-power spectrum The calculation formula and method are based on the following ratio (i.e., cross-power spectrum). With calibration self-power spectrum The ratio between them), to obtain the user's skull in the first Transmission characteristics at each frequency point The ratio formula is as follows:
[0106] ;
[0107] in, Let f be the skull transfer function at frequency f. To determine the cross-power spectral density of the response signal and the calibration signal, The calibration power spectral density is used to calibrate the signal. Further, for post-processing, [the following is done]... The amplitude and phase curves are smoothed, unreliable frequency points are removed, and the spectrum is divided into 16–24 sub-bands to facilitate subsequent parameter generation. Specifically, this is achieved through spectral coherence coefficients. The reliability of the frequency data is assessed to ensure that points affected by noise or motion interference are removed. Finally, the reliable frequency points are aggregated into sub-band characteristic values. This facilitates the subsequent implementation of low power consumption.
[0108] Cross power spectrum With calibration self-power spectrum The calculation formula is as follows:
[0109] Calibration formula for power spectrum: ;in, For the first Input calibration power spectrum at each frequency point The number of segments used for spectral averaging. For segmented indexes, For the first The segment calibration signal is in the first The complex spectrum at a frequency point It is a complex conjugate operator. Frequency point index;
[0110] Cross-power spectrum formula: ;in, For the first The cross-power spectrum of the response at each frequency point and the input. The number of segments used for spectral averaging. For segmented indexes, For the first The segment response signal at the first The complex spectrum at a frequency point For the first The segment calibration signal is in the first Complex conjugate of the frequency spectrum, For frequency point index.
[0111] Calibration parameter generation step: Based on the skull transmission characteristics, calibration parameters are generated to compensate for individual differences. Specifically, the skull transmission characteristics of the current user and wearing state have been obtained in the transmission characteristic estimation step. (Complex frequency response, including amplitude and phase); The target frequency response curve is pre-set in product design or customized configuration. (This can be a flat response or a curve conforming to speech enhancement principles); simultaneously, weighting factors are set based on the reliability and importance of each frequency. The task of this step is to determine the calibration compensation function / filter. This enables the skull to transmit signals. × Generate compensation parameters The joint response should approximate the target frequency response curve as closely as possible. The relationship to achieve the design goal is as follows, and its mathematical essence is to find a calibration compensation function. , that is, the calibration parameters, such that:
[0112] ;
[0113] in, This is a preset target frequency response curve, which can be a flat curve or a curve that enhances speech clarity; the obtained... Discretized sub-band gain / phase compensation tables (e.g., 16–24 frequency bands, one complex coefficient per band) serve as lookup data for parameter application steps. Pre-compensation is achieved by multiplying the input audio band-by-band during real-time playback. This is to ensure the actual output of the hearing aid... Approaching the expected target frequency response curve This case proposes three algorithms, which are computed by the following three solvers:
[0114] Solver 1: Weighted Least Squares Optimization (Full-Frequency Approximation of the Target under High-Confidence Data)
[0115] ;
[0116] in, In frequency The compensation weights (i.e., the frequency domain form of the calibration parameters) are used. As a weighting factor, This refers to the skull's transmission characteristics (bone conduction pathway characteristics). The target frequency response curve. This is for complex number modular arithmetic; min indicates that the solution that minimizes the error is required.
[0117] Solver with skull transport characteristics × Generate compensation parameters Target frequency response curve Minimizing the weighted error energy directly produces frequency-dependent compensation weights. Finally, the first calibration parameter can be obtained:
[0118] ;
[0119] in, In frequency The compensation weight below, for The complex conjugate, For power terms, This represents the target frequency response curve. The final fractional formula obtained is used to find the most suitable compensation adjustment value, i.e., the compensation weight. This cancels out bone distortion, making the final sound closer to the target frequency response curve. When the skull transmits signals... The estimation quality is high (coherent, good SNR) and we hope to achieve a uniform approximation across the entire frequency band. In this case, solver one should be used first.
[0120] Solver 2: Regularized Least Squares (Suppressing Gain Boom in Weak Response / Noise Bands)
[0121] ;
[0122] The second calibration parameter was finally obtained:
[0123] ;
[0124] in, As a regularization factor, The squared L2 norm of the compensation weights (sum of squared weights) is used, with the other symbols remaining the same. This formula maintains robustness in the trapped frequency band, preventing sudden changes in volume. This formula has an additional term. Add a strength penalty term to the least squares method. This is used to prevent excessive amplification in weak frequency bands, that is, to prevent... Excessive amplification in very small frequency bands can lead to distortion, howling, or increased power consumption. When there are channel dropout bands, motion artifacts, or high ambient noise, solver two should be used preferentially.
[0125] Solver 3: The third calibration parameter is obtained by regularized inverse filtering (a generalized inverse approximation with low complexity, suitable for fast implementation at the end).
[0126] ;
[0127] in, In frequency The calibration compensation function is as follows. for The complex conjugate, This is the power term (amplitude squared). This is the regularization factor. The formula directly gives the compensation weight. Closed-form approximate calibration compensation function ( This is equivalent to the transmission characteristics of the skull. The generalized inverse of the stabilized model is computationally simple and parallel-friendly. When the end-side computing power / latency budget is tight, and rapid micro-calibration or frequent re-wearing adaptation is required, solver three is preferred because it has low computational cost and is suitable for real-time implementation at the hearing aid end.
[0128] Ultimately, the output of these solver algorithms is a subband compensation table (such as a subband gain / phase compensation table for 16–24 frequency bands), with each frequency band storing compensation parameters, such as the gain or phase parameters for 16–24 frequency bands.
[0129] Subband Implementation and Engineering Constraints (Transforming Formula Solutions into Usable Parameter Tables): In actual implementation, continuous frequencies are divided into... Each sub-band (denserly distributed at lower frequencies) has a compensation coefficient calculated separately. (or by) Take representative values within the band and apply engineering boundaries: upper / lower limits for single-band gain (e.g., +6dB / −12dB), maximum slope of adjacent sub-bands (e.g., 2dB / band), and phase / group delay limits (to control the overall system delay). The resulting sub-band gain / phase compensation table satisfies both the optimization objective and the auditory and hardware safety boundaries, making it easy to apply in real-time in S5 using an analysis → weighting → synthesis process. Ultimately, it brings the bone path differences between different users and under different wearing conditions back to a unified target sound quality.
[0130] Parameter application steps: During subsequent audio playback, the amplitude and / or phase of the audio signal to be played are adjusted according to the calibration parameters, and output through the bone conduction transducer. Specifically, during subsequent audio playback, the processor adjusts the input audio in real time according to the calibration parameters. Details are as follows:
[0131] S1. Divide the audio signal into frames (frame length such as 256 points, 50% overlap).
[0132] S2. Perform a Fast Fourier Transform (FFT) on each frame;
[0133] S3. Adjust the amplitude and / or phase of each frequency band according to the compensation table;
[0134] S4. Perform an inverse IFFT on the adjusted signal and then sum them to obtain a continuous output.
[0135] S5, output to bone conduction transducer.
[0136] It also includes a parameter update step: when a change in wearing conditions or periodic triggering is detected, the above calibration process is rerun. A smooth transition strategy is used when writing new parameters, and the interpolation transition is completed within 100–300ms to avoid abrupt changes in sound quality.
[0137] In addition, an output limiter can be added to ensure that the average output over a long period of time does not exceed the hearing safety threshold, while a temperature sensor is used to control power consumption and temperature rise.
[0138] This invention can re-estimate the bone conduction path and generate compensation parameters for different users, when the same user wears the device repeatedly, and under different contact pressures, ensuring stable and consistent sound quality. In particular, the mathematical optimization and regularization design in the calibration parameter generation step enable the algorithm to both closely match the target response and avoid overcompensation in weak frequency bands, thus balancing sound quality and stability.
[0139] In another embodiment: the reference sensor includes an eye-tracking acquisition unit for acquiring eye-tracking response data triggered by the calibration signal before the calibration parameter generation step, so as to time-align the eye-tracking response data with the calibration signal and use it as an auxiliary basis for estimating the skull transmission characteristics;
[0140] The eye movement acquisition unit is a camera-type eye tracker or an electrophysiological electrooculogram sensor, which records pupil position, corneal reflex, eyelid state or horizontal and vertical eye movement potential difference at a sampling rate not lower than a preset rate, and synchronously records the trigger time stamp of the calibration signal in the response acquisition step to ensure a one-to-one correspondence between the two in the time domain.
[0141] The transmission characteristic estimation step further includes: performing quality assessment and artifact removal on the eye movement response data, extracting at least one feature from the eye movement response data, such as micro-eye movement amplitude, velocity peak, pupil micro-change, periocular micro-muscle activity proxy index, or saccade incidence rate, and generating a confidence weight based on the stability and repeatability of the features extracted from the eye movement response data, which is used for weighted processing in the process of obtaining representative transmission characteristics at the sub-band level.
[0142] The sub-band-level representative transmission characteristics are obtained by taking the characteristics of the eye-tracking response data and the energy and timing relationship of the calibration signal in each sub-band as input, and combining the sub-band weights determined by the confidence weights to generate sub-band-level representative transmission characteristics that can be used in the calibration parameter generation step.
[0143] The calibration parameter generation step outputs the complex compensation parameters for each sub-band and sets constraints on the amplitude range of the complex compensation parameters and the slope of change between adjacent sub-bands to ensure the stability of the output sound quality and the naturalness of the listening experience in the parameter application step.
[0144] The parameter application step involves performing frequency domain analysis on each frame of the input audio, mapping and applying the complex compensation parameters corresponding to the frequency points or sub-bands, and using a time-smooth transition strategy to perform a weighted transition between the old and new parameters when the complex compensation parameters are updated to avoid sudden changes in output sound quality.
[0145] The signal injection step employs low-amplitude, short-duration, and imperceptible calibration micro-vibration to inject calibration signals. Before injection, a safety threshold check is performed to ensure that the driving intensity of the bone conduction transducer does not exceed a preset threshold. During the injection process, overlap with audible playback is avoided to reduce interference to the user.
[0146] A bone conduction hearing aid includes a bone conduction transducer, a reference sensor, and a processor. The bone conduction transducer injects a preset calibration signal into the wearer's skull and plays an adjusted audio signal. The reference sensor collects the response signal of the calibration signal in the skull and includes an eye-tracking acquisition unit for acquiring eye-tracking response data triggered by calibration micro-vibrations. The processor is electrically connected to the bone conduction transducer and the reference sensor and is configured to: compare the calibration signal with the response signal to estimate the wearer's skull transmission characteristics; generate complex compensation parameters (also called calibration parameters) based on the skull transmission characteristics to compensate for individual differences; adjust the amplitude and / or phase of the audio signal to be played according to the complex compensation parameters during audio playback and output the adjustment to the bone conduction transducer; simultaneously perform framing, time alignment, feature extraction, and quality assessment on the eye-tracking response data, and use the processing results as input to generate the complex compensation parameters to achieve personalized automatic calibration based on eye-tracking micro-response. (Details follow.)
[0147] Currently, most hearing aids require a visit to a hospital or retail store where a technician conducts a hearing test in a soundproof environment and then makes adjustments based on subjective feedback. This process is time-consuming, has poor user compliance, and even slight changes in wearing position can render the original settings unsuitable, necessitating a return to the store for readjustment. To address the pain point of requiring a hospital / store visit for fitting, we introduce an objective self-calibration pathway that can be performed in everyday situations, while maintaining the existing bone conduction hearing aid process:
[0148] The hearing aid delivers a subtle, imperceptible calibration micro-vibration to the skull during a quiet moment. Simultaneously, an eye-tracking unit in a reference sensor observes the subtle eye movements and periocular micromuscular responses triggered by this micro-vibration. These objective responses are time-aligned with the stimulus, serving as an auxiliary basis for estimating individual bone conduction characteristics. In this way, users do not need to enter a soundproof room or answer questions or raise their hands; the device can automatically obtain personalized information at home equivalent to a hearing test at a hospital, and then automatically generate and apply compensation settings.
[0149] To ensure that this at-home self-calibration is truly usable, rather than a one-off guesswork, eye-tracking data isn't used directly. The device first performs a quality assessment on each eye-tracking response: blinks, occlusions, rapid head turns, and low-light frames are labeled and downweighted or even eliminated. Then, from the remaining valid segments, it extracts elements that stably reflect stimulus responses, such as the amplitude of slight eye movement, the fastest possible movement speed, the delay between stimulus and response, subtle pupillary constriction changes, the amount of micromuscular activity in the periocular region, and the frequency of saccades. The system also examines whether these elements are consistent after repeating the same type of stimulus several times, assigning higher confidence weights to segments with good consistency. Through this step, video or electrical signals that are easily contaminated by everyday actions are purified into a set of usable and repeatable objective indicators.
[0150] With these reliable eye-tracking data points, the processor doesn't simply provide a crude conclusion of overall volume being higher or lower. Instead, it maps the energy and timing of each data point and calibration signal across different frequency ranges, merging them into several sub-bands. This results in an input quantity for each sub-band that represents the user's bone conduction pathway characteristics. Subsequently, the system calculates a complex compensation parameter for each sub-band—combining amplitude and phase / delay corrections—truly adding where needed, subtracting where necessary, and aligning the timing. To avoid the side effect of over-amplifying adjacent bands due to automatically adjusted settings at home, the calculated parameters are subject to engineering constraints: each sub-band has upper and lower limits for intensity, and the slope of change between adjacent sub-bands is limited to a comfortable range. The parameters don't change instantaneously but rather through a smooth transition, allowing the old and new parameters to connect naturally within a short time, without any abrupt jumps or metallic sounds.
[0151] The entire process corresponds one-to-one with the original method's main chain: still injecting calibration signals—acquiring responses—estimating transmission characteristics—generating compensation parameters—applying to the playback link. The only difference is in the implementation of the reference sensor; in addition to the bone surface accelerometer or reference microphone, an eye-tracking acquisition unit is added as an objective observation channel. The two channels can either replace each other or be integrated in parallel: when the home environment is noisy, eye movements are not sensitive to sound noise and can still provide a physiological proxy response strongly correlated with the stimulus; when there is insufficient light or the camera is obstructed, the signal can fall back to bone surface sensing to ensure uninterrupted self-calibration. This multi-source evidence collection approach makes automatic at-home adaptation a viable alternative to an ideal.
[0152] Compared to going to a hospital / store, the direct benefits of this approach are obvious: users don't need appointments, dedicated personnel, or subjective cooperation; the device quickly completes a minor calibration in a quiet moment without disturbing daily use; sound quality drift caused by changes in wearing position, tightness, or environment can be automatically detected and corrected step by step by the system, stabilizing the output at a comfortable level over a long period. More importantly, all of this is based on objective response, not on the user's subjective description of whether it sounds loud or not, thus resulting in better repeatability and consistency, and making it easier to demonstrate its technical effectiveness and engineering feasibility during licensing.
[0153] In summary, by incorporating the eye-tracking acquisition unit into the reference sensor system, an objective response source is provided that does not require attachment to the bone surface, directly addressing the pain point of having to go to a hospital / store from the perspectives of scenario and compliance. Through quality assessment, feature extraction, and weighted aggregation, raw eye-tracking data from everyday environments is transformed into stable sub-band level input, ensuring reliable individual characteristics even under home conditions. By reliably implementing the input into constrained and smoothly updatable complex compensation parameters, overcompensation and abrupt changes in sound quality are avoided, ensuring a natural final hearing experience. With these three elements working together, the hearing aid has the ability to automatically tailor a solution for you, substantially reducing reliance on hospitals / stores while maintaining stable and consistent sound quality performance in daily use.
[0154] To avoid requiring a separate pair of glasses, this embodiment prioritizes an integrated form factor for eyeglass-style bone conduction hearing aids: the left and right temples already house the bone conduction transducers, while eye-tracking acquisition units (miniature near-infrared cameras or EOG electrodes) are integrated into the left and right temples or bridge of the same frame. These units share power and clock with the internal processor, resulting in the shortest signal routing, most stable synchronization, a natural appearance, and no compromise on wearing comfort. To facilitate stable acquisition of eye-tracking response data without altering the existing shape and wearing experience, the acquisition units are preferentially positioned near the frame and temples. Practice has shown that the area near the bridge of the nose and the inner front of the temples is closer to the eyeball, resulting in more stable relative posture, and is also beneficial for mass production assembly and concealment.
[0155] In a preferred embodiment, a miniature near-infrared camera is embedded between the nose pads or on the inner side of the upper edge of the nose bridge. The lens is slightly tilted upwards or inwards, operates in the 850–940nm wavelength band, and is equipped with low-power infrared illumination. The frame rate is no less than 120 frames per second to meet the temporal resolution requirements for minute eye movements, and the illumination intensity meets the eye exposure limits. The camera shares a clock with the hearing aid's digital signal processor or is marked with a unified timestamp to ensure a one-to-one correspondence with the trigger time of calibration micro-vibrations. A six-axis inertial sensor can be added to the inner side of the frame to subtract the influence of head micro-movements in post-processing, avoiding misinterpreting head movements as eye movements. When the overall shape or style of the device restricts the nose bridge module, a small camera can be placed on the inner front of each temple, with the lens facing inwards to cover the area from the inner canthus to the pupil, and the field of view can be moderate. If eyelashes obstruct the view or the lens reflects light, the other side can maintain tracking. For sports models or narrow-bezel models, the camera can also be embedded on the inner side of the lower edge of the frame to observe upwards, and a safe distance from the cheek and eyelashes can be reserved in the structure to reduce obstruction. The aforementioned camera-based path only extracts features such as pupil center, eyelid opening and closing, and corneal reflection locally, and immediately discards the original image without long-term storage to meet privacy compliance and low power consumption requirements. In scenarios with poor lighting or when wearing tinted lenses, electrophysiological electrooculography (EOG) can serve as a low-power supplement: dry electrodes in the horizontal channel are positioned on the inner side of the left and right temples near the outer canthus; vertical channel electrodes are located at the brow bone (inner side of the upper edge of the frame) and below the lower eyelid (near the nose pad or inner side of the lower edge); reference electrodes are located on the inner side of the nose pad, behind the ear, on the mastoid process, or in the middle of the temple. The electrode spacing should be 2–3 cm. High input impedance amplification is used at the front end, with sampling at no less than 250 Hz. Baseline drift correction and blink artifact detection are implemented in the firmware. The temple electrodes use spring contacts or flexible conductive rubber to adapt to different face shapes and avoid poor contact caused by hair or cosmetics. For different product shapes, equivalent deformation is also allowed without changing the technical essence.
[0156] Firstly, when the main unit is an ear-hook or headband bone conduction hearing aid, a lightweight frame or clip-on that only carries the eye-tracking acquisition unit can be fitted. It is connected to the main unit via a short ribbon cable, spring contact, or low-power wireless connection. The timestamp synchronization is still controlled by the main unit. It looks like a fake frame, but in the system, it belongs to the same hearing aid device as the sensing accessory.
[0157] Secondly, a camera or EOG electrode can be built into only one temple, while the other side only houses the transducer or serves as redundancy. The software automatically selects the effective side based on the quality score, thus reducing structural complexity and power consumption. The acquisition strategy does not require any one channel to always be active; instead, it adaptively switches based on the scene: when the pupil is clearly visible, the camera on the bridge of the nose or the inside of the temple is prioritized; in strong backlight, obstruction, or low light, the camera weight is reduced and the EOG weight is activated or increased. When both channels are available simultaneously, they are weighted and fused according to the real-time quality score. The calibration micro-vibration injection window is selected after a period of silence or low noise, and begins only after the inertial sensor determines that the head posture is stable. Before injection, drive strength and temperature rise threshold checks are performed. During injection, it is mutually exclusive or downgraded with audible playback to ensure the user is unaware of the vibration.
[0158] Through the aforementioned integrated setup and adaptive strategy, eye-tracking response data can be stably obtained without altering the wearing experience: when conditions permit, high-precision pupil and eyelid features are provided by a near-infrared camera on the bridge of the nose or the inner side of the temples; when the scene is unfavorable, the EOG seamlessly takes over. Both complete time alignment, quality scoring, and feature output under the same processing and clock system, thereby reliably transforming imperceptible small stimuli into objectively observable small responses. This provides consistent and repeatable basic data for subsequent subband input construction and complex compensation parameter generation, giving the automatically customized at-home self-calibration path a clear location and feasibility for implementation.
[0159] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A bone conduction hearing aid method, characterized by, The method comprises the following steps: a signal injection step of injecting a preset calibration signal into the skull of the wearer through at least one bone conduction transducer; a response acquisition step of acquiring a response signal formed in the skull by the calibration signal through a reference sensor arranged on the surface of the skull of the wearer; a transmission characteristic estimation step of comparing the calibration signal with the response signal to estimate the skull transmission characteristic of the wearer; a calibration parameter generation step of generating a calibration parameter for compensating for individual differences based on the skull transmission characteristic; a parameter application step of adjusting the amplitude and / or phase of an audio signal to be played according to the calibration parameter and outputting the audio signal through the bone conduction transducer; The transmission characteristic estimation step comprises: performing energy averaging on the spectrum of the calibration signal to obtain a calibration autocorrelation spectrum of the calibration signal; multiplying the spectrum of the response signal with the spectrum of the calibration signal frame by frame and performing averaging to obtain a cross-correlation spectrum; and dividing the cross-correlation spectrum by the input calibration autocorrelation spectrum to obtain a skull transmission function representing the amplitude and phase characteristics of the skull path; performing energy averaging on the spectrum of the response signal to obtain a response autocorrelation spectrum of the response signal; calculating a spectral coherence coefficient of each frequency point based on the energy proportion relationship among the calibration autocorrelation spectrum, the response autocorrelation spectrum and the cross-correlation spectrum, and using the spectral coherence coefficient to judge the reliability of the skull transmission function at different frequency points; dividing adjacent frequency points into a plurality of sub-bands according to a preset rule, performing weighted summation on the skull transmission function in each sub-band range to obtain a sub-band level representative transmission characteristic, and using the sub-band level representative transmission characteristic as the input of the calibration parameter generation step. The calibration signal is a linear chirp signal or a multi-tone signal, the linear chirp signal linearly changes from a starting frequency to an ending frequency within a set time length, and the multi-tone signal is composed of sine components of different frequencies, and the amplitudes and initial phases of the components are set; 2. The bone conduction hearing assistance method of claim 1, wherein, The response acquisition step further comprises: performing frame division and windowing processing on the calibration signal and the response signal, and performing frequency domain transformation in each frame to obtain a group of discrete frequency spectrum components under a predetermined sampling rate and transformation length, so that a group of frequency points corresponding to each other are formed, and the input and the output have a corresponding relationship at the frequency points, which are used for estimating the skull transmission characteristic. The calibration parameter generation step comprises: setting a target frequency response curve, and using an optimization solving strategy including a weighted least square method, a regularized least square method and a regularized inverse filtering method to generate a compensation parameter, so that the joint response of the skull transmission characteristic and the compensation parameter is close to the target frequency response curve; in the sub-band level implementation, the sub-band level representative transmission characteristic is used as the input, and the spectral coherence coefficient is used to determine the sub-band weight to solve the complex compensation parameter of each sub-band, and the amplitude upper and lower limit constraints and the change slope limit between adjacent sub-bands of the complex compensation parameter are set to ensure the stability and naturalness of the output signal.
3. The bone conduction hearing assistance method of claim 2, wherein, 4. The bone conduction hearing assistance method of claim 3, wherein, The parameter application step includes: performing frequency domain analysis on the input audio frame by frame, and mapping and applying the complex compensation parameters corresponding to each frequency point or each sub-band to complete frequency domain multiplication processing, and then performing inverse transformation from frequency domain to time domain and inter-frame overlap synthesis to reconstruct a time domain compensation signal and output to the bone conduction transducer; and when the complex compensation parameters need to be updated online, the old and new complex compensation parameters are weighted and transitioned through time smoothing transition to avoid sudden changes in output sound quality.
5. The bone conduction hearing assistance method of claim 4, wherein, The reference sensor includes an eye movement acquisition unit, which is used to acquire eye movement response data triggered by the calibration signal before the calibration parameter generation step, so as to time-align the eye movement response data with the calibration signal and use the eye movement response data as an auxiliary basis for estimating the skull transmission characteristic; The eye movement acquisition unit is a video eye tracker or an electro-physiological electro-oculogram sensor, and records the pupil position, corneal reflection, eyelid state or horizontal and vertical eye movement potential difference at a sampling rate not less than a preset sampling rate, and synchronously records the trigger time mark of the calibration signal in the response acquisition step to ensure the one-to-one correspondence of the calibration signal and the response signal in the time domain.
6. The bone conduction hearing assistance method of claim 5, wherein, The transmission characteristic estimation step further includes: quality evaluation and artifact rejection of the eye movement response data, extraction of at least one of the microsaccade amplitude, speed peak, pupil micro-change, eye muscle activity proxy indicator or saccade occurrence rate in the eye movement response data, and generation of a credibility weight based on the stability and repeatability of the extracted features in the eye movement response data for weighted processing in the sub-band level representative transmission characteristic calculation process; The sub-band level representative transmission characteristic calculation takes the features of the eye movement response data and the energy and timing relationship of the calibration signal in each sub-band as inputs, and generates a sub-band level representative transmission characteristic that can be used in the calibration parameter generation step in combination with the sub-band weight determined by the credibility weight.
7. The bone conduction hearing assistance method of claim 6, wherein, The calibration parameter generation step outputs complex compensation parameters for each sub-band, and sets constraints on the amplitude range of the complex compensation parameters and the change slope between adjacent sub-bands to ensure the stability and naturalness of the output sound quality in the parameter application step; The parameter application step maps and applies the complex compensation parameters corresponding to the frequency points or sub-bands after performing frequency domain analysis on the input audio frame by frame, and adopts a time smoothing transition strategy to weight the new and old parameters when the complex compensation parameters are updated to avoid sudden changes in output sound quality; The signal injection step injects a low-amplitude short-duration and imperceptible calibration micro-vibration injection calibration signal, and performs a safety threshold check before injection to ensure that the driving strength of the bone conduction transducer does not exceed a preset threshold, and avoids overlapping with audible playback during injection to reduce interference with the user.
8. A bone conduction hearing aid, characterized in that It includes: a bone conduction transducer for injecting a preset calibration signal into the skull of a wearer and playing an adjusted audio signal; a reference sensor for acquiring a response signal of the calibration signal in the skull; and a processor electrically connected with the bone conduction transducer and the reference sensor, configured to: compare the calibration signal with the response signal to estimate the skull transmission characteristic of the wearer; generating calibration parameters for compensating individual differences based on the skull transmission characteristics; adjusting the amplitude and / or phase of the audio signal to be played according to the calibration parameters during the audio playing process, and outputting to the bone conduction transducer; The bone conduction hearing aid is configured to execute the bone conduction hearing method of any one of claims 1-7.
9. The bone conduction hearing device according to claim 8, when claim 8 refers to claim 3, characterized in that, The reference sensor comprises an eye movement acquisition unit. The processor is further configured to perform frame division, alignment, feature extraction and quality evaluation on the eye movement response data, and use the processing result as an input of the compensation parameters of the calibration parameter generation step to realize the eye movement micro-reaction-based personalized automatic calibration.
Citation Information
Patent Citations
Calibration of bone conduction transducer assembly
US10658995B1
Bone conduction headset and audio processing method thereof
CN105721973A
Method and apparatus for regulation of hearing air device and computer program
CN109729485A