Signal enhancement methods and systems for bone conduction headphones
Patent Information
- Application Number
- CN202610245675.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-26
Smart Images

Figure CN122093702A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of signal processing technology, and in particular to a signal enhancement method and system for bone conduction headphones. Background Technology
[0002] Currently, bone conduction headphones are irreplaceable in sports communication and hearing aids due to their open ear canal and ability to simultaneously perceive ambient sound via voice sensors. However, the cranial bone conduction pathway exhibits severe selective attenuation of low-frequency vibrations. A significant amount of fundamental frequency energy dissipates before reaching the inner ear, resulting not simply in insufficient loudness, but in the fundamental frequency itself failing to effectively stimulate an auditory response, leading to a lack of fullness in music rhythms and movie sound effects. Simply increasing the driving power can cause overheating and strain on battery life, and cannot compensate for the physical losses in the conduction medium. Therefore, there is an urgent need to develop methods that optimize vibration coupling efficiency or employ targeted signal processing within the constraints of a compact structure and low power consumption, enabling the low-frequency fundamental frequency to be fully transmitted to the cochlea, achieving a deep and powerful low-frequency perception.
[0003] In one existing technology, increasing the output voltage of the audio driver chip allows the oscillator to produce a larger physical displacement under the same signal; or replacing the diaphragm material with a more elastic one allows for greater amplitude vibration. Another approach is to increase the driving power, i.e., using a power amplifier component with a higher rated power and increasing the battery discharge rate to enhance the transient current supply, so that the vibration unit obtains greater electromagnetic driving force in the low-frequency range. At the same time, a digital equalizer is used to pre-emphasize the low-frequency signal, further increasing the low-frequency gain in the amplification stage. Traditional technologies rely solely on amplifying the amplitude and increasing the power to compensate for low frequencies, which leads to a large amount of energy loss due to the low conduction efficiency of the skull, and causes overheating, shortened battery life, and signal distortion.
[0004] Therefore, existing technologies cannot solve the signal attenuation and distortion problems of bone conduction headphones in complex usage scenarios. Summary of the Invention
[0005] This invention provides a signal enhancement method and system for bone conduction headphones to solve the problems of signal attenuation and distortion in bone conduction headphones under complex usage scenarios.
[0006] In a first aspect, to solve the above-mentioned technical problems, the present invention provides a signal enhancement method for bone conduction headphones, comprising: Acquire audio input signal; Based on the audio input signal, fundamental frequency localization and harmonic extraction are performed to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components. Based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components, dynamic range statistics are performed to obtain low frequency feature vectors. Based on the low-frequency feature vector, a nonlinear curve fitting operation is performed on the harmonic intensity segments to generate a piecewise nonlinear mapping function. Then, based on the low-frequency feature vector and the piecewise nonlinear mapping function, amplitude nonlinear mapping and phase distortion smoothing are performed to obtain an enhanced amplitude-phase sequence. The impulse response data is extracted from the enhanced amplitude-phase sequence to generate an inverse attenuation function, and the corrected response data is calculated. The enhanced amplitude-phase sequence is then updated based on the corrected response data to obtain a low-frequency enhanced signal. Based on the low-frequency enhanced signal, a skull conduction model simulation is performed to obtain simulated vibration transmission data. Then, based on the simulated vibration transmission data, standard data deviation is calculated to obtain amplitude difference sequence and phase deviation sequence. Frequency domain compensation modulation is performed based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence. Then, impedance current compensation is performed by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuned current sequence. Power distribution is estimated based on the fine-tuned current sequence to obtain the output power spectrum, and resonance drift tracking and frequency compensation are performed based on the output power spectrum to generate optimized vibration control parameters. Based on the optimized vibration control parameters, parameter fusion calculation is performed to obtain vibration control commands, and digital interface driving is performed based on the vibration control commands to obtain tactile vibration signals.
[0007] In a second aspect, the present invention provides a signal enhancement system for bone conduction headphones, comprising: The data acquisition module is used to acquire audio input signals; The low-frequency feature module is used to perform fundamental frequency localization and harmonic extraction based on the audio input signal to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components, and to perform dynamic range statistics based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components to obtain a low-frequency feature vector; The amplitude-phase enhancement module is used to perform nonlinear curve fitting operation on harmonic intensity segments based on the low-frequency feature vector to generate a piecewise nonlinear mapping function, and to perform amplitude nonlinear mapping and phase distortion smoothing based on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence. The low-frequency enhancement module is used to extract impulse response data from the enhanced amplitude-phase sequence, generate an inverse attenuation function, calculate corrected response data, update the enhanced amplitude-phase sequence based on the corrected response data, and obtain a low-frequency enhanced signal. The conduction simulation module is used to simulate the skull conduction model based on the low-frequency enhanced signal, obtain simulated vibration transmission data, and perform standard data deviation calculation based on the simulated vibration transmission data to obtain amplitude difference sequence and phase deviation sequence. The fine-tuning current module is used to perform frequency domain compensation modulation based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence, and to perform impedance current compensation by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuning current sequence. The vibration control module is used to estimate the power distribution based on the fine-tuning current sequence, obtain the output power spectrum, and perform resonance drift tracking and frequency compensation based on the output power spectrum to generate optimized vibration control parameters. The output module is used to perform parameter fusion calculation based on the optimized vibration control parameters to obtain vibration control commands, and to drive the digital interface based on the vibration control commands to obtain tactile vibration signals.
[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention obtains low-frequency feature vectors through fundamental frequency localization and harmonic extraction, and segments the harmonics according to their proportions, configuring gain and compression coefficients for weak, medium, and strong harmonic intervals respectively. A piecewise nonlinear mapping function is generated by using cubic spline interpolation to map the fundamental frequency amplitude point by point and smooth the phase distortion using a Hanning window. This scheme avoids the harmonic distortion and overshoot caused by the global boosting of low frequencies in traditional equalizers, and maintains the naturalness and transient clarity of the sound while improving the perception of the fundamental frequency.
[0009] (2) This invention identifies abnormal impact peaks based on the enhanced amplitude-phase sequence, constructs a linearly decreasing weighted sequence as an inverse attenuation function through the overshoot deviation ratio, performs point-by-point weighted correction on the impact response data, and combines Hilbert envelope extraction and Gaussian smoothing to achieve transient shaping in the time domain. This scheme does not rely on hard truncation by the limiter, achieves gradual attenuation of the impact response, suppresses mechanical overshoot and residual vibration tailing of the vibration unit, and makes the low-frequency sound more solid and clean.
[0010] (3) This invention generates simulated vibration transmission data through a skull conduction model and calculates the amplitude and phase deviation with standard bone conduction data. It introduces PID compensation and time delay correction to generate gain compensation coefficients and complex compensation data in the frequency domain. At the same time, it tracks the resonance drift trend and generates frequency tracking quantity through PID frequency compensation and spectrum shifting. This scheme explicitly models and compensates for conduction loss and resonance offset in a closed loop, so that the drive signal is always aligned with the current resonance point, improving low-frequency transmission efficiency, wearing consistency and battery life. Attached Figure Description
[0011] Figure 1 This is a schematic flowchart of the signal enhancement method for bone conduction headphones provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the signal enhancement system structure of the bone conduction headphones provided in the second embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Reference Figure 1 The first embodiment of the present invention provides a signal enhancement method for bone conduction headphones, comprising the following steps: S11, acquire audio input signal; S12, perform fundamental frequency localization and harmonic extraction based on the audio input signal to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components, and perform dynamic range statistics based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components to obtain low frequency feature vector; S13, based on the low-frequency feature vector, a nonlinear curve fitting operation is performed on the harmonic intensity segment to generate a piecewise nonlinear mapping function, and amplitude nonlinear mapping and phase distortion smoothing are performed on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence; S14, extract the impact response data according to the enhanced amplitude-phase sequence, generate the inverse attenuation function, calculate the corrected response data, update the enhanced amplitude-phase sequence according to the corrected response data, and obtain the low-frequency enhanced signal; S15, perform skull conduction model simulation based on the low-frequency enhanced signal to obtain simulated vibration transmission data, and perform standard data deviation calculation based on the simulated vibration transmission data to obtain amplitude difference sequence and phase deviation sequence; S16, frequency domain compensation modulation is performed based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence, and impedance current compensation is performed by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuned current sequence. S17. Based on the fine-tuned current sequence, power distribution is estimated to obtain the output power spectrum, and based on the output power spectrum, resonance drift tracking and frequency compensation are performed to generate optimized vibration control parameters. S18, perform parameter fusion calculation based on the optimized vibration control parameters to obtain vibration control commands, and drive the digital interface based on the vibration control commands to obtain tactile vibration signals.
[0014] In step S11, the audio input signal is acquired.
[0015] Specifically, in step S11, an audio input signal is acquired. The audio input signal originates from an external audio source device connected to the bone conduction headphones, receiving raw sampled data through a digital audio interface. External audio source devices include smartphones, tablets, or personal computers, which transmit digital audio streams to the signal processing unit of the bone conduction headphones via wired or wireless means.
[0016] The specific receiving process involves the digital interface controller of the signal processing unit continuously monitoring the input buffer. When a valid data packet is detected, the audio frame is parsed according to a preset sampling rate (e.g., 48 kHz) and bit depth (e.g., 24 bits), converting it into a time-domain sample sequence in floating-point format. This time-domain sample sequence is arranged in chronological order to form a complete audio input signal, which serves as the basis for subsequent fundamental frequency localization and harmonic extraction.
[0017] Acquiring the audio input signal is the starting point of the entire signal enhancement process, providing the raw processing material for subsequent steps and ensuring that the system can respond promptly to various audio content played by the user. The audio input signal adopts a dual-channel stereo format, with each time point containing sample values from the left and right channels. The system performs equal-weighted mixing processing on the dual-channel signals, calculates the arithmetic mean of the sample values from the two channels at the corresponding time point, and generates a single-channel processed signal. This single-channel processed signal retains the time-domain waveform characteristics of the original audio, and the amplitude range is normalized to the range of -1 to +1 to avoid numerical overflow issues in subsequent processing.
[0018] In step S12, fundamental frequency localization and harmonic extraction are performed based on the audio input signal to obtain the fundamental frequency amplitude, fundamental frequency phase, and harmonic components. Dynamic range statistics are then performed based on the fundamental frequency amplitude, the fundamental frequency phase, and the harmonic components to obtain a low-frequency feature vector, including: Based on the audio input signal, a filtering operation is performed using the Butterworth low-pass filtering algorithm to obtain a low-frequency signal; Based on the low-frequency signal, a time-frequency conversion operation is performed using the short-time Fourier transform algorithm to obtain the spectrum matrix; Based on the spectrum matrix, locate the frequency and complex value corresponding to the amplitude peak in each frame to obtain the fundamental frequency amplitude sequence, fundamental frequency phase sequence and fundamental frequency sequence; Based on the fundamental frequency sequence, the amplitude of the corresponding frequency point is extracted within a preset octave range to obtain the harmonic component sequence; Based on the fundamental frequency amplitude sequence, the difference between the maximum and minimum values is calculated to obtain the global amplitude range; The fundamental frequency amplitude sequence, the fundamental frequency phase sequence, the global amplitude range, and the harmonic component sequence are combined and spliced to obtain the low-frequency feature vector.
[0019] Specifically, in step S12, the audio input signal obtained in step S11 is first used as the data source. This signal is a discrete-time sequence stored in pulse code modulation format, denoted as the audio input signal sequence. A Butterworth low-pass filter operation is performed on this sequence. The filter order is set to fourth order, and the cutoff frequency is statistically determined based on the attenuation inflection point distribution of the skull vibration transfer function of 120 groups of bone conduction headphones in historical tests. Specifically, the low-frequency attenuation start frequencies of each group are arranged in ascending order, and the 85th percentile frequency value is taken as the cutoff frequency, ensuring that the attenuation start point of 85% of the test samples is contained within the passband. The filtering operation first maps the transfer function of the analog filter to the coefficients of the digital filter using the bilinear transform method, constructing a cascaded structure of four second-order sections. Each second-order section performs a direct type II recursive operation. The current output sample value is equal to the input sample value multiplied by the first numerator coefficient, plus the input sample value of the previous step multiplied by the second numerator coefficient, plus the input sample value of the previous two steps multiplied by the third numerator coefficient, and then subtracts the output sample value of the previous step multiplied by the first denominator coefficient, and subtracts the output sample value of the previous two steps multiplied by the second denominator coefficient. This cascading process completes the fourth-order filtering, and the final output sequence is called the low-frequency signal sequence.
[0020] It should be noted that the historical test refers to the test conducted on 120 representative bone conduction headphone samples in a standard semi-anechoic chamber or simulated head model. During the test, a logarithmic sweep frequency signal was used as input, and the vibration response transmitted to the simulated point in the inner ear was measured by a simulated artificial ear or accelerometer attached to the temporal bone. The low-frequency attenuation start frequency is defined as the frequency point at which the amplitude drops by 3 dB (decibels) from the low-frequency reference value when the measured skull vibration transfer function amplitude-frequency curve transitions from the low-frequency flat region to the downward trend. By statistically analyzing this frequency point of all samples, the 85th percentile value after ascending order can be determined as the cutoff frequency of the Butterworth low-pass filter.
[0021] Next, a short-time Fourier transform (SFT) operation is performed on the low-frequency signal sequence to generate a spectrum matrix. First, the low-frequency signal sequence is divided into continuous data blocks according to frame length. The frame length is determined based on the fundamental frequency period distribution of speech data in a historical corpus: the fundamental frequency periods of 500 typical music and speech samples are measured, and 120% of the maximum period value is taken as the frame length, ensuring at least two complete periods are included. The frame shift is set to 50% of the frame length to achieve overlap analysis. Each frame of data is multiplied point-by-point by a Hanning window function. The Hanning window coefficients are calculated from the normalized angular frequency, and the window function length is the frame length. Each windowed frame sequence is input into a radix-2 decimation-time fast Fourier transform kernel. A butterfly operation maps the time-domain sequence to the frequency domain, outputting a set of complex arrays for each frame. The complex arrays of all frames are arranged in chronological order to form a two-dimensional complex matrix, called the spectrum matrix.
[0022] The fundamental frequency information of each frame is located based on the spectrum matrix. For the complex spectrum vector corresponding to each frame, the amplitude of each frequency point is extracted. The amplitude is calculated as the square root of the sum of the square of the real part and the square of the imaginary part. The frequency point corresponding to the maximum amplitude within the frame is searched; the frequency value of this point is the fundamental frequency, the modulus of the complex value at this frequency point is the fundamental frequency amplitude, and the argument of the complex value at this frequency point is the fundamental frequency phase. The argument is obtained by calculating the ratio of the imaginary part to the real part using the arctangent function and then performing quadrant correction. After traversing all frames, the fundamental frequency amplitude, fundamental frequency phase, and fundamental frequency are arranged in chronological order, generating the fundamental frequency amplitude sequence, fundamental frequency phase sequence, and fundamental frequency sequence, respectively.
[0023] Harmonic components are extracted based on the fundamental frequency sequence. The preset harmonic range is from the second to the fifth harmonic, with the upper limit determined statistically based on the nonlinear distortion spectrum of historical vibration units. Harmonic distortion components of eighty bone conduction oscillators under different driving voltages are measured, and the frequency range where the amplitudes of the second to fifth harmonics all exceed the noise floor is recorded. The harmonic order corresponding to the maximum value in this range is taken as the fifth harmonic. For each frame, the frequency points from the second to the fifth harmonic are calculated based on the fundamental frequency, and the amplitudes at these frequency points are extracted from the complex spectrum of the corresponding frame in the spectrum matrix. If a harmonic frequency point falls between frequency resolution grids, the linear weighted value of the amplitudes of two adjacent frequency points is taken. The harmonic amplitudes of all frames are arranged according to the frame number and harmonic order to form a harmonic component sequence.
[0024] Dynamic range statistics are performed based on the fundamental frequency amplitude sequence. All elements in the fundamental frequency amplitude sequence are traversed, and the current maximum and minimum values are compared and retained one by one. The extreme values are updated after processing each sampling point. After completing the full sequence scan, the difference between the maximum and minimum values is the global amplitude range. Finally, the fundamental frequency amplitude sequence, fundamental frequency phase sequence, global amplitude range scalar, and harmonic component sequence are concatenated end-to-end according to a preset splicing order to form a one-dimensional feature array, called the low-frequency feature vector. This vector condenses the fundamental frequency energy, phase information, harmonic structure, and dynamic range of the audio input signal in the low-frequency band, providing input parameters for subsequent piecewise nonlinear mapping.
[0025] In step S13, a piecewise nonlinear curve fitting operation is performed on the low-frequency feature vector through harmonic intensity segmentation to generate a piecewise nonlinear mapping function. Then, amplitude nonlinear mapping and phase distortion smoothing are performed based on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence, including: Based on the low-frequency feature vector, the ratio of the sum of the amplitude values of each harmonic component to the fundamental frequency amplitude is calculated to obtain the harmonic proportion sequence. The harmonic proportion sequence is then compared with the preset first proportion threshold and second proportion threshold to mark the harmonic components as weak harmonic range, medium harmonic range or strong harmonic range, and the harmonic range is determined. Based on the harmonic intervals, a set of preset gain coefficients and compression coefficients are configured for each interval, and the boundary values of the harmonic intervals are used as nodes. A continuous nonlinear curve is fitted using a cubic spline interpolation algorithm to generate a piecewise nonlinear mapping function. Based on the fundamental frequency amplitude sequence, the piecewise nonlinear mapping function is used to perform point-by-point calculations to map each fundamental frequency amplitude to a new amplitude, thereby obtaining the enhanced amplitude sequence. The amplitude difference between adjacent time frames in the enhanced amplitude sequence is calculated to obtain the amplitude change gradient. When the amplitude change gradient is greater than a preset phase distortion threshold, the fundamental frequency phase sequence is smoothed using a Hanning window weighted moving average method to obtain an adjusted phase sequence. The enhanced amplitude sequence and the adjusted phase sequence are combined according to time frames to obtain the enhanced amplitude-phase sequence.
[0026] Specifically, the harmonic proportion sequence is compared with preset first and second proportion thresholds to mark harmonic components as weak harmonic intervals, moderate harmonic intervals, or strong harmonic intervals, thereby determining the harmonic intervals, including: The harmonic proportion sequence is compared with a preset first proportion threshold and a preset second proportion threshold. When the harmonic proportion is lower than the preset first proportion threshold, the corresponding harmonic component is marked as a weak harmonic range. When the harmonic proportion is between the first proportion threshold and the preset second proportion threshold, the corresponding harmonic component is marked as a medium harmonic range. When the harmonic proportion is higher than the second proportion threshold, the corresponding harmonic component is marked as a strong harmonic range, thus determining the harmonic range.
[0027] Specifically, in step S13, the data source is the low-frequency feature vector output in step S12. This vector is composed of the fundamental frequency amplitude sequence, the fundamental frequency phase sequence, the global amplitude range, and the harmonic component sequence. First, the fundamental frequency amplitude sequence, the fundamental frequency phase sequence, and the harmonic component sequence are separated from the low-frequency feature vector. For each time frame, the amplitudes of all octave points in the harmonic component sequence are accumulated, starting from the second harmonic amplitude and successively adding up to the fifth harmonic amplitude, to obtain the total harmonic amplitude of that frame. The total harmonic amplitude is divided by the fundamental frequency amplitude of that frame, and the resulting ratio is the harmonic proportion of that frame. The same division operation is performed on all frames to generate a one-dimensional harmonic proportion sequence.
[0028] The determination of the first and second proportion thresholds is based on the harmonic distribution analysis of four hundred bone conduction headphone wearing test samples during the offline statistical phase. The collected samples covered music, voice, and mixed scenarios, with audio input signals and vibration unit output signals recorded simultaneously for each sample. First, the harmonic proportion sequence for each sample was calculated and its mean was taken. The four hundred means were arranged in ascending order, and the 30th percentile value was used as the first proportion threshold, and the 70th percentile value was used as the second proportion threshold. These two thresholds divide the harmonic intensity into three intervals: weak, medium, and strong. Each element in the harmonic proportion sequence was compared sequentially with the first and second proportion thresholds: if the harmonic proportion was less than the first proportion threshold, the corresponding harmonic component of that frame was marked as a weak harmonic interval; if the harmonic proportion was greater than or equal to the first proportion threshold and less than or equal to the second proportion threshold, it was marked as a medium harmonic interval; if the harmonic proportion was greater than the second proportion threshold, it was marked as a strong harmonic interval. This yielded the harmonic interval label for each frame.
[0029] A set of gain and compression coefficients was pre-configured for each harmonic range, determined through offline calibration experiments. During calibration, one hundred audio clips covering all harmonic ranges were selected. Each clip underwent amplitude mapping processing with different combinations of gain and compression coefficients. The processed signals were then input to the bone conduction headphone vibration unit, while the output vibration signals were simultaneously acquired, and total harmonic distortion (THD) and power spectral flatness were calculated. The optimization objective was to keep the THD below a preset upper limit and maximize power spectral flatness. In the weak harmonic range, a gain coefficient maximizing flatness and a compression coefficient minimizing distortion were searched; in the strong harmonic range, a gain coefficient minimizing distortion and a compression coefficient maximizing flatness were searched; and in the intermediate harmonic range, a compromise between the two was taken. Ultimately, a high-gain, low-compression coefficient was configured for the weak harmonic range, a medium-gain, medium-compression coefficient for the intermediate harmonic range, and a low-gain, high-compression coefficient for the strong harmonic range.
[0030] The piecewise nonlinear mapping function is constructed with the fundamental frequency amplitude as the independent variable and the enhanced amplitude as the dependent variable. First, the fundamental frequency amplitude distribution of all frames is statistically analyzed. The maximum value of the fundamental frequency amplitude in each frame marked as the weak harmonic interval is identified as the first boundary node, and the minimum value of the fundamental frequency amplitude in each frame marked as the strong harmonic interval is identified as the second boundary node. Simultaneously, the global minimum and maximum fundamental frequency amplitudes are identified as the first and last nodes, respectively. These four nodes divide the fundamental frequency amplitude domain into three intervals. Within each interval, a local mapping relationship is defined based on the corresponding gain coefficient and compression coefficient: the input amplitude is multiplied by the gain coefficient, and then a power operation is performed with the compression coefficient as the exponent to obtain the output amplitude at that node. After each of the four nodes obtains its corresponding output amplitude, four interpolation points are formed with the node amplitude as the x-coordinate and the output amplitude as the y-coordinate. A cubic spline interpolation algorithm is used to construct a cubic polynomial between each adjacent node in the sub-interval. This polynomial is equal to the given y-coordinate value at the node, and its first and second derivatives are continuous at the node. Solving the tridiagonal linear equations yields the coefficients of the polynomials in each subinterval. All the subinterval polynomials together constitute a piecewise nonlinear mapping function whose domain covers the entire range of the fundamental frequency amplitude.
[0031] Each element in the fundamental frequency amplitude sequence is input into the mapping function as an independent variable. First, the interval to which the amplitude belongs is determined. Then, a cubic polynomial for the corresponding sub-interval is called. The difference between the amplitude and the node value is substituted into the polynomial for calculation, and the new amplitude is output. After performing point-by-point calculations on all frames, an enhanced amplitude sequence is generated.
[0032] Phase distortion smoothing first calculates the amplitude difference between adjacent frames in the enhanced amplitude sequence, i.e., the absolute value of the amplitude of the subsequent frame minus the amplitude of the preceding frame, to obtain the amplitude change gradient sequence. The phase distortion threshold is determined based on the critical value of audibly perceptible phase abrupt changes in historical tests: subjective listening tests and objective phase deviation measurements are performed on 300 sets of bone conduction headphone output signals to mark the samples that are perceived as distorted by listeners. The smallest amplitude change gradient among these samples is taken as the threshold, ensuring that 95% of the distorted samples are detected. When the amplitude change gradient of a frame exceeds the phase distortion threshold, a Hanning window weighted moving average is applied to the fundamental frequency phase sequence of that frame and the two frames before and after it (a total of five frames). The Hanning window coefficients are generated from five points of the window length, with the center point having the highest weight and decreasing towards both sides. The phase values of the five frames are multiplied by their respective window coefficients, summed, and then divided by the sum of the window coefficients to obtain the adjusted phase value of that frame. After all frames are processed, an adjusted phase sequence is generated.
[0033] The enhanced amplitude sequence and the adjusted phase sequence are aligned with the same frame number. Each frame consists of a complex number formed by the enhanced amplitude and the adjusted phase, with the real part being the amplitude multiplied by the phase cosine and the imaginary part being the amplitude multiplied by the phase sine. This complex number sequence is called the enhanced amplitude-phase sequence. This sequence retains the amplitude information after the fundamental frequency energy is enhanced and suppresses the phase jump caused by abrupt amplitude changes, providing a clean low-frequency vibration characterization for subsequent impulse response correction.
[0034] In step S14, impulse response data is extracted from the enhanced amplitude-phase sequence to generate an inverse attenuation function, and corrected response data is calculated. The enhanced amplitude-phase sequence is then updated based on the corrected response data to obtain a low-frequency enhanced signal, including: Based on the enhanced amplitude-phase sequence, the time point when the amplitude exceeds the preset dynamic range compression threshold is identified to obtain the abnormal peak point, and the corresponding enhanced amplitude-phase sequence within a preset time window before and after the abnormal peak point is taken as the impact response data. Calculate the difference between the amplitude of the abnormal peak point in the impact response data and the dynamic range compression threshold, and divide the difference by the dynamic range compression threshold to obtain the overshoot deviation ratio. Based on the overshoot ratio, a weight sequence that decreases linearly from the abnormal peak point to both sides is generated using a linear interpolation algorithm as a reverse decay function. The impact response data is then multiplied point by point with the reverse decay function to obtain the corrected response data. The original data in the corresponding window of the enhanced amplitude-phase sequence is replaced with the corrected response data, and the envelope of the updated enhanced amplitude-phase sequence is extracted by Hilbert transform to obtain the low-frequency envelope profile. The low-frequency envelope contour is convolved using a Gaussian smoothing kernel to obtain the low-frequency enhanced signal.
[0035] Specifically, the data source for step S14 is the enhanced amplitude-phase sequence output in step S13. This sequence is stored in complex form, with each time frame containing the enhanced amplitude and adjusted phase. The real part is the amplitude multiplied by the phase cosine, and the imaginary part is the amplitude multiplied by the phase sine. First, the absolute values of the amplitude of each frame are extracted from the enhanced amplitude-phase sequence to form the enhanced amplitude monitoring sequence. The dynamic range compression threshold is determined based on offline statistics. Five hundred sets of enhanced amplitude-phase sequences output by bone conduction headphones in music, speech, and mixed scenarios are collected, each set being thirty seconds long. The amplitude values of all time frames are extracted and merged into a total sample. This sample set is sorted in ascending order of value, and the amplitude value corresponding to the 95th percentile is taken as the dynamic range compression threshold to ensure that 95% of normal vibration amplitudes are not marked, and only the 5% with the largest amplitude are corrected.
[0036] The enhanced amplitude monitoring sequence is traversed, and the amplitude of each frame is compared with the dynamic range compression threshold. When the amplitude exceeds the threshold, the frame index is recorded as an abnormal peak point. For each abnormal peak point, a time window of preset length is taken before and after it. The half-width of the window is determined statistically based on the vibration unit's impact response decay time. One hundred sets of vibration decay waveforms triggered by pulse excitation are measured, and the number of sampling frames required for the peak to decrease to steady state is recorded. The maximum value of all measurement results is taken and rounded up as the half-width of the window. The total length of the window is twice the half-width plus one. The corresponding enhanced amplitude-phase sequence segment is extracted within this window, which is called the impact response data.
[0037] For each abnormal peak point, the difference between its amplitude and the dynamic range compression threshold is calculated. The difference is the peak amplitude minus the threshold. This difference is then divided by the dynamic range compression threshold, and the resulting quotient is called the overshoot ratio. This ratio quantifies the relative degree to which the peak exceeds the threshold.
[0038] Based on the overshoot ratio, a weight sequence for the inverse decay function is constructed. First, the weight coefficient at the peak point is determined and set as the threshold divided by the peak amplitude, i.e., the reciprocal of the overshoot ratio plus one, so that the peak amplitude exactly drops to the threshold after the multiplication operation. Using the peak point index as the center, extend half a width of indices to the left and right, constructing a weight sequence with a length equal to the total window size. The weights at the left and right boundaries of the window are preset to one, while the weight at the peak point is the aforementioned coefficient. A linear interpolation algorithm is used to generate the weight values for each index within the interval from the left boundary to the peak point. The weights in this interval decrease linearly from one at the left boundary to the coefficient at the peak point; the weights for each index within the interval from the peak point to the right boundary increase linearly from the coefficient at the peak point to one at the right boundary. This forms a complete weight sequence that first decreases and then increases from the peak point outwards, which is the inverse decay function.
[0039] The impact response data is multiplied point by point with the weight sequence of the inverse decay function. The complex amplitude of each frame is multiplied by the corresponding weight to obtain the corrected complex sequence, which is called the corrected response data.
[0040] The original data within the corresponding window of the enhanced amplitude-phase sequence is completely replaced with the corrected response data, covering all windows containing abnormal peak points, to generate an updated enhanced amplitude-phase sequence. A Hilbert transform is then performed on this updated sequence to extract the time-domain envelope. First, a Fast Fourier Transform (FFT) is performed on the sequence to obtain a complex array in the frequency domain. The amplitudes of the positive frequency components are doubled, and the amplitudes of the negative frequency components are set to zero. Then, an Inverse Fast Fourier Transform (IFFT) is performed to obtain the analytic signal sequence. The magnitude of each frame of the analytic signal is calculated by taking the square root of the sum of the squared real and imaginary parts. This frame-by-frame calculation yields a one-dimensional real number sequence, called the low-frequency envelope profile.
[0041] This profile reflects the overall trend of vibration amplitude changes over time, but may retain residual pulse spikes. To obtain a smooth driving reference, Gaussian smoothing convolution is performed on the low-frequency envelope profile. The standard deviation of the Gaussian smoothing kernel parameters is determined based on the envelope noise level of historical stationary segments: stationary segments are extracted from 300 sets of low-frequency envelope profiles without impact response, and the standard deviation of their high-frequency fluctuation components is calculated. Three times this standard deviation is taken as the standard deviation of the Gaussian kernel to ensure that the main profile is preserved while suppressing random fluctuations after smoothing. The kernel length is set to the nearest odd number after six times the standard deviation plus one. The kernel coefficients are calculated from the values of the Gaussian function at each discrete position and normalized by dividing by the sum of all coefficients. The low-frequency envelope profile sequence is convolved one-dimensionally with the normalized Gaussian kernel, and symmetrical boundary expansion is used to process the beginning and end of the sequence. The convolution result is the low-frequency enhanced signal.
[0042] This signal achieves progressive compression and envelope shaping of abnormal impact peaks in the time domain, which not only suppresses the mechanical overshoot of the vibration unit but also maintains the natural envelope shape of the low-frequency rhythm, providing a stable amplitude and clean transient excitation input for subsequent skull conduction model simulation.
[0043] In step S15, a skull conduction model simulation is performed based on the low-frequency enhanced signal to obtain simulated vibration transmission data. Standard data deviation is then calculated based on the simulated vibration transmission data to obtain an amplitude difference sequence and a phase deviation sequence, including: Based on the low-frequency enhanced signal, the instantaneous amplitude envelope and instantaneous frequency are extracted by Hilbert transform to obtain the amplitude envelope sequence and instantaneous frequency offset; The amplitude envelope sequence and the instantaneous frequency offset are input into a pre-established skull conduction frequency response model to obtain simulated vibration transmission data. The amplitude values of the simulated vibration transmission data and the pre-stored standard bone conduction data are subtracted at the same frequency point to obtain an amplitude difference sequence. The cross-correlation function between the simulated vibration transmission data and the standard bone conduction data is calculated to obtain the phase deviation sequence.
[0044] Specifically, the data source for step S15 is the low-frequency enhancement signal output in step S14. This signal is a one-dimensional real number sequence, with each element corresponding to the vibration amplitude envelope value of a time sampling point, reflecting the time-domain profile of the low-frequency vibration after impact shaping.
[0045] First, a Hilbert transform is performed on the signal to extract instantaneous characteristics. A fast Fourier transform is then performed on the low-frequency enhanced signal sequence to obtain a complex array in the frequency domain. The amplitude of the positive frequency components is doubled, and the amplitude of the negative frequency components is set to zero. Then, an inverse fast Fourier transform is performed to obtain the analytic signal sequence. The magnitude of each frame of the analytic signal is the instantaneous amplitude envelope, and these frames are arranged in order to form an amplitude envelope sequence. The instantaneous phase of each frame of the analytic signal is obtained by taking the arctangent of the ratio of the imaginary part to the real part. The difference between the instantaneous phases of adjacent frames divided by the sampling interval is the instantaneous angular frequency. The difference between the instantaneous angular frequency and the standard center frequency constitutes the instantaneous frequency offset sequence. This offset represents the degree of deviation of the current vibration frequency from the ideal driving frequency.
[0046] The cranial conduction frequency response model is a pre-constructed linear time-invariant system model used to simulate the amplitude and phase response of bone conduction headphone vibration signals transmitted through the skull to the cochlea. The model structure adopts a zero-pole-gain form, and its transfer function consists of six cascaded second-order sections, each containing a pair of conjugate complex poles and a pair of conjugate complex zeros. Model parameters were determined through offline system identification experiments. One hundred and twenty sets of bone conduction headphone vibration input and inner ear received vibration output data were collected from wearers of different ages and genders. The input signal was a logarithmically swept frequency signal, and the output signal was picked up by an accelerometer attached to the temporal bone. A fast Fourier transform was performed on each set of data to calculate the input-output spectrum ratio, obtaining the measured frequency response curve. The median amplitude and median phase of all measured frequency response curves were taken at the same frequency points to generate a standard frequency response template. A least-squares iterative algorithm was used, with the standard frequency response template as the target, to adjust the zero-pole positions to minimize the deviation between the model's frequency response and the template. The iteration terminated when the mean square error change between two consecutive iterations was less than a preset tolerance.
[0047] It should be noted that the offline system identification experiment uses the generalized least squares method for iterative optimization of model parameters until the mean square error between the model output spectrum and the measured standard frequency response template is less than 1e-6, at which point the iteration terminates. The higher-order polynomial fitting refers to the process of using a 15th-order polynomial to perform least squares fitting on the amplitude-frequency curve and phase-frequency curve after obtaining 300 sets of extended sample standard frequency response template data points (frequency, gain / phase). The order is chosen such that the root mean square of the fitting residual no longer decreases significantly with increasing order. The measurement conditions, sensor types, and data post-processing procedures for the 300 sets of extended samples are exactly the same as those followed for the 120 sets of samples used to construct the skull conduction frequency response model, to ensure the consistency of the data basis.
[0048] Once the model is built, its inputs are the amplitude envelope value and instantaneous frequency offset at the current moment, and the output is the simulated vibration transmission data at the corresponding moment. The specific calculation process is as follows: the amplitude envelope value is considered as the baseband excitation amplitude; the instantaneous frequency offset is used to correct the current resonant frequency; based on the offset, the amplitude gain and phase delay corresponding to that frequency point are linearly interpolated from the model's frequency response table; the excitation amplitude is multiplied by the gain to obtain the output amplitude; the output phase is generated by accumulating the input phase and the delay; and the final output is a sequence of simulated vibration transmission data in complex form.
[0049] Standard bone conduction data is a reference frequency response dataset pre-stored in non-volatile memory. Its generation process is the same as the model calibration process described above, but the data source is expanded to include three hundred sets of measured samples covering a wider population, and higher-order polynomial fitting is used to obtain smooth amplitude and phase frequency curves.
[0050] The data is indexed by frequency, storing the ideal output amplitude at each frequency point. The deviation calculation first aligns the simulated vibration transmission data and the standard bone conduction data to the same frequency grid. A short-time Fourier transform is performed on the simulated vibration transmission data sequence to convert the time-domain complex sequence to the frequency domain, obtaining the amplitude and phase at each frequency point. At each common frequency point, the amplitude of the simulated vibration transmission data is subtracted from the amplitude of the standard bone conduction data; the difference is the amplitude difference value at that frequency point. All frequency point differences are arranged into an amplitude difference sequence. The phase deviation sequence is calculated using a cross-correlation function. The phase sequences of the simulated vibration transmission data and the standard bone conduction data are slidably aligned in the frequency domain. At each frequency point, the phase offset corresponding to the cross-correlation peak in a small neighborhood of that frequency point is calculated; this offset is the phase deviation. All frequency point phase deviations are arranged into a phase deviation sequence.
[0051] This step explicitly quantifies the loss and offset of the physical conduction path as frequency domain error, providing a precise correction target for subsequent current compensation.
[0052] In step S16, frequency domain compensation modulation is performed based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence. Then, based on the current amplitude sequence, impedance current compensation is performed by calculating the impedance offset to obtain a fine-tuned current sequence, including: Based on the amplitude difference sequence, the gain compensation coefficient is calculated using a PID algorithm, and the phase deviation sequence is input into a preset phase-delay conversion function to calculate the delay compensation amount. The gain compensation coefficient and the time delay compensation amount are combined and encapsulated to generate frequency point compensation data; The frequency compensation data and the pre-stored sinusoidal drive command are multiplied in the frequency domain to generate a current amplitude sequence. The feedback voltage and drive current are collected, the real-time impedance value is calculated using Ohm's law, and the difference between the preset reference impedance value and the real-time impedance value is calculated to obtain the impedance offset value. The current compensation coefficient is obtained by querying the preset current compensation table according to the impedance offset value, and then the current compensation coefficient is multiplied by the current amplitude sequence to obtain the fine-tuning current sequence.
[0053] Specifically, the data source for step S16 is the amplitude difference sequence and phase deviation sequence output from step S15, with both sequences corresponding one-to-one by frequency index. First, proportional-integral-derivative (PID) control is performed on each frequency point in the amplitude difference sequence. The proportional, integral, and derivative coefficients of the PID controller are determined through an offline step response calibration experiment. Fifty bone conduction headphones were selected for the experiment, each fixed on a rigid fixture. A standard step current excitation was applied to the vibration unit, and vibration acceleration response curves were collected using a laser vibrometer. The rise time (the interval between 10% and 90% of the steady-state value) and the percentage of the maximum overshoot relative to the steady-state value were read from each response curve. The rise time and overshoot were substituted into the Ziegler-Nichols empirical formula to calculate the initial proportional, integral, and derivative coefficients, respectively. The controller was then connected to the closed-loop test system. The amplitude difference sequence was used as the error input. With the goal of minimizing the root mean square of the residual deviation, the fixed step size traversal method was used to search for the optimal combination within the range of ±50% of the initial coefficients. Finally, the coefficient value that minimizes the root mean square residual deviation was selected and fixed into the algorithm.
[0054] During real-time operation, the amplitude difference at the current frequency is used as the error value at the current moment. This error value is multiplied by a scaling factor to obtain the proportional term. The error values are accumulated on the time axis, and the sum is multiplied by an integral factor to obtain the integral term. The error change is obtained by subtracting the error value from the previous moment from the current error value. This change is divided by the sampling interval to obtain the error change rate, and the error change rate is multiplied by a differential factor to obtain the differential term. The proportional term, integral term, and differential term are added together, and the sum is then subjected to upper and lower amplitude limiting to obtain the gain compensation coefficient for that frequency. The above operations are performed sequentially for all frequencies to generate a sequence of gain compensation coefficients.
[0055] The phase deviation sequence is input into a phase delay conversion function to calculate the delay compensation. This conversion function is pre-constructed based on the principle of a linear phase system. The construction process first collects input and output phase response data of 300 sets of bone conduction headphones under standard wearing conditions to verify that the phase deviation at each frequency point has a linear relationship with the reciprocal of the frequency. The function form is defined as the delay compensation equal to the phase deviation divided by the product of twice pi and the frequency. During real-time operation, for each frequency point, the phase deviation value of that frequency point is extracted from the phase deviation sequence, and the frequency value of that frequency point is obtained simultaneously. The phase deviation value is divided by twice pi and then by the frequency value; the quotient is the delay compensation amount for that frequency point. All delay compensation amounts for all frequency points are arranged by frequency index to generate a delay compensation amount sequence.
[0056] The gain compensation coefficient sequence and the time delay compensation sequence are combined and encapsulated in the same frequency order. At each frequency, the gain compensation coefficient is used as the real part, and the time delay compensation is first converted into a phase offset by multiplying the time delay compensation by the corresponding angular frequency of that frequency. The angular frequency is calculated by multiplying twice pi by the frequency. The resulting phase offset is used as the imaginary part. The real and imaginary parts form a complex number, called the frequency compensation data for that frequency. All frequency compensation data are arranged by frequency index to form a frequency compensation data sequence.
[0057] The pre-stored sinusoidal drive instruction is a reference complex number sequence stored in non-volatile memory at the factory. This sequence is indexed by frequency, with each frequency point storing a complex number whose real part is one and whose imaginary part is zero, representing a unit amplitude zero-phase sine wave in the frequency domain. The frequency point compensation data sequence and the sinusoidal drive instruction sequence are multiplied at the corresponding frequency points using complex number multiplication. The multiplication rule is: multiplying the real parts of the two complex numbers and subtracting the multiplication of their imaginary parts yields the new real part; multiplying the real part and imaginary part and adding the imaginary part and real part yields the new imaginary part. The resulting complex number sequence is the current amplitude sequence. The modulus of each complex number in this sequence represents the amplitude of the drive current at that frequency point, and the argument represents the phase shift relative to the reference sine wave.
[0058] Real-time impedance calculation is achieved by acquiring feedback voltage and drive current. The feedback voltage is continuously captured at a fixed sampling rate by an analog-to-digital converter connected in parallel to the output of the bone conduction headphone drive circuit. The feedback current is synchronously captured by another analog-to-digital converter after the voltage across a precision sampling resistor connected in series in the vibration unit circuit is conditioned by a differential amplifier. At each sampling moment, the instantaneous value of the feedback voltage is divided by the instantaneous value of the feedback current to obtain the real-time impedance value at that moment. The reference impedance value is determined during the factory calibration phase of each headphone. During the calibration process, a standard sinusoidal sweep signal is applied to the headphone, and the voltage and current are recorded synchronously, while the impedance amplitude at each frequency point is calculated. The impedance amplitude at the resonant frequency point is taken as the reference impedance value and written to the storage area. During real-time operation, the calculated real-time impedance value is subtracted from the reference impedance value, and the difference is recorded as the impedance offset value. If the real-time impedance value is greater than the reference impedance value, the offset is positive; otherwise, it is negative.
[0059] The current compensation table was pre-constructed through offline temperature rise and aging experiments. One hundred bone conduction vibration units were selected and placed in a constant temperature chamber, with twenty equally spaced current amplitudes ranging from the minimum starting current to the rated maximum current applied. Excitation was maintained at each current amplitude level for thirty minutes, and the impedance value of the vibration unit was recorded every five seconds. Steady state was considered reached when the impedance value changed less than the preset tolerance three times consecutively, and the percentage deviation of the steady-state impedance value from the initial reference impedance value was recorded. The impedance deviation values of all experimental samples were summarized along with their corresponding current compensation coefficients. The current compensation coefficient was defined as the ratio of the current required to achieve standard vibration output to the current driving current. The impedance deviation values were divided into 256 equally spaced intervals, ranging from 30% negative to 80% positive. Within each interval, the median current compensation coefficient of all samples was taken to form a lookup table with the impedance deviation value as input and the current compensation coefficient as output.
[0060] During real-time operation, the current impedance offset value is used as input to locate its corresponding interval in the lookup table. If the offset value falls within an interval, the current compensation coefficient for that interval is directly output. If the offset value lies between two intervals, linear interpolation is used to calculate the coefficient value. A scalar multiplication operation is performed between the current amplitude at each frequency point in the current amplitude sequence and the corresponding current compensation coefficient; the product is the fine-tuning current amplitude at that frequency point. All fine-tuning current amplitudes at all frequencies are arranged in the original frequency index order to generate a fine-tuning current sequence.
[0061] This sequence simultaneously performs closed-loop amplitude and phase correction of conduction loss and open-loop gain compensation of impedance drift in the frequency domain, enabling the drive current to adapt to the real-time electroacoustic state of the vibrating unit.
[0062] In step S17, power distribution estimation is performed based on the fine-tuned current sequence to obtain the output power spectrum. Resonance drift tracking and frequency compensation are then performed based on the output power spectrum to generate optimized vibration control parameters, including: Based on the fine-tuned current sequence, time-frequency analysis is performed using short-time Fourier transform to obtain the instantaneous peak sequence and the instantaneous frequency offset sequence; Based on the instantaneous peak sequence, instantaneous frequency offset sequence, and pre-constructed equivalent circuit model of the vibration unit, power distribution estimation is performed to obtain the output power spectrum; Based on the output power spectrum, dynamic range compression is performed using a soft limiting function to limit instantaneous peak values; The position of the main peak of the output power spectrum is identified by sliding window peak detection combined with linear regression analysis to obtain the resonant frequency offset and drift rate. The frequency compensation value is calculated using a PID algorithm based on the resonant frequency deviation and the drift rate. Based on the frequency compensation value, a spectrum shifting operation is performed on the preset reference sinusoidal drive command in the frequency domain, and combined with the preset amplitude limit value, optimized vibration control parameters including the amplitude limit value and frequency tracking amount are generated.
[0063] Specifically, the data source for step S17 is the fine-tuning current sequence output from step S16. This sequence is stored in frequency domain complex form, with each frequency point corresponding to a driving current amplitude and phase. First, a short-time Fourier transform is performed on the fine-tuning current sequence to obtain its time-frequency distribution. The fine-tuning current sequence is then transformed to the time domain via an inverse Fourier transform to obtain a time-domain current signal sequence. This sequence is divided into continuous data blocks with a fixed frame length. The frame length is determined based on the mechanical time constant of the vibration unit, which is obtained by averaging the rise times of one hundred sets of vibration unit step responses. The frame length is set to twice this average rise time to ensure that each frame contains a complete transient process. The frame shift is set to half the frame length to achieve 50% overlap.
[0064] Each frame of data is multiplied pointwise by a Hanning window, with the Hanning window coefficients generated from the frame length. After windowing, a Fast Fourier Transform is performed on each frame to obtain the complex vector of the frame's spectrum. After traversing all frames, the spectra of each frame are arranged in chronological order into a three-dimensional matrix, with dimensions of frequency points, time frames, and complex values. The frequency value corresponding to the maximum amplitude in each frame's spectrum is extracted from this matrix and arranged in frame order to generate an instantaneous peak sequence; simultaneously, the difference between the frequency corresponding to this maximum amplitude and the reference center frequency is extracted and arranged in frame order to generate an instantaneous frequency offset sequence.
[0065] The equivalent circuit model of the vibration unit is a pre-constructed linear time-invariant system model used to describe the electromechanical conversion relationship from current excitation to vibration acceleration. The model structure adopts the form of a second-order bandpass filter in series with a pure time delay. The second-order bandpass part is equivalent to a series resonant circuit composed of resistors, inductors, and capacitors. Its transfer function is uniquely determined by three parameters: resonant frequency, quality factor, and DC gain.
[0066] The parameters were obtained through an offline system identification experiment: Eighty bone conduction vibration units were selected, each rigidly fixed to an impedance head. A sweep frequency current source applied logarithmic sweep frequency excitation from low to high frequencies, and the output acceleration of the vibration units was simultaneously acquired. A fast Fourier transform was performed on each set of input and output data to calculate the amplitude-frequency response and phase-frequency response curves. The measured peak frequency of the amplitude-frequency response was used as the initial value of the resonant frequency. The quality factor was calculated using the √2 / 2 bandwidth of the peak value, and the average gain in the low-frequency band was used as the initial value of the DC gain. A nonlinear least squares algorithm was used to iteratively optimize the three parameters until the error converged, with the goal of minimizing the mean square error between the model output and the actual output spectrum. The median of the parameters for all samples was taken to obtain the standard model parameters, which were then fixed.
[0067] When using the model, the instantaneous peak frequency of the current frame is used as the model's center frequency. The power distribution near the instantaneous peak frequency is estimated by multiplying the squared current amplitude corresponding to each frame of the instantaneous peak sequence by the squared gain of the model at that frequency point to obtain the vibration power estimate for that frame. The power estimates of all frames are sorted by frequency, and the contributions of each frame are accumulated at each frequency point to generate the output power spectrum. The output power spectrum is a two-dimensional array, with frequency on the horizontal axis and power density on the vertical axis.
[0068] The output power spectrum is input into a soft-limiting function for dynamic range compression. The threshold of the soft-limiting function is determined statistically based on the maximum safe power of the headphone oscillator: incremental power excitation is applied to two hundred vibration units until the total harmonic distortion exceeds 10%, and the power value at this point is recorded as the critical power for damage. The fifth percentile of the critical power of all samples is taken as the soft-limiting threshold, ensuring that 95% of the units operate linearly within this threshold.
[0069] The soft-limiting function takes the form that when the input power is below a threshold, the output equals the input; when the input power exceeds the threshold, the output equals the threshold plus the product of the compression coefficient and the difference between the threshold and the input power. The compression coefficient is determined by an auditory masking experiment: thirty listeners are selected to rate the signals before and after compression, and the compression coefficient corresponding to the highest rating is used as the final value and fixed. Applying the soft-limiting function to the power value at each frequency point of the output power spectrum results in a compressed power spectrum, with the instantaneous peak value limited to a safe range.
[0070] The drift trend of the main peak position in the compressed output power spectrum is identified. The main peak position is defined as the frequency point corresponding to the maximum power value in the power spectrum. A sliding window peak detection method is adopted, and the window width is determined statistically based on the natural drift rate of the resonant frequency of the vibration unit: the main peak frequency of the output power spectrum of fifty headphones is recorded every ten seconds during one hour of continuous operation, the absolute value of the drift between two adjacent times is calculated, and the 95th percentile of the absolute value of the drift of all samples is taken as the half-width of the window. The total length of the window is twice the half-width plus one.
[0071] At each current time, the main peak frequency sequence of each frame within the window is extracted, and a univariate linear regression analysis is performed on this sequence and the time index. The regression slope is calculated by dividing the covariance of the independent and dependent variables by the variance of the independent variable; the slope value is the drift rate. The difference between the regression intercept and the current main peak frequency is filtered and taken as the resonant frequency offset. A positive drift rate indicates an increase in frequency, while a negative drift rate indicates a decrease in frequency.
[0072] The resonant frequency offset and drift rate are input to a proportional-integral-derivative (PID) controller to calculate the frequency compensation value. The proportional, integral, and derivative coefficients of the controller are determined through offline frequency tracking experiments: a closed-loop test platform is built to simulate resonant drift scenarios with different rates and amplitudes. With the goal of minimizing the residual frequency offset, a grid search method is used to traverse coefficient combinations within a preset range, and the coefficient combination that minimizes the root mean square residual frequency offset is selected and fixed.
[0073] During real-time calculation, the resonant frequency deviation is used as the error value, which is multiplied by the proportional coefficient to obtain the proportional term; the error value is accumulated over time and then multiplied by the integral coefficient to obtain the integral term; the difference between the current error and the error at the previous moment is divided by the sampling interval to obtain the error change rate, which is multiplied by the differential coefficient to obtain the differential term; the sum of the three terms is then limited and output, which is the frequency compensation value.
[0074] The reference sine wave drive instruction is a complex number sequence pre-stored in non-volatile memory. Its construction process is the same as the sine wave drive instruction in step S16, but it additionally stores multiple sets of reference complex numbers at different frequency points for spectrum shifting operations. Based on the frequency compensation value, the reference sine wave drive instruction is spectrum-shifted in the frequency domain: if the frequency compensation value is positive, the real and imaginary parts of the reference complex number are shifted to the left by the number of frequency points corresponding to the compensation value; if it is negative, it is shifted to the right. Missing frequency points are filled with zeros, and portions exceeding the original frequency band are truncated. The shifted complex number sequence is the frequency tracking quantity.
[0075] The preset amplitude limit is the maximum linear amplitude of the vibration unit determined during the factory calibration phase for each earphone. The testing method involves gradually increasing the drive current amplitude while monitoring the total harmonic distortion (THD). The current amplitude at which the THD reaches one percent is taken as the amplitude limit and written to the storage area. The amplitude limit and the frequency tracking sequence are encapsulated into the same data structure to generate optimized vibration control parameters.
[0076] This parameter includes both the maximum allowable amplitude after dynamic range compression and the drive frequency after tracking resonance drift, providing a closed-loop optimization benchmark for amplitude and frequency for vibration control command generation.
[0077] In step S18, parameter fusion calculation is performed based on the optimized vibration control parameters to obtain vibration control commands, and digital interface driving is performed based on the vibration control commands to obtain tactile vibration signals, including: Based on the amplitude limit value in the optimized vibration control parameters, the vibration unit current amplitude is obtained by nonlinear lookup through a preset current mapping relationship; Calculate the frequency deviation between the frequency tracking value and the standard resonant frequency, and input the frequency deviation into a preset phase deviation mapping function to obtain the phase correction value; The mechanical resonance suppression coefficient is obtained by querying a preset resonance suppression lookup table based on the frequency tracking value. The vibration unit current amplitude, the phase correction amount, and the mechanical resonance suppression coefficient are encapsulated to generate vibration control commands; The vibration control command is transmitted to the drive circuit in the bone conduction headphone hardware via a digital audio interface, driving the vibration unit to output a tactile vibration signal calibrated for amplitude, phase, and resonance characteristics.
[0078] Specifically, the data source for step S18 is the optimized vibration control parameters output in step S17. These parameters include two core components: amplitude limit value and frequency tracking quantity. The amplitude limit value is a scalar, and the frequency tracking quantity is a complex sequence in the frequency domain.
[0079] First, an amplitude limit value is extracted from the optimized vibration control parameters. This value represents the maximum allowable driving current amplitude of the vibration unit at the current moment. A pre-defined current mapping relationship, constructed using an offline nonlinear calibration experiment, is used to convert the amplitude limit value into the actual vibration unit current amplitude. The calibration experiment selects one hundred bone conduction vibration units, each fixed on a rigid fixture. A DC bias gradually increasing from zero, superimposed with a sinusoidal sweep current, is applied, while a laser vibrometer simultaneously measures the displacement amplitude of the vibration unit. At each frequency point, the current amplitude that causes the displacement amplitude to reach the mechanical limit critical point is recorded, and this critical current amplitude is taken as the maximum allowable current at that frequency point. The median of the maximum allowable currents for all samples at each frequency point is used to obtain a curve showing the current amplitude versus frequency.
[0080] Considering that bone conduction headphones primarily operate in the low-frequency range, the median of the curve in the low-frequency range is taken as the reference mapping point. When constructing the mapping table, the amplitude limit value is normalized to the zero-to-one interval. Using the normalized value as input and the measured current amplitude as output, piecewise cubic Hermitian interpolation is used to generate 256 evenly spaced mapping nodes. During real-time operation, the normalized amplitude limit value from the optimized vibration control parameters is used as input to locate its corresponding interval in the mapping table. Hermitian interpolation is then performed using the function values and derivative values of the nodes at both ends of the interval to output the corresponding vibration unit current amplitude. This current amplitude is a dimensionless digital quantity, which will subsequently be converted into a digital-to-analog converter code value for the drive circuit.
[0081] The frequency tracking sequence is a complex sequence obtained from the spectrum shifting operation in step S17. Each frequency point represents the driving frequency component after resonance drift compensation. The standard resonant frequency is a fixed value determined by impedance spectrum analysis during the headphone's factory calibration stage. During the calibration process, a sweep voltage excitation is applied to each headphone, and the current response is collected and the impedance amplitude curve is calculated. The frequency corresponding to the minimum impedance amplitude is taken as the standard resonant frequency of the headphone and written into the non-volatile memory.
[0082] During real-time operation, the peak frequency value is extracted from the frequency tracking sequence, which is the frequency corresponding to the frequency index of the maximum amplitude in the sequence. This value is the current frequency tracking value. The difference between this frequency tracking value and the standard resonant frequency is calculated. Subtracting the standard resonant frequency from the frequency tracking value yields the frequency deviation. This deviation value is a signed scalar; a positive value indicates that the current driving frequency is higher than the standard resonant frequency, and a negative value indicates that it is lower.
[0083] The phase deviation mapping function is a pre-constructed linear function model used to map frequency deviation into a phase correction amount for the drive signal. This function is constructed through offline phase response experiments: fifty bone conduction vibration units are selected, and sinusoidal drive signals of different frequencies are applied at equal intervals near the standard resonant frequency of each unit. Simultaneously, the phase difference between the output acceleration of the vibration unit and the drive current is measured. The phase difference value is recorded at each test frequency point. With frequency deviation as the independent variable and phase difference as the dependent variable, a linear equation is fitted using the least squares method for all sample data to obtain the slope and intercept. The goodness-of-fit test shows that the linear relationship holds, and the slope and intercept are fixed parameters. During real-time operation, the frequency deviation value is substituted into this linear function, multiplied by the slope, and then the intercept is added to calculate the output value, which is the phase correction amount. This correction amount represents the phase shift required to make the output vibration in phase with the drive current, expressed in radians.
[0084] A resonance suppression lookup table was constructed through offline mechanical impedance measurement experiments to suppress mechanical resonance peaks near the resonant frequency of vibrating elements. Eighty bone conduction vibrating elements were selected for the experiment. The acceleration admittance curve of each element was recorded under swept-frequency excitation to identify the quality factor at the resonant frequency. A higher quality factor indicates a sharper resonance peak, requiring stronger suppression.
[0085] When constructing the lookup table, frequency is used as the independent variable and the resonance suppression coefficient as the dependent variable. The resonance suppression coefficient is determined as follows: For each frequency point, the driving current amplitude is set as the reference value, and the notch filter depth is adjusted to flatten the acceleration response of the vibration unit at that frequency point. The attenuation amount of the notch gain that optimizes the response flatness is recorded as the resonance suppression coefficient. The coefficients of all samples at each frequency point are averaged to form a one-dimensional array indexed by frequency. The array covers the headphone's operating frequency band, and the frequency resolution is aligned with the frequency resolution of subsequent digital interface drive commands. During real-time operation, the peak frequency value of the frequency tracking quantity is used as input, and the two nearest frequency nodes are located in the resonance suppression lookup table. The mechanical resonance suppression coefficient corresponding to that frequency is calculated by linear interpolation. This coefficient is a scalar between zero and one, representing the attenuation ratio of the driving current amplitude.
[0086] The calculated vibration unit current amplitude, phase correction, and mechanical resonance suppression coefficient are fused and encapsulated. The vibration unit current amplitude serves as the driving amplitude reference. It is first multiplied by the mechanical resonance suppression coefficient to obtain the actual driving amplitude after resonance attenuation. Then, the driving waveform is phase-shifted based on the phase correction. The phase shift is achieved by adding the driving amplitude to the reference sine wave phase to generate new sine wave phase parameters.
[0087] The packaged vibration control command is a data structure containing an amplitude code value, a phase offset, and a synchronization clock tag. This command is transmitted via a digital audio interface, which uses an integrated circuit-built-in audio bus protocol with a fixed clock frequency and a data width of 24 bits. Upon receiving the vibration control command, the drive circuit sends the amplitude code value to a digital-to-analog converter to generate a corresponding analog voltage. The phase offset is used to adjust the carrier phase of the pulse width modulation driver, ultimately driving the vibration unit to output a tactile vibration signal.
[0088] The signal is strictly limited to a safe range in amplitude, achieves adaptive alignment with the resonant frequency in phase, and suppresses the mechanical resonance peak in frequency response, allowing the wearer to perceive a pure, powerful, and unblemished low-frequency vibration.
[0089] Reference Figure 2 The second embodiment of the present invention provides a signal enhancement system for bone conduction headphones, comprising: The data acquisition module is used to acquire audio input signals; The low-frequency feature module is used to perform fundamental frequency localization and harmonic extraction based on the audio input signal to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components, and to perform dynamic range statistics based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components to obtain a low-frequency feature vector; The amplitude-phase enhancement module is used to perform nonlinear curve fitting operation on harmonic intensity segments based on the low-frequency feature vector to generate a piecewise nonlinear mapping function, and to perform amplitude nonlinear mapping and phase distortion smoothing based on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence. The low-frequency enhancement module is used to extract impulse response data from the enhanced amplitude-phase sequence, generate an inverse attenuation function, calculate corrected response data, update the enhanced amplitude-phase sequence based on the corrected response data, and obtain a low-frequency enhanced signal. The conduction simulation module is used to simulate the skull conduction model based on the low-frequency enhanced signal, obtain simulated vibration transmission data, and perform standard data deviation calculation based on the simulated vibration transmission data to obtain amplitude difference sequence and phase deviation sequence. The fine-tuning current module is used to perform frequency domain compensation modulation based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence, and to perform impedance current compensation by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuning current sequence. The vibration control module is used to estimate the power distribution based on the fine-tuning current sequence, obtain the output power spectrum, and perform resonance drift tracking and frequency compensation based on the output power spectrum to generate optimized vibration control parameters. The output module is used to perform parameter fusion calculation based on the optimized vibration control parameters to obtain vibration control commands, and to drive the digital interface based on the vibration control commands to obtain tactile vibration signals.
[0090] It should be noted that the signal enhancement system for bone conduction headphones provided in this embodiment of the invention is used to execute all the process steps of the signal enhancement method for bone conduction headphones in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0091] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0092] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for signal enhancement in bone conduction headphones, characterized in that, include: Acquire audio input signal; Based on the audio input signal, fundamental frequency localization and harmonic extraction are performed to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components. Based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components, dynamic range statistics are performed to obtain low frequency feature vectors. Based on the low-frequency feature vector, a nonlinear curve fitting operation is performed on the harmonic intensity segments to generate a piecewise nonlinear mapping function. Then, based on the low-frequency feature vector and the piecewise nonlinear mapping function, amplitude nonlinear mapping and phase distortion smoothing are performed to obtain an enhanced amplitude-phase sequence. The impulse response data is extracted from the enhanced amplitude-phase sequence to generate an inverse attenuation function, and the corrected response data is calculated. The enhanced amplitude-phase sequence is then updated based on the corrected response data to obtain a low-frequency enhanced signal. Based on the low-frequency enhanced signal, a skull conduction model simulation is performed to obtain simulated vibration transmission data. Then, based on the simulated vibration transmission data, standard data deviation is calculated to obtain amplitude difference sequence and phase deviation sequence. Frequency domain compensation modulation is performed based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence. Then, impedance current compensation is performed by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuned current sequence. Power distribution is estimated based on the fine-tuned current sequence to obtain the output power spectrum, and resonance drift tracking and frequency compensation are performed based on the output power spectrum to generate optimized vibration control parameters. Based on the optimized vibration control parameters, parameter fusion calculation is performed to obtain vibration control commands, and digital interface driving is performed based on the vibration control commands to obtain tactile vibration signals.
2. The signal enhancement method for bone conduction headphones according to claim 1, characterized in that, The process involves fundamental frequency localization and harmonic extraction based on the audio input signal to obtain fundamental frequency amplitude, fundamental frequency phase, and harmonic components. Dynamic range statistics are then performed based on the fundamental frequency amplitude, fundamental frequency phase, and harmonic components to obtain a low-frequency feature vector, including: Based on the audio input signal, a filtering operation is performed using the Butterworth low-pass filtering algorithm to obtain a low-frequency signal; Based on the low-frequency signal, a time-frequency conversion operation is performed using the short-time Fourier transform algorithm to obtain the spectrum matrix; Based on the spectrum matrix, locate the frequency and complex value corresponding to the amplitude peak in each frame to obtain the fundamental frequency amplitude sequence, fundamental frequency phase sequence and fundamental frequency sequence; Based on the fundamental frequency sequence, the amplitude of the corresponding frequency point is extracted within a preset octave range to obtain the harmonic component sequence; Based on the fundamental frequency amplitude sequence, the difference between the maximum and minimum values is calculated to obtain the global amplitude range; The fundamental frequency amplitude sequence, the fundamental frequency phase sequence, the global amplitude range, and the harmonic component sequence are combined and spliced to obtain the low-frequency feature vector.
3. The signal enhancement method for bone conduction headphones according to claim 2, characterized in that, The step of performing nonlinear curve fitting based on the low-frequency feature vector through harmonic intensity segmentation to generate a piecewise nonlinear mapping function, and then performing amplitude nonlinear mapping and phase distortion smoothing based on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence includes: Based on the low-frequency feature vector, the ratio of the sum of the amplitude values of each harmonic component to the fundamental frequency amplitude is calculated to obtain the harmonic proportion sequence. The harmonic proportion sequence is then compared with the preset first proportion threshold and second proportion threshold to mark the harmonic components as weak harmonic range, medium harmonic range or strong harmonic range, and the harmonic range is determined. Based on the harmonic intervals, a set of preset gain coefficients and compression coefficients are configured for each interval, and the boundary values of the harmonic intervals are used as nodes. A continuous nonlinear curve is fitted using a cubic spline interpolation algorithm to generate a piecewise nonlinear mapping function. Based on the fundamental frequency amplitude sequence, the piecewise nonlinear mapping function is used to perform point-by-point calculations to map each fundamental frequency amplitude to a new amplitude, thereby obtaining the enhanced amplitude sequence. The amplitude difference between adjacent time frames in the enhanced amplitude sequence is calculated to obtain the amplitude change gradient. When the amplitude change gradient is greater than a preset phase distortion threshold, the fundamental frequency phase sequence is smoothed using a Hanning window weighted moving average method to obtain an adjusted phase sequence. The enhanced amplitude sequence and the adjusted phase sequence are combined according to time frames to obtain the enhanced amplitude-phase sequence.
4. The signal enhancement method for bone conduction headphones according to claim 3, characterized in that, The step of comparing the harmonic proportion sequence with preset first and second proportion thresholds to label harmonic components as weak harmonic intervals, medium harmonic intervals, or strong harmonic intervals, and determining the harmonic intervals, includes: The harmonic proportion sequence is compared with a preset first proportion threshold and a preset second proportion threshold. When the harmonic proportion is lower than the preset first proportion threshold, the corresponding harmonic component is marked as a weak harmonic range. When the harmonic proportion is between the first proportion threshold and the preset second proportion threshold, the corresponding harmonic component is marked as a medium harmonic range. When the harmonic proportion is higher than the second proportion threshold, the corresponding harmonic component is marked as a strong harmonic range, thus determining the harmonic range.
5. The signal enhancement method for bone conduction headphones according to claim 1, characterized in that, The step of extracting impulse response data from the enhanced amplitude-phase sequence, generating an inverse attenuation function, calculating corrected response data, and updating the enhanced amplitude-phase sequence based on the corrected response data to obtain a low-frequency enhanced signal includes: Based on the enhanced amplitude-phase sequence, the time point when the amplitude exceeds the preset dynamic range compression threshold is identified to obtain the abnormal peak point, and the corresponding enhanced amplitude-phase sequence within a preset time window before and after the abnormal peak point is taken as the impact response data. Calculate the difference between the amplitude of the abnormal peak point in the impact response data and the dynamic range compression threshold, and divide the difference by the dynamic range compression threshold to obtain the overshoot deviation ratio. Based on the overshoot ratio, a weight sequence that decreases linearly from the abnormal peak point to both sides is generated using a linear interpolation algorithm as a reverse decay function. The impact response data is then multiplied point by point with the reverse decay function to obtain the corrected response data. The original data in the corresponding window of the enhanced amplitude-phase sequence is replaced with the corrected response data, and the envelope of the updated enhanced amplitude-phase sequence is extracted by Hilbert transform to obtain the low-frequency envelope profile. The low-frequency envelope contour is convolved using a Gaussian smoothing kernel to obtain the low-frequency enhanced signal.
6. The signal enhancement method for bone conduction headphones according to claim 1, characterized in that, The process involves simulating the skull conduction model based on the low-frequency enhanced signal to obtain simulated vibration transmission data, and then calculating the standard data deviation based on the simulated vibration transmission data to obtain an amplitude difference sequence and a phase deviation sequence, including: Based on the low-frequency enhanced signal, the instantaneous amplitude envelope and instantaneous frequency are extracted by Hilbert transform to obtain the amplitude envelope sequence and instantaneous frequency offset; The amplitude envelope sequence and the instantaneous frequency offset are input into a pre-established skull conduction frequency response model to obtain simulated vibration transmission data. The amplitude values of the simulated vibration transmission data and the pre-stored standard bone conduction data are subtracted at the same frequency point to obtain an amplitude difference sequence. The cross-correlation function between the simulated vibration transmission data and the standard bone conduction data is calculated to obtain the phase deviation sequence.
7. The signal enhancement method for bone conduction headphones according to claim 1, characterized in that, The step of performing frequency domain compensation modulation based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence, and then performing impedance current compensation by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuned current sequence, includes: Based on the amplitude difference sequence, the gain compensation coefficient is calculated using a PID algorithm, and the phase deviation sequence is input into a preset phase-delay conversion function to calculate the delay compensation amount. The gain compensation coefficient and the time delay compensation amount are combined and encapsulated to generate frequency point compensation data; The frequency compensation data and the pre-stored sinusoidal drive command are multiplied in the frequency domain to generate a current amplitude sequence. The feedback voltage and drive current are collected, the real-time impedance value is calculated using Ohm's law, and the difference between the preset reference impedance value and the real-time impedance value is calculated to obtain the impedance offset value. The current compensation coefficient is obtained by querying the preset current compensation table according to the impedance offset value, and then the current compensation coefficient is multiplied by the current amplitude sequence to obtain the fine-tuning current sequence.
8. The signal enhancement method for bone conduction headphones according to claim 1, characterized in that, The process of estimating power distribution based on the fine-tuned current sequence to obtain the output power spectrum, and then performing resonance drift tracking and frequency compensation based on the output power spectrum to generate optimized vibration control parameters includes: Based on the fine-tuned current sequence, time-frequency analysis is performed using short-time Fourier transform to obtain the instantaneous peak sequence and the instantaneous frequency offset sequence; Based on the instantaneous peak sequence, instantaneous frequency offset sequence, and pre-constructed equivalent circuit model of the vibration unit, power distribution estimation is performed to obtain the output power spectrum; Based on the output power spectrum, dynamic range compression is performed using a soft limiting function to limit instantaneous peak values; The position of the main peak of the output power spectrum is identified by sliding window peak detection combined with linear regression analysis to obtain the resonant frequency offset and drift rate. The frequency compensation value is calculated using a PID algorithm based on the resonant frequency deviation and the drift rate. Based on the frequency compensation value, a spectrum shifting operation is performed on the preset reference sinusoidal drive command in the frequency domain, and combined with the preset amplitude limit value, optimized vibration control parameters including the amplitude limit value and frequency tracking amount are generated.
9. The signal enhancement method for bone conduction headphones according to claim 8, characterized in that, The process of performing parameter fusion calculation based on the optimized vibration control parameters to obtain vibration control commands, and then performing digital interface driving based on the vibration control commands to obtain tactile vibration signals, includes: Based on the amplitude limit value in the optimized vibration control parameters, the vibration unit current amplitude is obtained by nonlinear lookup through a preset current mapping relationship; Calculate the frequency deviation between the frequency tracking value and the standard resonant frequency, and input the frequency deviation into a preset phase deviation mapping function to obtain the phase correction value; The mechanical resonance suppression coefficient is obtained by querying a preset resonance suppression lookup table based on the frequency tracking value. The vibration unit current amplitude, the phase correction amount, and the mechanical resonance suppression coefficient are encapsulated to generate vibration control commands. The vibration control command is transmitted to the drive circuit in the bone conduction headphone hardware via a digital audio interface, driving the vibration unit to output a tactile vibration signal calibrated for amplitude, phase, and resonance characteristics.
10. A signal enhancement system for bone conduction headphones, characterized in that, include: The data acquisition module is used to acquire audio input signals; The low-frequency feature module is used to perform fundamental frequency localization and harmonic extraction based on the audio input signal to obtain fundamental frequency amplitude, fundamental frequency phase and harmonic components, and to perform dynamic range statistics based on the fundamental frequency amplitude, the fundamental frequency phase and the harmonic components to obtain a low-frequency feature vector; The amplitude-phase enhancement module is used to perform nonlinear curve fitting operation on harmonic intensity segments based on the low-frequency feature vector to generate a piecewise nonlinear mapping function, and to perform amplitude nonlinear mapping and phase distortion smoothing based on the low-frequency feature vector and the piecewise nonlinear mapping function to obtain an enhanced amplitude-phase sequence. The low-frequency enhancement module is used to extract impulse response data from the enhanced amplitude-phase sequence, generate an inverse attenuation function, calculate corrected response data, update the enhanced amplitude-phase sequence based on the corrected response data, and obtain a low-frequency enhanced signal. The conduction simulation module is used to simulate the skull conduction model based on the low-frequency enhanced signal, obtain simulated vibration transmission data, and perform standard data deviation calculation based on the simulated vibration transmission data to obtain amplitude difference sequence and phase deviation sequence. The fine-tuning current module is used to perform frequency domain compensation modulation based on the amplitude difference sequence and the phase deviation sequence to generate a current amplitude sequence, and to perform impedance current compensation by calculating the impedance offset based on the current amplitude sequence to obtain a fine-tuning current sequence. The vibration control module is used to estimate the power distribution based on the fine-tuning current sequence, obtain the output power spectrum, and perform resonance drift tracking and frequency compensation based on the output power spectrum to generate optimized vibration control parameters. The output module is used to perform parameter fusion calculation based on the optimized vibration control parameters to obtain vibration control commands, and to drive the digital interface based on the vibration control commands to obtain tactile vibration signals.