Self-adaptive conversation noise reduction system based on bone conduction and air conduction dual-mode fusion

By using an adaptive call noise reduction system that integrates bone conduction and air conduction modes, the problems of anti-interference and natural sound quality in headphone call noise reduction technology have been solved, achieving efficient call clarity and sound quality improvement in complex scenarios.

CN121908179APending Publication Date: 2026-04-21陵水安声科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
陵水安声科技有限公司
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing headphone call noise reduction technology suffers from the problem that a single transmission method cannot simultaneously achieve both anti-interference and natural sound quality, and lacks dynamic scene adaptation, resulting in fluctuations in call clarity in complex scenarios.

Method used

An adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion is adopted. Through dual-mode signal acquisition module, scene dynamic classification and weight allocation module, and dual-mode noise suppression and signal enhancement algorithm, hardware coordination and algorithm scheduling are achieved to dynamically adjust signal weight, specifically suppress noise and compensate for high-frequency components.

Benefits of technology

Significantly improves call quality in complex scenarios, with strong anti-interference capabilities, natural sound quality, dynamic adaptation to different environments and motion states, noise reduction of 15-20dB, transient noise suppression accuracy of ≥95%, high-frequency compensation to ensure clear speech, and reduced echo.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908179A_ABST
    Figure CN121908179A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of call noise reduction, in particular to a self-adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion, and adopts the technical scheme that a bone conduction sensor acquires a low-frequency voice signal generated by vocal cord vibration, and an air conduction acquisition unit enhances a voice signal of a main microphone; in bone conduction vibration noise suppression, vibration noise is stripped from bone conduction signals, air conduction environment noise suppression separates voice and residual noise, high-frequency compensation is superposed to a high-frequency missing area of the bone conduction signals, phase alignment carries out delay compensation on lagged signals, and dynamic gain balance adjusts signal gains. Dynamic fusion and noise reduction optimization of bone conduction and air conduction dual-mode signals are realized through hardware cooperation and algorithm scheduling, the advantages of environmental noise resistance of the bone conduction signals and high-frequency and abundant advantages of the air conduction signals are utilized, the weights of the two signals are dynamically distributed in combination with real-time characteristics of a scene, and the noise reduction and noise reduction of the two signals are realized. Bone conduction vibration noise and air conduction environment noise are restrained in a targeted mode, and finally clear and natural communication voice is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of call noise reduction technology, and in particular to an adaptive call noise reduction system based on the fusion of bone conduction and air conduction dual modes. Background Technology

[0002] Headphone call noise reduction technology uses microphone arrays, algorithm optimization, and environmental awareness to eliminate background noise during calls and improve voice clarity. Current headphone call noise reduction technologies suffer from two main technical problems: limitations of a single transmission method and lagging scene adaptation. These are detailed below: 1. A single conduction method cannot simultaneously guarantee anti-interference and natural sound quality. Air conduction microphones are easily affected by strong winds and transient noise (such as subway announcements and screams from the crowd), which can cause the voice signal to be submerged. Although bone conduction microphones can isolate environmental noise, they lack high-frequency components of the voice they collect (the sound is muffled) and are easily mixed with non-voice bone conduction noise such as chewing and walking vibrations.

[0003] 2. The scene adaptation lacks dynamism. Existing technologies can only adapt to the scene by using fixed parameters or manually switching modes. They cannot automatically adjust the noise reduction strategy according to the real-time environmental noise type (steady state, transient, strong wind) and the user's movement state (stationary, walking, running), resulting in fluctuations in call clarity in complex scenarios (such as outdoor cycling, walking while talking). Summary of the Invention

[0004] The purpose of this invention is to provide an adaptive call noise reduction system based on the fusion of bone conduction and air conduction dual modes, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion, comprising a dual-mode signal acquisition module, a scene dynamic classification and weight allocation module, and a dual-mode noise suppression and signal enhancement algorithm; wherein, the dual-mode signal acquisition module includes a bone conduction acquisition unit and an air conduction acquisition unit, which are triggered collaboratively; the scene dynamic classification and weight allocation module performs multi-dimensional scene feature extraction and adopts a dynamic weight allocation algorithm, the multi-dimensional scene feature extraction including environmental noise type identification, user motion state identification, and data fusion; the dual-mode noise suppression and signal enhancement algorithm performs bone conduction vibration noise suppression, air conduction environmental noise suppression, and dual-mode signal fusion enhancement, the dual-mode signal fusion enhancement performing high-frequency compensation, phase alignment, and dynamic gain balancing.

[0006] Furthermore, in the dual-mode signal acquisition module, for the bone conduction acquisition unit, a piezoelectric miniature bone conduction sensor is integrated at the contact point between the earphone earbud or ear hook and the temporal bone, and a miniature accelerometer is simultaneously built in as a vibration noise detection submodule.

[0007] Furthermore, in the dual-mode signal acquisition module, for the air conduction acquisition unit, a dual-microphone directional array is set at the headphone microphone opening, wherein the main microphone faces the user's mouth and the reference microphone faces outward.

[0008] Furthermore, in the dual-mode signal acquisition module, the triggering logic for collaborative triggering includes: when the bone conduction sensor detects that the vocal cord vibration characteristic value is greater than or equal to the voice activation threshold, the dual-mode acquisition unit is automatically activated; in non-call state, both dual-mode units are in a listening mode with power consumption less than or equal to 50μA.

[0009] Furthermore, in the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction and environmental noise type identification, the noise spectrum is analyzed by air conduction microphone array, and the noise type is divided into steady-state noise, transient noise and strong wind noise according to the spectrum.

[0010] Furthermore, in the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction and user motion state recognition, the motion acceleration is calculated by the accelerometer, and the motion state is divided into stationary, slow motion, and fast motion according to the acceleration.

[0011] Furthermore, in the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction, data fusion inputs noise type and motion state into the classification model and outputs 8 core scene labels; among them, the 8 core scene labels include indoor office, subway commuting, outdoor walking, outdoor cycling, running, restaurant, airport, and quiet home.

[0012] Furthermore, in the scene dynamic classification and weight allocation module, the weight logic of the dynamic weight allocation algorithm includes: adjusting the fusion weight of bone conduction signal and air conduction signal in real time according to the scene label.

[0013] Furthermore, in the dual-mode noise suppression and signal enhancement algorithm, for bone conduction vibration noise suppression, an adaptive filtering algorithm is adopted. The non-speech vibration signal collected by the accelerometer is used as reference noise and input with the bone conduction speech signal into the filtering model. The spectral characteristics of vibration noise are learned in real time to generate an inverse cancellation signal and remove vibration noise from the bone conduction signal. For air conduction environmental noise suppression, beamforming and deep learning noise reduction are combined to perform dual processing. First, beamforming is used to weaken the side noise, and then the main microphone signal is input into the pre-trained LSTM speech separation model to separate speech from residual noise.

[0014] Furthermore, in the dual-mode noise suppression and signal enhancement algorithm, for high-frequency compensation of dual-mode signal fusion enhancement, high-frequency components above 3kHz in the air conduction signal are extracted and superimposed onto the high-frequency missing region of the bone conduction signal through a spectrum interpolation algorithm; for phase alignment of dual-mode signal fusion enhancement, the time difference between the two types of signals is calculated through cross-correlation analysis, and delay compensation is performed on the lagging signal; for dynamic gain balancing of dual-mode signal fusion enhancement, the signal gain is adjusted according to the scene noise intensity.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention achieves dynamic fusion and noise reduction optimization of bone conduction and air conduction dual-mode signals through hardware collaboration and algorithm scheduling. Leveraging the advantages of bone conduction signals in resisting environmental noise and the high-frequency richness of air conduction signals, and combining real-time scene characteristics, it dynamically allocates the weights of the two signals while specifically suppressing bone conduction vibration noise and air conduction environmental noise, ultimately outputting clear and natural speech. Specifically, the bone conduction sensor collects low-frequency speech signals generated by vocal cord vibration, while the accelerometer specifically captures non-speech vibrations during chewing and walking, providing reference signals for subsequent noise suppression and avoiding the vibration-induced distortion of bone conduction signals found in existing technologies. The problem of noise pollution is addressed by using beamforming technology in the air conduction acquisition unit to enhance the voice signal from the main microphone and reduce lateral ambient noise. Simultaneously, it references ambient noise collected by the microphone to provide an environmental baseline for subsequent noise reduction, thus solving the problem of weak resistance to transient and strong wind noise in existing air conduction solutions. Bone conduction vibration noise suppression learns the spectral characteristics of vibration noise in real time, generating an inverse cancellation signal to remove vibration noise from the bone conduction signal, achieving a noise reduction of 15-20 dB, thus solving the vibration pollution problem of existing bone conduction solutions. Air conduction ambient noise suppression first reduces lateral noise through beamforming, then inputs the main microphone signal into a pre-trained LSTM. The speech separation model separates speech from residual noise, achieving a transient noise suppression accuracy of ≥95%, thus addressing the weakness of existing air conduction schemes in resisting transient noise. High-frequency compensation extracts high-frequency components above 3kHz from the air conduction signal and superimposes them onto the high-frequency missing regions of the bone conduction signal using a spectral interpolation algorithm, ensuring that consonants in the speech are clearly distinguishable. Phase alignment calculates the time difference between the two types of signals through cross-correlation analysis, compensating for the delay of the lagging signal to avoid an echo effect after fusion. Dynamic gain balancing adjusts the signal gain according to the noise intensity of the scene, ensuring stable speech volume after fusion. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the adaptive call noise reduction system based on the fusion of bone conduction and air conduction modes of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0018] The inventors discovered through research that the mainstream headphone call noise reduction technologies in the industry currently fall into two categories, with the specific implementation schemes as follows: For the air conduction single-mode solution: On the hardware side, 1-2 omnidirectional or directional MEMS microphones are used to enhance the speech signal in the direction of the user's mouth through beamforming technology and reduce the ambient noise from the side; on the algorithm side, adaptive filtering or deep learning noise reduction models (such as LSTM-based speech separation algorithms) are combined to suppress steady-state noise (such as air conditioner noise and engine noise), but the suppression effect on transient noise (such as car horns) and strong wind noise is limited.

[0019] For the bone conduction single-mode solution: On the hardware side, a piezoelectric bone conduction sensor is integrated at the contact point between the earphone and the skull (such as the concha or temporal bone) to directly collect the speech signal transmitted through the skull by the vibration of the vocal cords, and isolate the environmental noise transmitted through the air; the algorithm removes high-frequency interference in the bone conduction signal through low-pass filtering, but cannot supplement the inherent high-frequency components of speech (such as the consonants "s" and "sh"), and cannot effectively filter out non-speech vibration noise generated by chewing and walking.

[0020] Some high-end headphones use a simple superposition of air conduction and bone conduction, but they only fuse the signals in a fixed ratio (such as 40% bone conduction and 60% air conduction) without dynamically adjusting the weights according to the scenario or addressing the inherent defects of the two types of signals.

[0021] For air conduction single-mode solutions: In strong wind and transient noise scenarios, the ambient noise and speech signal spectrum overlap, and the noise reduction algorithm is prone to accidentally deleting speech details, resulting in the other party not being able to hear clearly; In open outdoor environments, the speech signal is easily affected by air attenuation, and after the ambient noise is superimposed, the call signal-to-noise ratio (SNR) can be as low as 5-8dB (the industry standard is above 15dB); Therefore, there are problems of weak anti-interference ability and poor outdoor applicability.

[0022] For bone conduction single-mode solutions: high-frequency components (above 3kHz) are missing, making the speech sound dull, unclear, and with low recognizability; when users move (such as running) or chew, the vibration of the skull is collected by the bone conduction sensor, forming low-frequency noise of 200-500Hz, which masks the speech signal; therefore, there are problems of severe sound quality distortion and vibration noise pollution.

[0023] For a simple superposition of air conduction and bone conduction: a fixed fusion ratio cannot cope with variable scenarios. For example, in low-noise indoor scenarios, a high proportion of bone conduction is still relied upon, resulting in redundant sound quality. In strong wind outdoor scenarios, the proportion of air conduction is too high, resulting in insufficient anti-interference. The problem of time difference and phase deviation between the two types of signals is not solved, and echoes, frequency band breaks, and fragmented listening experience are likely to occur after fusion. Therefore, there are problems of rigid scene adaptation and rough signal fusion.

[0024] In view of this, we propose an adaptive call noise reduction system based on the fusion of bone conduction and air conduction dual modes to solve the existing problems. Example 1

[0025] This application is mainly applied to products such as TWS in-ear headphones and ear-hook sports headphones. It can be integrated into the existing headphone hardware architecture to improve call quality in all scenarios, especially for high-frequency call scenarios such as commuting, outdoor sports, and daily office work.

[0026] like Figure 1 As shown, the adaptive call noise reduction system based on the fusion of bone conduction and air conduction dual modes includes a dual-mode signal acquisition module, a scene dynamic classification and weight allocation module, and a dual-mode noise suppression and signal enhancement algorithm.

[0027] The dual-mode signal acquisition module includes a bone conduction acquisition unit and an air conduction acquisition unit, which are triggered in tandem.

[0028] For the bone conduction acquisition unit, a piezoelectric miniature bone conduction sensor is integrated at the contact point between the earphone earbud or ear hook and the temporal bone, and a miniature accelerometer is simultaneously built in as a vibration noise detection submodule. The piezoelectric miniature bone conduction sensor measures 4mm × 3mm × 1mm and has a sensitivity of -26dBV / Pa; the miniature accelerometer has a sampling rate of 1kHz and a measurement range of ±2g. The bone conduction sensor acquires low-frequency speech signals generated by vocal cord vibration (resistant to environmental noise), while the accelerometer specifically captures non-speech vibrations during chewing and walking (such as 1-3Hz skull vibrations during walking), providing a reference signal for subsequent noise suppression and avoiding the problem of vibration contamination of bone conduction signals in existing technologies.

[0029] For the air conduction acquisition unit, a dual-microphone directional array is set at the headphone microphone opening, with the main microphone facing the user's mouth and the reference microphone facing outwards. The spacing between the dual-microphone directional arrays is 5mm, using AAC Technologies MEMS microphones with a frequency response range of 20Hz-20kHz. Beamforming technology (e.g., a beamwidth of 30° and a sidelobe suppression ratio ≥18dB) is used to enhance the voice signal of the main microphone and reduce lateral ambient noise. At the same time, the reference microphone collects ambient noise, providing an environmental benchmark for subsequent noise reduction and solving the problem of weak resistance to transient and strong wind noise in existing air conduction solutions.

[0030] The triggering logic of the collaborative triggering includes: when the bone conduction sensor detects a vocal cord vibration characteristic value ≥0.3g (voice activation threshold), the dual-mode acquisition unit is automatically activated; in non-call state, both dual-mode units are in low-power listening mode (power consumption ≤50μA) to avoid battery drain, which is different from the problem of excessive power consumption caused by continuous acquisition in existing technologies.

[0031] The scene dynamic classification and weight allocation module extracts multi-dimensional scene features and employs a dynamic weight allocation algorithm. Multi-dimensional scene feature extraction includes environmental noise type identification, user motion state identification, and data fusion.

[0032] Environmental noise type identification involves analyzing the noise spectrum using an air-conducting microphone array. Based on the spectrum, noise types are categorized into steady-state noise, transient noise, and strong wind noise. For steady-state noise, the spectral fluctuation is ≤5dB, such as air conditioner noise; for transient noise, the spectral peak is ≥20dB and the duration is <0.5s, such as car horns; for strong wind noise, the frequency range of broadband random noise is 100-500Hz, and the frequency band energy accounts for ≥60%.

[0033] User motion state recognition calculates motion acceleration using an accelerometer and categorizes motion states into stationary, slow-moving, and fast-moving states based on acceleration. An acceleration ≤0.1g is considered a stationary state; for slow-moving motion, 0.1g < acceleration ≤0.5g, such as walking; for fast-moving motion, acceleration >0.5g, such as running.

[0034] Data fusion inputs noise type and motion state into the classification model and outputs 8 core scene labels; the 8 core scene labels include indoor office, subway commuting, outdoor walking, outdoor cycling, running, restaurant, airport, and quiet home.

[0035] The dynamic weight allocation algorithm's weight logic includes: adjusting the fusion weights of bone conduction signal B and air conduction signal A in real time based on scene labels. The fusion weight of bone conduction signal B is Wb, and the fusion weight of air conduction signal A is Wa, with Wa + Wb = 1. The specific weight strategy is shown in the table below, which differs from the limitations of fixed weights in existing technologies.

[0036] The dual-mode noise suppression and signal enhancement algorithm performs bone conduction vibration noise suppression, air conduction environmental noise suppression, and dual-mode signal fusion enhancement. The dual-mode signal fusion enhancement performs high-frequency compensation, phase alignment, and dynamic gain balancing.

[0037] For bone conduction vibration noise suppression, an adaptive filtering algorithm is adopted. The non-speech vibration signal collected by the accelerometer is used as the reference noise and input into the filtering model along with the bone conduction speech signal. The spectral characteristics of vibration noise (such as periodic noise of 200-500Hz during running) are learned in real time to generate an inverse cancellation signal, which removes vibration noise from the bone conduction signal. The noise reduction can reach 15-20dB, solving the vibration contamination problem of existing bone conduction solutions.

[0038] For air conduction environmental noise suppression, beamforming and deep learning noise reduction are combined in an algorithm for dual processing. First, beamforming weakens side noise, and then the main microphone signal is input into a pre-trained LSTM speech separation model (training data contains 100,000+ environmental noise samples) to separate speech from residual noise. The transient noise suppression accuracy is ≥95%, which solves the problem of weak transient noise resistance in existing air conduction solutions.

[0039] For high-frequency compensation in dual-mode signal fusion enhancement, high-frequency components above 3kHz in the air conduction signal are extracted and superimposed onto the high-frequency missing region of the bone conduction signal through a spectrum interpolation algorithm to ensure that consonants (such as "s" and "sh") in speech are clearly distinguishable.

[0040] For phase alignment in dual-mode signal fusion enhancement, the time difference (0.5-2ms) between the two types of signals is calculated through cross-correlation analysis, and delay compensation is performed on the lagging signal to avoid echo after fusion.

[0041] For dynamic gain balancing of dual-mode signal fusion enhancement, the signal gain is adjusted according to the noise intensity of the scene. For example, in outdoor scenes, the bone conduction gain is increased by 3-5dB to avoid being masked by environmental noise and to ensure stable voice volume after fusion.

[0042] In summary, this application achieves the complementary advantages of bone conduction anti-interference and air conduction natural sound quality through three major modules: hardware collaborative acquisition, scene dynamic classification, and dual-mode noise suppression and fusion.

[0043] Compared to competitors' single-mode and simple overlay solutions, this application significantly improves call quality in complex scenarios. For example, in outdoor cycling (strong wind) scenarios, the signal-to-noise ratio (SNR) reaches 22-25dB (industry average 10-12dB), and in running scenarios, vibration noise suppression reaches 18dB (industry average 8-10dB). Subjective listening tests verify that the high-frequency clarity of indoor speech is improved by 40%, demonstrating core performance advantages. Users do not need to manually switch noise reduction modes; the system automatically adapts to all scenarios, solving the pain points of unclear outdoor calls and muffled indoor call quality. It especially meets the high-frequency call needs of commuters and sports enthusiasts, upgrading the user experience. The hardware is based on the existing headphone architecture, with only the addition of a miniature bone conduction sensor, increasing the cost by ≤5 yuan. The algorithm can run on existing DSP chips without replacing the main control chip. Compared to solutions that redesign the hardware, it has lower mass production costs and a faster deployment cycle. Example

[0044] This embodiment provides a detailed explanation of the specific working method of the dual-mode noise suppression and signal enhancement algorithm.

[0045] For bone conduction vibration noise suppression, a bone conduction sensor acquires a bone conduction speech signal mixed with vibration noise, and an accelerometer acquires a reference signal for the vibration noise. The signal processing module performs the following steps: preprocessing the bone conduction speech signal and the reference signal; inputting the two preprocessed signals into an adaptive filtering model; using the adaptive filtering model, learning the spectral characteristics of the vibration noise in real time based on the reference signal, and generating a corresponding inverse cancellation signal; subtracting the inverse cancellation signal from the bone conduction speech signal, and outputting the noise-reduced bone conduction speech signal. The signal processing module is also configured to: decompose the bone conduction speech signal and the reference signal into multiple non-overlapping or partially overlapping frequency sub-bands through a sub-band decomposition unit; the adaptive filtering model includes multiple parallel sub-band adaptive filters, each sub-band adaptive filter processing the signal within its corresponding sub-band.

[0046] The signal processing module also includes a speech activity detection unit, used to determine whether the bone conduction speech signal of the current frame contains valid speech. When the speech activity detection unit determines that there is no speech, it controls the sub-band adaptive filter to update the filter coefficients according to a first strategy; when it determines that there is speech, it controls the sub-band adaptive filter to update the filter coefficients according to a second strategy. The coefficient update rate under the first strategy is higher than that under the second strategy. The first strategy updates the filter coefficients with a normal step size, and the second strategy is one of the following operations: freezing the filter coefficients, or updating the filter coefficients with a step size significantly smaller than the normal step size.

[0047] The adaptive filtering model employs a variable step-size adaptive filtering algorithm, with the step-size factor adaptively adjusted according to the magnitude of the filtering error signal. Bone conduction vibration noise suppression utilizes at least two accelerometers positioned at different spatial locations to acquire multiple reference signals. The adaptive filtering model is a multi-input adaptive filtering model used to fuse multiple reference signals to generate a reverse cancellation signal. The adaptive filtering model is a nonlinear adaptive filter, which establishes a nonlinear mapping relationship between the reference signal and vibration noise through a neural network or Volterra series.

[0048] For air conduction noise suppression, noisy speech signals are received through a microphone array. A deep neural network model is used to process the microphone array signals to generate beamforming weights, and beamforming is performed on the microphone array signals based on these weights to obtain a pre-enhanced speech signal. The deep neural network model is trained to optimize the beamforming weights to maximize the signal-to-noise ratio of the pre-enhanced speech signal. The pre-enhanced speech signal is then input into a pre-trained speech separation model, which is built based on a recurrent neural network or Transformer architecture, to separate the target speech signal from the pre-enhanced speech signal and residual noise. The separated target speech signal is then output.

[0049] During the speech separation process, spatial feature parameters generated in the beamforming stage are also used as auxiliary inputs to the speech separation model. These spatial feature parameters include at least one of the following: source arrival direction estimation, coherence coefficient between microphone signals, and spatial covariance matrix characteristics of the noise field. Air-conducted environmental noise suppression also utilizes the signal or intermediate features output by the speech separation model to generate control feedback signals for the beamforming module. Based on these control feedback signals, the internal parameters or output weights of the deep neural network beamforming model are dynamically adjusted.

[0050] The speech separation model is a convolutional recurrent network model. Its front end extracts the time-frequency local features of the input signal through convolutional layers, while the back end captures long-term dependencies in the time series through recurrent neural network layers. The speech separation model is also a Transformer model based on a self-attention mechanism, which dynamically focuses on the information region for speech reconstruction by calculating the correlation weights between time-frequency points. Air conduction environmental noise suppression also includes an online adaptive step: during device operation, pure noise segments and / or clean speech segments are detected and acquired; the acquired segments are used to adjust the speech separation model and / or deep neural network beamforming model online to adapt to the current usage environment.

[0051] For high-frequency compensation in dual-mode signal fusion enhancement, bone conduction speech signals and air conduction speech signals are acquired synchronously; high-frequency components above a first preset frequency threshold are extracted from the air conduction speech signal; perceptual weighting is applied to the high-frequency components to generate a weighted high-frequency compensation signal; wherein, the perceptual weighting weights are dynamically calculated based on the human ear auditory perception model and / or the instantaneous signal-to-noise ratio of the air conduction signal in the high-frequency band; the weighted high-frequency compensation signal and the bone conduction speech signal are fused in the frequency domain to generate a high-frequency enhanced bone conduction speech signal.

[0052] After the extraction step and before the perceptual weighting step, speech enhancement processing is performed on the high-frequency components to suppress ambient noise and obtain clean high-frequency components; the perceptual weighting process targets the clean high-frequency components. Speech enhancement processing is implemented using a pre-trained deep learning model trained to decompose the high-frequency components of the air conduction signal into clean speech components and noise components.

[0053] The weight calculation for perceptual weighting is as follows: In sub-bands with a signal-to-noise ratio (SNR) higher than a second preset threshold, higher fusion weights are assigned; in sub-bands with an SNR lower than a third preset threshold, lower fusion weights or zero weights are assigned. Before the fusion step, the relative time delay between the bone conduction speech signal and the air conduction speech signal is estimated; based on the relative time delay, the weighted high-frequency compensation signal is phase-aligned in the time or frequency domain. The fusion step includes: estimating the spectral envelope of the bone conduction speech signal; and adaptively shaping the amplitude spectrum of the weighted high-frequency compensation signal based on the spectral envelope of the bone conduction speech signal to match the low-to-mid-frequency energy attenuation trend of the bone conduction signal.

[0054] For phase alignment enhancement of dual-mode signal fusion, the first and second signals are acquired. The relative time delay between the first and second signals is calculated based on the generalized cross-correlation function. The calculation of the generalized cross-correlation function includes: calculating the cross-power spectrum of the two signals, applying a frequency domain weighting function to the cross-power spectrum, and then transforming the weighted result back to the time domain to obtain the generalized cross-correlation function; performing curve interpolation near the peak of the generalized cross-correlation function, and determining the subsampling precision time delay estimate based on the peak of the interpolated curve; and performing delay compensation on the lagging signal based on the time delay estimate to achieve phase alignment of the two signals.

[0055] The frequency domain weighting function is a phase transform weighting function. Curve interpolation is parabolic interpolation or sinc function interpolation. The signal is streamed in frames, and the delay estimation and compensation steps are executed independently for each frame to track time-varying delays. The steps are: calculating the confidence level of the delay estimation result for the current frame; when the confidence level is lower than a preset threshold, one of the following strategies is adopted: using the effective delay estimate from the previous frame, re-estimating the delay based on the non-cross-correlation characteristics of the signal, or outputting an alignment failure flag. The confidence level is calculated based on at least one of the following factors: the ratio of the main peak amplitude to the mean of the generalized cross-correlation function, the ratio of the main peak amplitude to the second-highest peak amplitude, and the sharpness of the main peak.

[0056] For dynamic gain balancing in dual-mode signal fusion enhancement, the current noise intensity of the input signal is estimated; the target gain is calculated based on the current noise intensity and the target level; the response time constant for smooth transition from the current gain to the target gain is adaptively adjusted based on the characteristics of the input signal; wherein, when a sharp increase in noise intensity is detected, a first shorter time constant is used; when a decrease in noise intensity is detected or during speech activity, a second longer time constant is used; the adjusted time constant is used to smooth the gain change and is applied to the input signal.

[0057] The steps for estimating the current noise intensity include: decomposing the input signal into multiple frequency sub-bands; independently estimating the noise intensity of each sub-band; assigning a perceptual weight to each sub-band based on the human auditory perception model; and combining the weighted noise intensities of each sub-band into the current noise intensity of the entire band. When estimating the noise intensity within each sub-band, a recursive averaging algorithm based on speech activity detection is used, updating the noise estimate only in frames determined to have no speech.

[0058] The steps for calculating the target gain include: setting a lower threshold and an upper threshold for noise intensity; when the current noise intensity is below the lower threshold, the target gain is a first fixed value; when the current noise intensity is above the upper threshold, the target gain is a second fixed value; when the current noise intensity is between the upper and lower thresholds, the target gain is calculated by interpolation between the first and second fixed values. After applying the smoothed gain to the input signal, the level of the output signal after gain adjustment is monitored; when the output signal level exceeds the preset maximum level, a limiter is activated to limit the signal peak value, wherein the limiter uses a soft clipping algorithm. The steps for calculating the target gain are based on a multi-dimensional feature set, obtained through a pre-trained mapping function or machine learning model. The multi-dimensional feature set includes at least two of the following: instantaneous signal-to-noise ratio, historical gain value, input signal level, and signal spectral flatness.

[0059] The above specific embodiments are merely several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. An adaptive call noise reduction system based on the fusion of bone conduction and air conduction dual modes, characterized in that: It includes a dual-mode signal acquisition module, a scene dynamic classification and weight allocation module, and a dual-mode noise suppression and signal enhancement algorithm. The dual-mode signal acquisition module includes a bone conduction acquisition unit and an air conduction acquisition unit, which are triggered collaboratively. The scene dynamic classification and weight allocation module performs multi-dimensional scene feature extraction and adopts a dynamic weight allocation algorithm. The multi-dimensional scene feature extraction includes environmental noise type identification, user motion state identification, and data fusion. The dual-mode noise suppression and signal enhancement algorithm performs bone conduction vibration noise suppression, air conduction environmental noise suppression, and dual-mode signal fusion enhancement. The dual-mode signal fusion enhancement performs high-frequency compensation, phase alignment, and dynamic gain balancing.

2. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that: In the dual-mode signal acquisition module, for the bone conduction acquisition unit, a piezoelectric miniature bone conduction sensor is integrated at the contact point between the earphone earbud or ear hook and the temporal bone, and a miniature accelerometer is simultaneously built in as a vibration noise detection submodule.

3. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that: In the dual-mode signal acquisition module, for the air conduction acquisition unit, a dual-microphone directional array is set at the headphone microphone opening, wherein the main microphone faces the user's mouth and the reference microphone faces outward.

4. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that, In the dual-mode signal acquisition module, the triggering logic of the coordinated triggering includes: when the bone conduction sensor detects that the characteristic value of vocal cord vibration is greater than or equal to the voice activation threshold, the dual-mode acquisition unit is automatically activated; in non-call state, both dual-mode units are in a listening mode with power consumption less than or equal to 50μA.

5. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that: In the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction and environmental noise type identification, the noise spectrum is analyzed by air conduction microphone array, and the noise type is divided into steady-state noise, transient noise and strong wind noise according to the spectrum.

6. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that: In the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction, user motion state recognition calculates motion acceleration through an accelerometer and classifies motion state into stationary, slow motion, and fast motion based on acceleration.

7. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion according to claim 1, characterized in that: In the scene dynamic classification and weight allocation module, for multi-dimensional scene feature extraction, data fusion inputs noise type and motion state into the classification model and outputs 8 core scene labels; among them, the 8 core scene labels include indoor office, subway commuting, outdoor walking, outdoor cycling, running, restaurant, airport, and quiet home.

8. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion according to claim 1, characterized in that, In the scene dynamic classification and weight allocation module, the weight logic of the dynamic weight allocation algorithm includes: adjusting the fusion weight of bone conduction signal and air conduction signal in real time according to the scene label.

9. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion as described in claim 1, characterized in that: In the dual-mode noise suppression and signal enhancement algorithm, for bone conduction vibration noise suppression, an adaptive filtering algorithm is adopted. The non-speech vibration signal collected by the accelerometer is used as the reference noise and input with the bone conduction speech signal into the filtering model. The spectral characteristics of vibration noise are learned in real time to generate an inverse cancellation signal and remove vibration noise from the bone conduction signal. For suppressing air-conducted environmental noise, beamforming and deep learning noise reduction are combined in an algorithm for dual processing. First, beamforming is used to reduce side noise, and then the main microphone signal is input into a pre-trained LSTM speech separation model to separate speech from residual noise.

10. The adaptive call noise reduction system based on bone conduction and air conduction dual-mode fusion according to claim 1, characterized in that: In the dual-mode noise suppression and signal enhancement algorithm, for high-frequency compensation of dual-mode signal fusion enhancement, high-frequency components above 3kHz in the air conduction signal are extracted and superimposed onto the high-frequency missing region of the bone conduction signal through a spectrum interpolation algorithm; for phase alignment of dual-mode signal fusion enhancement, the time difference between the two types of signals is calculated through cross-correlation analysis, and delay compensation is performed on the lagging signal; for dynamic gain balancing of dual-mode signal fusion enhancement, the signal gain is adjusted according to the scene noise intensity.