Real-time bone-air fusion communication glasses and methods with narrowband noise tracking cancellation

CN118050916BActive Publication Date: 2026-09-01NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410177234.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-08
Publication Date
2026-09-01
Estimated Expiration
2044-02-08

AI Technical Summary

Technical Problem

[0012]目前市场上的产品都是单独使用骨传导技术或者气导,在对语音信号的提取有着不少的干扰因素,例如每个人的发音都独特的,中国每个地区的方言以及传感器与人体摩擦产生的噪声,这些都对语音的识别产生了干扰

Benefits of technology

[0070]第一、骨振动噪声追踪提取通过一个不与人体接触的振动传感器实时捕获穿透噪声,线性预测滤波器连接的并行结构优化输出信号和气导语音;并重新生成误差信号以监督在线骨-气融合部分的控制器更新,从而降低骨-气相关噪声的影响,提高语音增强质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118050916B_ABST
    Figure CN118050916B_ABST
Patent Text Reader

Abstract

This invention relates to a real-time bone-air fusion communication glasses and method with narrowband noise tracking and cancellation, belonging to the field of bone-air guided speech communication technology. The glasses include a left bone-guided microphone, a right bone-guided microphone, a reference noise bone-guided microphone, and an air-guided microphone connected to a main controller. The method introduces a reference bone-guided noise channel and reconstructs the error feedback loop of the traditional system. Through a bone vibration noise tracking and extraction module, penetration noise is captured in real time. A parallel structure connected to a linear prediction filter optimizes the output signal, and an error signal is regenerated to supervise the controller update of the online bone-air fusion section, thereby reducing the impact of bone-air related noise and improving speech enhancement quality. Parameter constraints are imposed on the real-time bone fusion training system through signal energy estimation, improving the system's robustness in strong narrowband noise environments. This system does not rely on bone-guided speech and air-guided speech databases and can be widely applied to various language environments, including dialects and minority languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to real-time bone-air fusion communication glasses and methods with narrowband noise tracking cancellation, belonging to the field of bone-air guided speech communication technology. Background Technology

[0002] With the development of technology and the widespread use of electronic products, children are exposed to electronic devices, especially mobile phones, shortly after birth. This has led to a common problem—the age of people wearing glasses is getting younger and the number of people wearing glasses is increasing. Furthermore, industrial and agricultural production environments are filled with noise pollution, affecting workers' normal communication. Therefore, there is a great need for smart glasses that can achieve communication while eliminating the impact of noise on communication quality. Currently, existing products use bone conduction microphones or air conduction microphones built into the temples for voice communication. However, bone conduction microphones are in close contact with the head, and if the user moves during communication, friction is inevitable between the head and the glasses. This friction generates noise, affecting communication quality. Moreover, everyone has different accents, especially in China where each region has its own unique dialect, which further complicates voice signal extraction.

[0003] Our designed smart glasses utilize a bone-air fusion algorithm. Currently, bone-air fusion algorithms are implemented under ideal conditions, resulting in significantly worse performance in real-world applications compared to expectations. In practical applications, bone conduction and air conduction signals share common noise sources, a feature not addressed by previous algorithms, causing many to lose functionality in extreme situations. Speech attenuates as it travels through human tissue, losing high-frequency components and reducing intelligibility. Furthermore, bone conduction microphones are closely fitted to the skin, generating friction and vibration noise during wear, meaning bone-guided speech inevitably contains inherent noise. These inherent limitations pose significant challenges to practical applications. Therefore, much research focuses on improving bone-guided speech quality from several angles, such as optimizing measurement positions to compensate for high-frequency components and upgrading sensor materials. While bone-air fusion technology can improve the intelligibility of bone-guided speech, air conduction signals contain limited useful information in environments with extremely low signal-to-noise ratios, limiting their practical value.

[0004] Existing technologies and products can be broadly categorized into the following areas:

[0005] Compensation and fusion of high-frequency components of bone conduction speech;

[0006] Compensation for high-frequency components of bone conduction speech;

[0007] High-frequency compensation technology mainly includes bone-air fusion, big data, and deep trained neural networks.

[0008] Air conduction speech is acquired using traditional microphones, but the low-frequency components are easily lost in noisy environments. Bone conduction speech, on the other hand, is acquired using bone conduction sensors, but due to its inherent limitations, the high-frequency components are severely lost. Based on these two points, we propose to combine the two technologies effectively through a series of algorithms to improve speech intelligibility and enhance transmission capabilities. This is the so-called bone-air fusion.

[0009] Fusion of bone-guided and air-guided speech to improve speech quality is one of the most attractive methods currently available, primarily encompassing offline training and online real-time enhancement. Research in the former direction is based on neural network models and deep learning methods. Thang Vu Tat et al. used linear predictive analysis of speech signals, employing line spectral frequencies as suitable prediction parameters, and designed a frame-based equalization filter. They then trained a recurrent neural network to map the line spectral frequencies of bone-guided speech to those of air-guided speech. Phung et al. used Gaussian mixture models to enhance the prediction of line spectral frequencies in air-guided speech. SzuWei Fu proposed an end-to-end application architecture based on fully convolutional neural networks for speech enhancement. Most of these methods use bone-guided speech as an auxiliary or primary signal source to improve speech quality in noisy environments. Another direction is real-time online speech enhancement using adaptive noise cancellers. Most of these methods use bone-guided speech as the reference signal and air-guided speech as the desired signal. Yu et al. proposed using an FIR filter, which resulted in good recovery of low-frequency components of the speech signal, but with almost no high-frequency components. Xiao et al. began to consider Volterra filters. Compared with linear systems based on FIR filters, nonlinear systems based on Volterra filters have better convergence and computational efficiency. In general, bone-air conduction speech fusion methods based on adaptive noise cancellation have good real-time performance, meeting communication requirements. However, bone-air conduction fusion is performed under the assumption that bone-guided speech is pure and noise-free. Due to the strong low-frequency narrowband periodic noise in the environment, extensive experiments have shown that during the acquisition of bone-guided speech, the bone conduction sensor can collect narrowband periodic noise, and this noise is strongly correlated with the contaminated air conduction speech. Therefore, simple bone-air conduction fusion and the introduced nonlinear algorithms cannot remove noise when compensating for the high-frequency components of bone-guided speech, rendering the bone-air conduction fusion method ineffective in this noisy environment.

[0010] Extensive experiments revealed that low-frequency periodic noise, such as noise from cutting or forging machines in a factory, can penetrate bone conduction sensors, resulting in slight narrowband noise in bone-guided speech. Since bone and air conduction signals are correlated, this leads to undesirable noise enhancement during the adaptation process, severely impacting the quality of the output signal. Simulations show that conventional systems fail even when bone-guided speech noise energy is as low as 0.1% relative to air conduction noise. To effectively address this issue, measures must be taken to reduce or eliminate noise in bone-guided speech.

[0011] II. Smart Communication Glasses

[0012] Currently, most products on the market use bone conduction or air conduction technology alone, which introduces many interfering factors into speech signal extraction. These include the unique pronunciation of each individual, the dialects of different regions in China, and noise generated by the friction between the sensor and the human body, all of which interfere with speech recognition.

[0013] If bone conduction and air conduction technologies could be combined and applied to smart glasses, the above problems could be solved. Currently, no one has proposed this idea. Summary of the Invention

[0014] To address the shortcomings of existing technologies, this invention provides a real-time bone-air fusion communication glasses and method with narrowband noise tracking and cancellation. The glasses include a left bone conduction microphone, a right bone conduction microphone, a reference noise bone conduction microphone, and an air conduction microphone connected to a main controller. The method introduces a reference bone conduction noise channel and reconstructs the error feedback loop of the traditional system. Penetration noise is captured in real time by a bone vibration noise tracking extraction module, and the output signal is optimized by a parallel structure connected to a linear prediction filter. An error signal is then regenerated to supervise the controller update of the online bone-air fusion component, thereby reducing the impact of bone-air related noise and improving speech enhancement quality. Signal energy estimation is used to impose parameter constraints on the real-time bone fusion training system, improving the system's robustness in environments with strong narrowband noise. This system does not rely on bone-guided speech and air-guided speech databases and can be widely applied to various language environments, including dialects and minority languages.

[0015] The objective of this invention is achieved as follows:

[0016] Real-time bone-air fusion communication glasses with narrowband noise tracking cancellation include a left bone conduction microphone, a right bone conduction microphone, a reference noise bone conduction microphone, an air conduction microphone, and a main controller; the left bone conduction microphone, the right bone conduction microphone, the reference noise bone conduction microphone, and the air conduction microphone are all connected to the main controller;

[0017] The left bone conduction microphone is located on the left side of the silicone adhesive surface at the nose wing position of the glasses body, and is used to receive bone voice signals from the user's nose wing;

[0018] The right bone conduction microphone is located on the right side of the silicone adhesive surface at the nose wing position of the glasses body, and is used to receive bone voice signals from the user's nose wing;

[0019] The reference bone conduction microphone is positioned on the outer side of the temple of the glasses, in a location that does not directly contact the human skull, and is used to collect environmental noise that contaminates the bone conduction microphone.

[0020] The air conduction microphone is located at the bottom of the eyeglass frame, in a position where it is positioned above the nose tip and does not touch the body when worn, and is used to collect air conduction voice signals.

[0021] The main controller is located in the interlayer of the temple of the eyeglasses and is used for the operation and calculation of the bone air conduction speech signal processing program.

[0022] The real-time bone-air fusion method with narrowband noise tracking cancellation implemented on the aforementioned real-time bone-air fusion communication glasses includes the following steps:

[0023] Step a: Reference bone conduction ambient noise is collected using a reference bone conduction microphone. Left bone conduction microphone for nasal bone conduction speech Right bone conduction microphone for nasal bone conduction speech Air conduction microphone collects air conduction speech signals ;

[0024] Step b: Refer to the bone conduction environment noise Adaptive linear predictive filtering is performed to enhance the eigenvalues ​​of the reference narrowband noise, and the output noise estimate is obtained. Specifically:

[0025]

[0026] Where L is the length of the adaptive linear prediction filter. It is the adaptive linear predictive filter system function, and D is the delay factor of the adaptive linear predictive filter;

[0027] Step c: Estimate the noise Adaptive depth denoising is performed to reduce air-conducted speech signals. The noise in the speech is used to obtain an estimated air-guided speech. This also reduces the output signal of the fusion section. The noise in the signal is used to recover the speech signal. This allows the error signal of the bone-air fusion section to be reconstructed, and the output error signal to be generated. ;

[0028] Step d: Based on the second-order Volterra filter, convert the nasal bone-guided speech... Nasal ala bone guided speech Sum of error signals Perform online bone-air conduction speech fusion and update the output of the second-order Volterra filter. .

[0029] In the aforementioned real-time bone-air fusion method with narrowband noise tracking cancellation, step c specifically includes:

[0030] Step c1: Reduce air conduction speech signal The noise in the speech is used to obtain an estimated air-guided speech.

[0031] Estimate air conduction speech It is given by the following formula:

[0032]

[0033] in, It is the narrowband noise estimate in the interactive speech, and the calculation formula is:

[0034]

[0035] in, It is the length of the first adaptive depth noise reduction filter. It is the first adaptive depth noise reduction filtering system function;

[0036] The The update pattern is as follows:

[0037]

[0038] in, It is the positive step size of the first adaptive depth noise reduction filter.

[0039] Step c2: Reduce the output signal of the fusion section The noise in the signal is used to recover the speech signal.

[0040] Restore voice signal It is given by the following formula:

[0041]

[0042] in, It is an estimate of the narrow band component that is incorrectly recovered during osteofusion, and the calculation formula is:

[0043]

[0044] in, It is the length of the second adaptive depth noise reduction filter. It is the second adaptive depth noise reduction filtering system function;

[0045] The The update pattern is as follows:

[0046]

[0047] in, It is the positive step size of the second adaptive depth noise reduction filter;

[0048] Step c3: Output error signal The calculation formula is:

[0049] .

[0050] In the aforementioned real-time bone-air fusion method with narrowband noise tracking cancellation, step d specifically includes:

[0051] Speech guidance for nasal ala bones and nasal ala bone guided speech Taking the average, we get:

[0052]

[0053] by The output of the second-order Volterra filter serves as the real-time input to the Volterra filter. for:

[0054]

[0055] in, It is the first-order filter length of the second-order Volterra filter. It is the second-order filter length of the second-order Volterra filter. These are linear training coefficients. These are non-linear training coefficients;

[0056] The linear training coefficients and nonlinear training coefficients The update pattern is as follows:

[0057]

[0058] in, and Both are the step size of the convergence rate of the second-order Volterra filter, and the error signal Used to update the coefficients of the second-order Volterra filter.

[0059] To prevent excessive noise energy from affecting the system's robustness, a variable limiting parameter is introduced. To adjust the step size of the second-order Volterra filter and And there are:

[0060]

[0061]

[0062]

[0063] in, It is the length of the time-averaged window, and , It is the energy of the error signal. It is the energy of bone conduction signals. It is a constant;

[0064] if:

[0065] >1, and No changes;

[0066] 1, and They will be respectively and replace;

[0067] in, , .

[0068] .

[0069] Compared with the prior art, the beneficial effects of the real-time bone-air fusion communication glasses and method with narrowband noise tracking cancellation of the present invention are as follows:

[0070] First, bone vibration noise tracking and extraction uses a vibration sensor that does not contact the human body to capture penetrating noise in real time, and optimizes the output signal using a parallel structure connected to a linear prediction filter. Harmony with guided speech; and regenerate error signals. The controller updates for the online bone-air fusion component are monitored to reduce the impact of bone-air related noise and improve speech enhancement quality.

[0071] Second, by applying parameter constraints to the real-time bone fusion training system through signal energy estimation, the robustness of the system in strong narrowband noise environments is improved.

[0072] Third, since the entire method does not rely on bone conduction and air conduction speech databases, it can be widely applied to various language environments such as dialects and minority languages. Attached Figure Description

[0073] Figure 1 This is a front view of the eyeglasses structure of the present invention;

[0074] Figure 2 This is a side view of the eyeglasses structure of the present invention;

[0075] Figure 3 This is a rear view of the eyeglasses structure of the present invention;

[0076] Figure 4 This is a structural diagram of the bone-qi fusion system of the present invention;

[0077] Figure 5 This is a flowchart of the bone-qi fusion method of the present invention;

[0078] Figure 6 Bone conduction speech and air conduction speech contaminated by frequency jumps;

[0079] Figure 7 The spectrogram for recovering speech under conditions of transitioning to real noise;

[0080] Figure 8 A comparison of subjective and objective evaluations of the system and previous systems was presented when SNR=-10dB.

[0081] The components include: 1. Left bone conduction microphone; 2. Right bone conduction microphone; 3. Reference bone conduction microphone; 4. Air conduction microphone; 5. Main controller. Detailed Implementation

[0082] The specific embodiments of the present invention will now be described in further detail with reference to the accompanying drawings. Specific Implementation Method 1

[0084] The following is a specific embodiment of the real-time bone-air fusion communication glasses with narrowband noise tracking cancellation of the present invention.

[0085] The narrowband noise tracking cancellation real-time bone-air fusion communication glasses in this specific embodiment have the following structures: Figure 1 , Figure 2 and Figure 3 As shown, it includes a left bone conduction microphone 1, a right bone conduction microphone 2, a reference bone conduction microphone 3, an air conduction microphone 4, and a main controller 5; the left bone conduction microphone 1, the right bone conduction microphone 2, the reference bone conduction microphone 3, and the air conduction microphone 4 are all connected to the main controller 5;

[0086] The left bone conduction microphone 1 is located on the left side of the silicone adhesive surface at the nose wing position of the glasses body, and is used to receive bone speech signals from the user's nose wing;

[0087] The right bone conduction microphone 2 is located on the right side of the silicone adhesive surface at the nose wing position of the glasses body, and is used to receive bone voice signals from the user's nose wing;

[0088] The reference bone conduction microphone 3 is positioned on the outer side of the temple of the glasses, in a location that does not directly contact the human skull, and is used to collect environmental noise that contaminates the bone conduction microphone.

[0089] The air conduction microphone 4 is located at the bottom of the eyeglass frame, in a position where it is turned over at the tip of the nose and does not touch the human body when worn, and is used to collect air conduction voice signals.

[0090] The main controller 5 is located in the interlayer of the temple of the eyeglasses and is used for the operation and calculation of the bone air conduction speech signal processing program. Specific Implementation Method Two

[0092] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0093] The real-time bone-air fusion method with narrowband noise tracking cancellation in this specific embodiment is implemented on the real-time bone-air fusion communication glasses with narrowband noise tracking cancellation as described in Specific Embodiment 1. The method structure diagram and flowchart are as follows: Figure 4 and Figure 5 As shown, it includes the following steps:

[0094] Step a: Collect reference bone conduction ambient noise using a reference bone conduction microphone 3. Left bone conduction microphone 1 collects nasal bone conduction speech. Right bone conduction microphone 2 collects nasal bone conduction speech. Air conduction microphone 4 collects air conduction speech signals ;

[0095] Step b: Refer to the bone conduction environment noise Adaptive linear predictive filtering is performed to enhance the eigenvalues ​​of the reference narrowband noise, and the output noise estimate is obtained. Specifically:

[0096]

[0097] Where L is the length of the adaptive linear prediction filter. It is the adaptive linear predictive filter system function, and D is the delay factor of the adaptive linear predictive filter;

[0098] Step c: Estimate the noise Adaptive depth denoising is performed to reduce air-conducted speech signals. The noise in the speech is used to obtain an estimated air-guided speech. This also reduces the output signal of the fusion section. The noise in the signal is used to recover the speech signal. This allows the error signal of the bone-air fusion section to be reconstructed, and the output error signal to be generated. ;

[0099] Step d: Based on the second-order Volterra filter, convert the nasal bone-guided speech... Nasal ala bone guided speech Sum of error signals Perform online bone-air conduction speech fusion and update the output of the second-order Volterra filter. . Specific Implementation Method 3

[0101] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0102] The real-time bone-air fusion method for narrowband noise tracking cancellation in this specific embodiment, based on specific embodiment two, further specifies that step c is as follows:

[0103] Step c1: Reduce air conduction speech signal The noise in the speech is used to obtain an estimated air-guided speech.

[0104] Estimate air conduction speech It is given by the following formula:

[0105]

[0106] in, It is the narrowband noise estimate in the interactive speech, and the calculation formula is:

[0107]

[0108] in, It is the length of the first adaptive depth noise reduction filter. It is the first adaptive depth noise reduction filtering system function;

[0109] The The update pattern is as follows:

[0110]

[0111] in, It is the positive step size of the first adaptive depth noise reduction filter.

[0112] Step c2: Reduce the output signal of the fusion section The noise in the signal is used to recover the speech signal.

[0113] Restore voice signal It is given by the following formula:

[0114]

[0115] in, It is an estimate of the narrow band component that is incorrectly recovered during osteofusion, and the calculation formula is:

[0116]

[0117] in, It is the length of the second adaptive depth noise reduction filter. It is the second adaptive depth noise reduction filtering system function;

[0118] The The update pattern is as follows:

[0119]

[0120] in, It is the positive step size of the second adaptive depth noise reduction filter;

[0121] Step c3: Output error signal The calculation formula is:

[0122] . Specific Implementation Method Four

[0124] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0125] The real-time bone-air fusion method for narrowband noise tracking cancellation in this specific embodiment, based on specific embodiment two, further specifies that step d specifically includes:

[0126] Speech guidance for nasal ala bones and nasal ala bone guided speech Taking the average, we get:

[0127]

[0128] by The output of the second-order Volterra filter serves as the real-time input to the Volterra filter. for:

[0129]

[0130] in, It is the first-order filter length of the second-order Volterra filter. It is the second-order filter length of the second-order Volterra filter. These are linear training coefficients. These are non-linear training coefficients;

[0131] The linear training coefficients and nonlinear training coefficients The update pattern is as follows:

[0132]

[0133] in, and Both are the step size of the convergence rate of the second-order Volterra filter, and the error signal Used to update the coefficients of the second-order Volterra filter. Detailed Implementation Method Five

[0135] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0136] The real-time bone-air fusion method for narrowband noise tracking cancellation in this specific implementation method, based on specific implementation method four, is further limited by introducing a variable limiting parameter to prevent excessive noise energy from affecting the robustness of the system. To adjust the step size of the second-order Volterra filter and And there are:

[0137]

[0138]

[0139]

[0140] in, It is the length of the time-averaged window, and , It is the energy of the error signal. It is the energy of bone conduction signals. It is a constant;

[0141] if:

[0142] >1, and No changes;

[0143] 1, and They will be respectively and replace;

[0144] in, , . Specific Implementation Method Six

[0146] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0147] The real-time bone-air fusion method with narrowband noise tracking cancellation in this specific implementation method is further limited based on specific implementation method five: . Detailed Implementation Method Seven

[0149] The following is a specific implementation of the real-time bone-air fusion method of the present invention, which includes narrowband noise tracking cancellation.

[0150] To verify the effectiveness of the real-time skeletal speech fusion method with narrowband noise tracking cancellation of the present invention in noisy environments, simulation experiments were conducted. The simulations simulated variations in narrowband noise in mechanical equipment under different operating conditions: from synthetic narrowband noise with two frequencies (300 Hz, 500 Hz) to synthetic narrowband noise with five frequencies (300 Hz, 500 Hz, 800 Hz, 1200 Hz, and 1600 Hz), with the synthetic noise added to the recorded speech signal at signal-to-noise ratios (SNR) of -10, -5, 0, 5, and 10 dB, respectively. 1% of the synthetic narrowband noise was added to the pure BC speech. This noise was also used as the BC reference noise. Random additive noise was added to AC speech with an SNR of -1 dB. Several objective evaluation metrics were used to assess the quality of speech enhancement at SNR = -10 dB, including Perceptual Speech Quality (PESQ), Short-Term Objective Intelligibility (STOI), and Extended Short-Term Objective Intelligibility (ESTOI). The PESQ score ranges from -0.5 to 4.5, while the STOI and ESTAI scores range from 0 to 1, with higher scores indicating better speech intelligibility after recovery. We also used the Mean Opinion Score (MOS), represented as a value between 1 and 5. To ensure the accuracy of the MOS, 50 males and 50 females were selected to rate the recovered speech. Since other methods struggle to handle real-world noisy environments, only spectrograms of the recovered speech from Volterra and the system proposed in this invention are shown. Figure 6 The spectrograms of contaminated air-guided speech and bone-guided speech are shown, where (a) is the spectrogram of contaminated air-guided speech; and (b) is the spectrogram of contaminated bone-guided speech. Figure 7 The spectrograms of the recovered speech are shown at an SNR of -10 dB. (a) is the spectrogram of the Volterra-enhanced speech; (b) is the spectrogram of the speech enhancement from the proposed system. Subjective and objective evaluations of all systems are shown below. Figure 8 As shown.

Claims

1. A real-time bone-air fusion method with narrowband noise tracking cancellation implemented on real-time bone-air fusion communication glasses with narrowband noise tracking cancellation, wherein the real-time bone-air fusion communication glasses with narrowband noise tracking cancellation include a left bone conduction microphone (1), a right bone conduction microphone (2), a reference noise microphone (3), an air conduction microphone (4), and a main controller (5); the left bone conduction microphone (1), the right bone conduction microphone (2), the reference noise microphone (3), and the air conduction microphone (4) are all connected to the main controller (5); The left bone conduction microphone (1) is located on the left side of the silicone adhesive surface at the nose wing position of the eyeglasses, and is used to receive bone speech signals from the user's nose wing. The right bone conduction microphone (2) is located on the right side of the silicone adhesive surface at the nose wing position of the eyeglasses, and is used to receive bone speech signals from the user's nose wing. The reference bone conduction microphone (3) is set on the outer side of the temple of the eyeglasses, in a position that does not directly contact the human skull, and is used to collect environmental noise from contaminated bone conduction microphones. The air conduction microphone (4) is located at the bottom of the eyeglass frame, in a position where it is turned over at the tip of the nose and does not touch the human body when worn, and is used to collect air conduction speech signals. The main controller (5) is located in the interlayer of the temple of the eyeglasses and is used for the operation and calculation of the bone air conduction speech signal processing program; Its features are, The real-time bone-air fusion method with narrowband noise tracking cancellation includes the following steps: Step a, Reference bone conduction environmental noise (3) Left bone conduction microphone (1) for collecting nasal bone conduction speech Right bone conduction microphone (2) for nasal bone conduction speech acquisition The air conduction microphone (4) collects air conduction speech signals. ; Step b: Refer to the bone conduction environment noise Adaptive linear predictive filtering is performed to enhance the eigenvalues ​​of the reference narrowband noise, and the output noise estimate is obtained. Specifically: Where L is the length of the adaptive linear prediction filter. It is the adaptive linear predictive filter system function, and D is the delay factor of the adaptive linear predictive filter; Step c: Estimate the noise Adaptive depth denoising is performed to reduce air-conducted speech signals. The noise in the speech is used to obtain an estimated air-guided speech. This also reduces the output signal of the fusion section. The noise in the signal is used to recover the speech signal. This allows the error signal of the bone-air fusion section to be reconstructed, and the output error signal to be generated. Step c specifically involves: Step c1: Reduce air conduction speech signal The noise in the speech is used to obtain an estimated air-guided speech. Estimate air conduction speech It is given by the following formula: in, It is the narrowband noise estimate in the interactive speech, and the calculation formula is: in, It is the length of the first adaptive depth noise reduction filter. It is the first adaptive depth noise reduction filtering system function; The The update pattern is as follows: in, It is the positive step size of the first adaptive depth noise reduction filter; Step c2: Reduce the output signal of the fusion section The noise in the signal is used to recover the speech signal. Restore voice signal It is given by the following formula: in, It is an estimate of the narrow band component that is incorrectly recovered during osteofusion, and the calculation formula is: in, It is the length of the second adaptive depth noise reduction filter. It is the second adaptive depth noise reduction filtering system function; The The update pattern is as follows: in, It is the positive step size of the second adaptive depth noise reduction filter; Step c3: Output error signal The calculation formula is: Step d: Based on the second-order Volterra filter, convert the nasal bone-guided speech... Nasal ala bone guided speech Sum of error signals Perform online bone-air conduction speech fusion and update the output of the second-order Volterra filter. .

2. The real-time bone-air fusion method with narrowband noise tracking cancellation according to claim 1, characterized in that, Step d specifically involves: Speech guidance for nasal ala bones and nasal ala bone guided speech Taking the average, we get: by The output of the second-order Volterra filter serves as the real-time input to the Volterra filter. for: in, It is the first-order filter length of the second-order Volterra filter. It is the second-order filter length of the second-order Volterra filter. These are linear training coefficients. These are non-linear training coefficients; The linear training coefficients and nonlinear training coefficients The update pattern is as follows: in, and Both are the step size of the convergence rate of the second-order Volterra filter, and the error signal Used to update the coefficients of the second-order Volterra filter.

3. The real-time bone-air fusion method with narrowband noise tracking cancellation according to claim 2, characterized in that, To prevent excessive noise energy from affecting the system's robustness, a variable limiting parameter is introduced. To adjust the step size of the second-order Volterra filter and And there are: in, It is the length of the time-averaged window, and , It is the energy of the error signal. It is the energy of bone conduction signals. It is a constant; if: >1, and No changes; 1, and They will be respectively and replace; in, , .

4. The real-time bone-air fusion method with narrowband noise tracking cancellation according to claim 3, characterized in that, 。

Citation Information

Patent Citations

  • Real-time bone and gas fusion voice communication mask in extreme environment and bone and gas fusion method

    CN116019275A

  • Active noise cancellation for bone conduction speaker of a head-mounted wearable device

    US10699691B1