Real-time Bone-Conduction and Air-Conduction Integrated Voice Communication Mask and Bone-Conduction and Air-Conduction Integration Method in Extreme Environments

By using notch-based frequency-chasing preprocessing active noise reduction technology and adaptive (Fe-Vb-LMS) update algorithm in extreme environments, the volterra model is trained in real time, solving the problem of bone-directed voice signal pollution in strong noise environments, achieving efficient backbone fusion voice communication, and improving voice communication quality and system stability.

CN116019275BActive Publication Date: 2025-06-27NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310039479.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-11
Publication Date
2025-06-27
Estimated Expiration
2043-01-11

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time voice communication in extreme environments, especially in strong noise environments. The voice signal of the bone-directed voice is easily polluted by environmental noise, resulting in the failure of the technology of bone-oriented fusion.

Method used

Adopting the frequency-chasing preprocessing active noise reduction technology based on notch, combined with online adaptive noise reduction and offline pre-training algorithm, an adaptive (Fe-Vb-LMS) update algorithm with filter error variable boundaries is proposed to train the volterra model in real time to achieve robust online intelligibility enhancement of bone-directed voice, and adaptively reduce narrowband vibration noise components.

Benefits of technology

In extreme environments, the quality of voice communication is significantly improved, the intelligibility and signal-to-noise ratio of bone-directed voice is improved, and the noise pollution is effectively reduced, ensuring the stability and real-time nature of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116019275B_ABST
    Figure CN116019275B_ABST
Patent Text Reader

Abstract

The real-time bone-air integrated voice communication mask and the bone-air integration method in extreme environments of the present invention belong to the technical field of bone-conduction voice communication; the mask includes: a mask body and silicone mask straps arranged on both sides of the mask body, a bone conduction microphone is arranged at the nasal wing position of the mask body, and the bone conduction microphone is used to receive bone voice signals from the nasal wings of the user. A reference noise bone microphone is arranged at a position on the outer side of the mask body that does not directly adhere to the human skull, and the reference noise bone microphone is used to collect environmental noise that may contaminate the bone conduction microphone. An air conduction microphone is arranged at a position below the lips and not in contact with the human body when worn inside the mask body. The notch filter-based frequency tracking preprocessing active noise reduction technology is introduced, and an adaptive update algorithm with variable boundaries of filtering error is proposed to train the volterra model in real time to achieve robust online intelligibility enhancement of bone-conducted speech, and at the same time adaptively reduce the narrowband vibration noise component homologous to air-conducted speech in bone-conducted speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The real-time bone-air integrated voice communication mask and the bone-air integration method under extreme environments of the present invention belong to the technical field of bone-conducted voice communication. Background Art

[0002] In the industrial and agricultural production and processing environments, it is inevitable to have pollutions such as dust and strong noise, so masks need to be worn. However, while masks isolate pollutions such as viruses and dust, they also isolate the voices of people speaking. Especially for protective masks, which are filled with a certain thickness of adsorption materials, it further affects the transmission of voices. Currently, existing products have air-conducted microphones built into the masks to achieve voice communication in special environments. However, the air-conducted microphones are too close to the lips, inevitably introducing friction and airflow noises. At the same time, in application scenarios with strong background noises, the air-conducted voices are also easily polluted by the environmental noises.

[0003] Currently, all bone-air integration algorithms consider that the bone-conducted signals are relatively clean, and the bone-conducted and air-conducted signals are algorithmically integrated under ideal conditions, lacking consideration for algorithm real-time performance and algorithm robustness. Therefore, it is difficult to be practically applied. In practical applications, the bone-conducted signals are often slightly polluted by environmental vibration noises homologous to the air-conducted signals in noisy factories, which causes the bone-air integration technology trained with the bone-conducted signals as the reference signals to fail. Due to problems such as the lack of oral radiation characteristics and the loss of high-frequency components caused by in-vivo conduction of bone-conducted voices, their intelligibility is significantly lower than that of air-conducted voices. The bone-air integration technology can improve the intelligibility of bone-conducted voices. However, in an environment with an extremely low signal-to-noise ratio, the useful information in the air-conducted signals is also less, and the utilization value is not high.

[0004] The existing technologies and products are roughly divided into the following aspects:

[0005] Compensation and integration of high-frequency components of bone-conducted voices;

[0006] Compensation of high-frequency components of bone-conducted voices;

[0007] The high-frequency compensation technology mainly includes bone-air integration, big data, deep training neural networks, etc.

[0008] Air-conducted voices are collected for voice signals through traditional voice microphones. The low-frequency part of air-conducted voices is easily interfered and submerged by strong noise environments; bone-conducted voices are collected for voice signals through bone-conducted sensors, and due to the constraints of their own properties, the high-frequency part is severely lost. Starting from these two points, considering combining the two through a series of algorithms to improve the intelligibility of voices and enhance the transmission ability, this is the so-called bone-air integration.

[0009] Air-conducted speech is very vulnerable to various noises during communication. The introduced bone-conducted speech signal collects the vibration of the skull through a sensor and then converts it into an audio signal. However, for the speech signal collected through the human body, the high-frequency attenuation is very serious. Therefore, the method of bone-air fusion is considered. However, bone-air fusion is carried out under the assumption that the bone-conducted speech is pure and noise-free. Due to the very strong low-frequency narrowband periodic noise in the environment, through a large number of experiments, it can be known that during the collection of bone-conducted speech, the bone-conducted sensor can collect narrowband periodic noise, and this noise is strongly correlated with the contaminated air-conducted speech. Therefore, the simple bone-air fusion and the introduced non-linear algorithm cannot remove the noise when compensating for the high-frequency components of bone-conducted speech, and the bone-air fusion method fails under this noise background:

[0010] The team of Shiming Zhang studied the characteristics of bone-conducted speech and fused the bone-conducted speech by the least squares method; quantified the noise robustness of bone-conducted speech in different noise environments, and obtained that bone-conducted speech provides an SNR gain of about 10 dB compared with air-conducted speech.

[0011] Hung-Ping Liu et al. studied bone-conducted microphones and air-conducted microphones, and found that the bone-conducted microphone collects the skull vibration of the speaker and transmits the speech signal, which has better anti-noise ability than the air-conducted microphone. A method of deep denoising autoencoder was proposed to combine bone-conducted speech and air-conducted speech to improve the intelligibility and clarity of speech.

[0012] Cheng Yu proposed a new time-domain multi-modal speech enhancement structure, which utilizes bone conduction and air conduction signals. And two ensemble learning-based strategies, early fusion (EF) and late fusion (LF), were studied to integrate two types of speech signals, and a deep learning-based fully convolutional network was used for enhancement.

[0013] The above studies have improved the intelligibility and clarity of speech to varying degrees by using filters with different structures and strategies such as deep learning, but they are limited to the laboratory level and cannot effectively solve the problem of low-frequency narrowband noise introduced in bone-conducted sensors for real application scenarios.

[0014] II. Intelligent communication mask

[0015] High-efficiency filtering masks often affect speech transmission. Therefore, there is a technology of installing microphones inside existing masks to collect speech signals, but these air-conducted microphones will be contaminated by environmental noise.

[0016] If the bone-conduction and air-conduction technologies can be fused and applied to masks, the above problems can be solved. However, this technology is currently in a blank stage. Summary of the Invention

[0017] Aiming at the deficiencies of the prior art, the present invention provides a real-time bone-air fusion voice communication mask and a bone-air fusion method under extreme environments. By combining bone conduction and air conduction microphones and using a combination of online adaptive noise reduction and offline pre-training algorithms, high-efficiency bone-air fusion is achieved, and the voice communication quality under extreme environments is improved. The present invention introduces a frequency tracking preprocessing active noise reduction technology based on notch filters and proposes an adaptive (Fe-Vb-LMS) update algorithm with variable filter error boundaries to train the Volterra model in real time to achieve robust online intelligibility enhancement of bone-conducted speech. At the same time, it adaptively reduces the narrowband vibration noise components homologous to air-conducted speech in bone-conducted speech. The air-conducted signal in the mask is the best measurement position for air-conducted signals in a strong noise environment, and the inner side of the nose clip of the mask is also the best acquisition position for bone-conducted speech. The fusion of the two signals can structurally improve the sound source quality, make up for each other's deficiencies, and greatly exert the respective advantages of bone-conducted and air-conducted speech, improving the voice communication quality in various complex environments.

[0018] To achieve the above object, the present invention provides the following technical solutions:

[0019] A real-time bone-air fusion voice communication mask under extreme environments, comprising: a mask body and silicone mask straps arranged on both sides of the mask body. A silicone fitting surface is provided on the side of the mask body close to the face. A left bone conduction microphone and a right bone conduction microphone are provided on the silicone fitting surface at the nose wing position of the mask body. The left bone conduction microphone and the right bone conduction microphone are used to receive bone voice signals from the user's nose wings. A reference noise bone microphone is provided at a position on the outer side of the mask body that does not directly fit the human skull, and the reference noise bone microphone is used to collect the environmental noise that contaminates the bone conduction microphone. An air conduction microphone is provided at a position inside the mask body that is below the lips and does not contact the human body when worn, and the air conduction microphone is used to collect air-conducted voice signals. Left and right vibrators are respectively provided at the intermediate positions where the two silicone mask straps can fit with the posterior ear bones when worn.

[0020] For the above real-time bone-air fusion voice communication mask under extreme environments, a filtering structure capable of circulating air in both directions is provided on the mask body; the filtering structure includes: a filter upper cover, a filter screen, and a filter lower cover, and the front and rear sides of the filter screen are respectively fastened by the filter upper cover and the filter lower cover.

[0021] For the above real-time bone-air fusion voice communication mask under extreme environments, a main controller for running and calculating the bone-air conduction voice signal processing program is provided in the mask body sandwich.

[0022] The left bone conduction microphone, the right bone conduction microphone, the reference noise bone microphone, the air conduction microphone, the left vibrator, and the right vibrator are all connected to the main controller.

[0023] Bone-air fusion method, applied to a bone-air fusion system, the bone-air fusion system comprising: a real-time intelligibility enhancement module, a reference frequency tracking module, an adaptive depth noise reduction module, and an adaptive error correction module, the bone-air fusion method comprising the following steps:

[0024] Step a, a reference bone-conduction microphone outside the mask collects reference bone-conduction environmental noise x N (n); the left and right bone-conduction microphones inside the mask collect nasal bone-conduction speech x L (n), x R (n); an air-conduction microphone inside the mask collects an air-conduction speech signal d(n);

[0025] Step b, the output reference bone-conduction environmental noise x N (n) passes through the reference frequency tracking module to obtain in real time the frequency components of the large-energy noise of the low-frequency period in the environment, and outputs f i (n);

[0026] Step c, the frequency component f i (n) passes through the adaptive error correction module, with orthogonal reference signals x Nai and x Nbi as the reference input, and in real time tracks and suppresses the low-frequency narrowband periodic components in the air-conduction speech d(n) in step a, and outputs e N (n);

[0027] Step d, the frequency component f i (n) passes through the adaptive depth noise reduction module, with orthogonal reference signals x Nai and x Nbi as the reference input, and in real time tracks and suppresses the low-frequency narrowband frequency information after speech fusion, and outputs as the speech signal with the optimal intelligibility and signal-to-noise ratio;

[0028] Step e, the nasal bone-conduction speech x L (n), x R (n) in step a and the output of step e The output e N (n) of step c pass through the real-time intelligibility enhancement module, and based on the second-order Volterra model, use the Fe-Vb-LMS algorithm to realize online compensation for the high-frequency components of robust bone-conduction speech.

[0029] For the above bone-air fusion method, step c is specifically: the frequency component f i (n) of step two passes through the adaptive error correction module, with the orthogonalized reference signals x Nai and x Nbi as the reference input,

[0030]

[0031] Taking the air-conducted voice signal d(n) as the target signal, the environmental periodic component estimation of the adaptive error correction module is:

[0032]

[0033] Among them, the i-th group of adaptive coefficients a Ni (n), b Ni (n) are updated as:

[0034]

[0035] Among them, μ N is the step size of coefficient update, generally set to be between, e N (n) is the output of the adaptive error correction module.

[0036] For the above bone-air fusion method, step e is specifically:

[0037] Taking the nasal bone-conducted voices x L (n), x R (n) in step a and taking the average value to get:

[0038]

[0039] Taking as the reference input of the real-time intelligibility enhancement module, the intelligibility-enhanced bone-conducted voice signal y0(n) output by the second-order Volterra model is:

[0040]

[0041] Among them, the linear training coefficient h 1,i (n + 1) and the non-linear training coefficient h 2,i,j (n + 1) are updated as

[0042]

[0043] The error e(n) of its adaptive update is the error output e N (n) of the adaptive error correction module and the output error of the adaptive depth noise reduction module

[0044]

[0045] and e N (n) are expressed as:

[0046]

[0047] Noise prediction of the adaptive depth noise reduction module is as follows:

[0048]

[0049] Its coefficient update formula is:

[0050]

[0051] To ensure the stability of the real-time intelligibility enhancement module, boundary constraint judgment is introduced. When the update error energy of the corrected prediction is greater than the bone-conducted speech reference input energy , that is:

[0052]

[0053] a variable boundary r is introduced b

[0054]

[0055] Step size update

[0056] μ(n + 1) = r b μ(n)

[0057] Otherwise

[0058] μ(n + 1) = μ(n).

[0059] Beneficial effects:

[0060] First, for the real-time bone-air integrated voice communication mask in extreme environments of the present invention, a frequency tracking preprocessing active noise reduction technology based on notch filters is introduced, and an adaptive (Fe-Vb-LMS) update algorithm with variable boundaries of filtering error is proposed to realize robust online intelligibility enhancement of bone-conducted speech by real-time training of the Volterra model, and at the same time adaptively reduce the narrowband vibration noise components homologous to air-conducted speech in bone-conducted speech.

[0061] Second, for the real-time bone-air integrated voice communication mask in extreme environments of the present invention, it can perform adaptive notch filter frequency estimation on the bone-conducted signals collected outside the mask and generate a noise reference signal, so as to realize real-time preprocessing and fusion processing of online bone-conducted and air-conducted speech.

[0062] Third, for the real-time bone-air integrated voice communication mask in extreme environments of the present invention, the boundary of the adaptive fusion system error update matrix is adjusted by using variable constraints to ensure the stability of the system in a strong noise environment.

[0063] Fourth, the present invention combines bone conduction and air conduction technologies and applies them to a mask. Based on the structure of a filtering mask, a bone-air integration communication system is designed. By using the bone conduction voice signal source in the nasal bone and the air conduction voice signal source inside the mask, cleaner voice signals are obtained, and the environmental reference signal source is collected by a bone microphone on the outside of the mask that does not come into contact with the human body.

[0064] Fifth, the present invention proposes a bone-air integration method, specifically a Fe-Vb-LMS algorithm, which realizes the online compensation of the high-frequency components of bone-conducted speech based on a second-order Volterra model. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the front view of the mask structure of the present invention;

[0066] Figure 2 is the side view of the mask structure of the present invention;

[0067] Figure 3 is the top view of the mask structure of the present invention;

[0068] Figure 4 is the disassembled view of the mask structure of the present invention;

[0069] Figure 5 is the schematic diagram of the system operation of the mask of the present invention;

[0070] Figure 6 is the structure diagram of the bone-air integration system of the present invention;

[0071] Figure 7 is the flow chart of the bone-air integration algorithm of the present invention;

[0072] Figure 8 is the spectrogram of bone-conducted speech in a strong noise environment;

[0073] Figure 9 is the spectrogram of the comparative literature with a signal-to-noise ratio of -2;

[0074] Figure 10 is the effect diagram of the present invention with a signal-to-noise ratio of -2;

[0075] Figure 11 is the effect diagram of the present invention with a signal-to-noise ratio of -10;

[0076] Figure 12 is the introduction of the stability boundary r b after that, the effect diagram of the present invention with a signal-to-noise ratio of -10;

[0077] Figure 13 is the introduction of the stability boundary r b after that, the effect diagram of the present invention with a signal-to-noise ratio of -30.

[0078] Wherein: 1. Left bone conduction microphone; 2. Right bone conduction microphone; 3. Reference noise bone microphone; 4. Air conduction microphone; 5. Left oscillator; 6. Right oscillator; 7. Main controller; 8. Upper filter cover; 9. Filter net; 10. Lower filter cover; 11. Mask body; 12. Soft silicone fitting surface; 13. Silicone mask strap; 100. Real-time intelligibility enhancement module; 200. Reference frequency tracking module; 300. Adaptive depth noise reduction module; 400. Adaptive error correction module. Specific embodiments

[0079] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Specific embodiment 1

[0081] The following are the specific embodiments of the real-time bone-air fusion voice communication mask of the present invention in extreme environments.

[0082] The real-time bone-air fusion voice communication mask in extreme environments under this specific embodiment, as Figures 1-4 shown, includes: a mask body 11 and silicone mask straps 13 arranged on both sides of the mask body 11. A silicone fitting surface 12 is provided on the side of the mask body 11 close to the face. A left bone conduction microphone 1 and a right bone conduction microphone 2 are provided on the silicone fitting surface 12 at the nose wing position of the mask body 11. The left bone conduction microphone 1 and the right bone conduction microphone 2 are used to receive bone voice signals from the user's nose wing. A reference noise bone microphone 3 is provided at a position on the outer side of the mask body 11 that does not directly contact the human skull. The reference noise bone microphone 3 is used to collect the environmental noise that contaminates the bone conduction microphone. An air conduction microphone 4 is provided at a position on the inner side of the mask body 11 that is below the lips and does not contact the human body when worn. The air conduction microphone 4 is used to collect air conduction voice signals. Left and right oscillators 5 and 6 are respectively provided at the middle positions where the two silicone mask straps 13 can be attached to the posterior ear bone when worn. Specific embodiment 2

[0084] The following are the specific embodiments of the real-time bone-air fusion voice communication mask of the present invention in extreme environments.

[0085] The real-time bone-air fusion voice communication mask in extreme environments under this specific embodiment is further defined on the basis of specific embodiment 1: A filtering structure capable of circulating air in both directions is provided on the mask body 11; The filtering structure includes: an upper filter cover 8, a filter net 9, and a lower filter cover 10. The front and rear sides of the filter net 9 are respectively buckled by the upper filter cover 8 and the lower filter cover 10. Specific embodiment 3

[0087] The following are the specific embodiments of the real-time bone-air fusion voice communication mask of the present invention in extreme environments.

[0088] The real-time bone-air integrated voice communication mask in extreme environments under this specific implementation further defines, on the basis of the first specific implementation: a main controller 7 for running and calculating the bone-conducted voice signal processing program is provided in the sandwich layer of the mask body 11.

[0089] The left bone conduction microphone 1, the right bone conduction microphone 2, the reference noise bone conduction microphone 3, the air conduction microphone 4, the left vibrator 5 and the right vibrator 6 are all connected to the main controller 7, as Figure 5 shown. Specific implementation four

[0091] The following are the specific implementations of the bone-air integration method of the present invention.

[0092] The bone-air integration method under this specific implementation is applied to a bone-air integration system. The bone-air integration system is as Figure 6 shown, and includes: a real-time intelligibility enhancement module 100, a reference frequency tracking module 200, an adaptive depth noise reduction module 300, and an adaptive error correction module 400. The flow chart of the bone-air integration method is as Figure 7 shown, and includes the following steps:

[0093] Step a: The reference noise bone conduction microphone 3 outside the mask collects the reference bone conduction environmental noise x N (n); the left bone conduction microphone 1 and the right bone conduction microphone 2 inside the mask collect the nasal bone conduction voice x L (n), x R (n); the air conduction microphone 4 inside the mask collects the air conduction voice signal d(n);

[0094] Step b: The output reference bone conduction environmental noise x N (n) of step a passes through the reference frequency tracking module 200 to obtain the frequency components of the large-energy noise with low-frequency cycles in real time in the environment, and outputs f i (n);

[0095] Step c: The frequency component f i (n) of step b passes through the adaptive error correction module 400, and uses the orthogonal reference signals x Nai and x Nbi as the reference input to track and suppress the low-frequency narrow-band periodic components in the air conduction voice d(n) of step a in real time, and outputs e N (n);

[0096] Step d: The frequency component f i (n) of step b passes through the adaptive depth noise reduction module 300, and uses the orthogonal reference signals x Nai and x Nbi as the reference input to track and suppress the low-frequency narrow-band frequency information after voice integration in real time, and outputs The voice signal with optimal final intelligibility and signal-to-noise ratio;

[0097] Step e: The nasal bone-conducted voice x in step a L (n), x R (n) and the output of step e The output e of step c N (n), through the real-time intelligibility enhancement module 100, based on the second-order Volterra model, using the Fe-Vb-LMS algorithm, realizes the online compensation of the high-frequency components of the robust bone-conducted voice. Specific Embodiment 5

[0099] The following is the specific implementation of the bone-air fusion method of the present invention.

[0100] Under this specific implementation, the bone-air fusion method, on the basis of Specific Embodiment 4, is further limited: Step c is specifically: The frequency component f in step two i (n) passes through the adaptive error correction module 400, with the orthogonalized reference signals x Nai and x Nbi as the reference inputs,

[0101]

[0102] Taking the air-conducted voice signal d(n) as the target signal, the environmental periodic component estimation of the adaptive error correction module 400 is:

[0103]

[0104] Among them, the i-th group of adaptive coefficients a of the adaptive error correction module 400 Ni (n), b Ni (n) are updated as:

[0105]

[0106] Among them, μ N is the step size of coefficient update, generally set between 0 and 1, and e N (n) is the output of the adaptive error correction module 400. Specific Embodiment 6

[0108] The following is the specific implementation of the bone-air fusion method of the present invention.

[0109] Under this specific implementation, the bone-air fusion method, on the basis of Specific Embodiment 4, is further limited: Step e is specifically:

[0110] Taking the average value of the nasal bone-conducted voices x L (n), x R (n) in step a gives:

[0111]

[0112] Taking as the reference input of the real-time intelligibility enhancement module 100, the bone-conducted speech signal y0(n) with enhanced intelligibility output through the second-order Volterra model is:

[0113]

[0114] where the linear training coefficient h 1,i (n + 1) and the non-linear training coefficient h 2,i,j (n + 1) are updated to

[0115]

[0116] The error e(n) of its adaptive update is the error output e N (n) of the adaptive error correction module 400 minus the output error of the adaptive deep noise reduction module 300

[0117]

[0118] and the expression of e N (n) is:

[0119]

[0120] The noise prediction of the adaptive deep noise reduction module 300 is:

[0121]

[0122] Its coefficient update formula is:

[0123]

[0124] To ensure the stability of the real-time intelligibility enhancement module 100, boundary constraint judgment is introduced. When the update error energy of the correction prediction is larger than the bone-conducted speech reference input energy , that is:

[0125]

[0126] a variable boundary r b

[0127]

[0128] step size update

[0129] μ(n + 1) = r b μ(n)

[0130] Otherwise

[0131] μ(n + 1) = μ(n).

[0132] The present invention introduces a coefficient update rule with variable boundary definition. If the estimated error energy is less than the bone-conducted speech reference input energy, the original stable step size is maintained and the coefficient is updated normally; if the error energy is greater than the bone-conducted speech input, it indicates that the signal-to-noise ratio is extremely low and the environmental noise is much greater than the input signal. There are two possibilities. One is that there is a gap in the speech, that is, there is only noise without speech. In this case, the boundary r b is a very small value, and the formula μ(n + 1) = r b μ(n) outputs a very small value approaching 0, making the volterra coefficient almost remain unchanged, eliminating the need for the traditional additional endpoint detection step for speech, and making the adaptive process more coherent; the second possibility is that when the speech exists normally, r b is approximate and less than 1. The greater the error, the smaller the r b step size, and the coefficient update can be adaptively adjusted to ensure the system performance. By introducing the boundary r b the system can be made more stable, capable of coping with extreme environments with a signal-to-noise ratio above -30 dB; moreover, it can automatically judge the speech interval and adjust automatically, enhancing the system's noise suppression effect and improving the signal-to-noise ratio of the output signal.

[0133] Combined Figures 8-13 As shown, for the rotating machinery noises of 500 Hz and 1000 Hz, when the signal-to-noise ratio is -2 dB ( -1 for broadband and narrowband noises respectively), if the energy of the periodic signal of 500 Hz sensed by the bone-conducted microphone is only 1% of the reference noise bone microphone 3, and the energy of 1000 Hz is only 0.5% of the environmental air-conducted microphone, as Figure 8 shown, the final trained speech effect is as Figure 9 shown. Since there is a very small periodic contamination in the bone-conducted signal, which is correlated with the periodic signal in the air-conducted signal, after training, the periodic signal is easily enhanced and amplified. After the algorithm of the present invention, the periodic signal is effectively suppressed as Figure 10 . If the signal-to-noise ratio is -10 dB in an extreme environment, the system cannot converge, but after the algorithm of the present invention, the system still converges normally with excellent stability, as Figure 11 . After introducing the stability boundary r b the system performance is further improved, and the effect is as Figure 12 shown as the output spectrogram of the improved system at a signal-to-noise ratio of -10; Figure 13 is the output spectrogram of the improved system at a signal-to-noise ratio of -30. It can be seen that the improved system is less affected by the signal-to-noise ratio and has high stability.

[0134] The present invention is summarized as follows:

[0135] The input signals of the present invention include: left bone conduction microphone 1, right bone conduction microphone 2, reference noise bone microphone 3, and air conduction microphone 4; the output signals include: left oscillator 5 and right oscillator 6. The present invention adopts a silicone mask structure to design an intelligent bone-air conduction voice communication system, introduces a frequency tracking preprocessing active noise reduction technology with an adaptive notch filter band ANFB, and proposes an adaptive (Fe-Vb-LMS) update algorithm with variable filter error boundaries to train the volterra model in real time to achieve robust online intelligibility enhancement of bone conduction voice. At the same time, through the real-time intelligibility enhancement module 100, reference frequency tracking module 200, adaptive depth noise reduction module 300, and adaptive error correction module 400, the narrowband vibration noise components homologous to air conduction voice in bone conduction voice are adaptively reduced in an environment with dust pollution, extremely strong noise, and mechanical processing equipment periodic vibration noise as the background.

[0136] The present invention uses the adaptive notch filter band ANFB to adaptively extract the noise frequency reference signal from the bone conduction microphones on the outer side of the mask, and synchronously tracks the air conduction and trained bone conduction signals, which can effectively cope with the influence of periodic mechanical processing vibration noise on the bone conduction voice microphone in a closed strong noise environment.

[0137] The present invention proposes an adaptive (Fe-Vb-LMS) update algorithm with variable filter error boundaries, and uses the error after real-time update and correction to update the volterra coefficients to improve the signal-to-noise ratio of the bone conduction voice online intelligibility enhancement system.

[0138] The present invention uses nasal bone conduction voice and air conduction voice inside the mask, which can complement each other to effectively compensate for the lost high-frequency components in bone conduction voice, effectively ensuring the real-time efficiency of the system while improving the intelligibility of the system.

[0139] The present invention introduces a variable boundary criterion in the adaptive volterra link to adjust the feedback error, ensuring the stability of the system in a low signal-to-noise ratio scenario in an extreme environment.

[0140] The present invention uses the vibration of the posterior auricular bone oscillator after the silicone mask strap is attached to achieve the perception of bone conduction voice. It is convenient to wear and wire, the silicone ear hook is comfortable, and the ears are liberated without wearing burden.

Claims

1. A real-time bone-air integrated voice communication mask for extreme environments, comprising: A mask body (11) and silicone mask straps (13) provided on both sides of the mask body (11), characterized in that a silicone fitting surface (12) is provided on the side of the mask body (11) close to the face, and a left bone conduction microphone (1) and a right bone conduction microphone (2) are provided on the silicone fitting surface (12) at the nose wing position of the mask body (11). The left bone conduction microphone (1) and the right bone conduction microphone (2) are used to receive bone voice signals from the nose wings of the user. A reference noise bone microphone (3) is provided at a position on the outer side of the mask body (11) that does not directly fit the human skull. The reference noise bone microphone (3) is used to collect the environmental noise that contaminates the bone conduction microphone. An air conduction microphone (4) is provided at a position below the lips and not in contact with the human body when worn on the inner side of the mask body (11). The air conduction microphone (4) is used to collect air conduction voice signals. Left oscillators (5) and right oscillators (6) are respectively provided at intermediate positions where the two silicone mask straps (13) can be attached to the posterior ear bones when worn.

2. The real-time bone-air integrated voice communication mask in an extreme environment according to claim 1, wherein A filtering structure capable of circulating air in both directions is provided on the mask body (11); the filtering structure includes: a filter upper cover (8), a filter net (9), and a filter lower cover (10), and the front and rear sides of the filter net (9) are respectively fastened by the filter upper cover (8) and the filter lower cover (10).

3. The real-time bone-air fusion voice communication mask in an extreme environment according to claim 1 or 2, characterized in that, A main controller (7) for running and calculating a bone-air conduction voice signal processing program is provided in the sandwich layer of the mask body (11).

4. The real-time bone-air integrated voice communication mask for extreme environments according to claim 3, characterized in that, The left bone conduction microphone (1), the right bone conduction microphone (2), the reference noise bone microphone (3), the air conduction microphone (4), the left oscillator (5), and the right oscillator (6) are all connected to the main controller (7).

5. A bone-air integration method for a real-time bone-air integrated voice communication mask in an extreme environment according to any one of claims 1-4, characterized in that, Applied to a bone-air fusion system, the bone-air fusion system includes: a real-time intelligibility enhancement module (100), a reference frequency tracking module (200), an adaptive depth noise reduction module (300), and an adaptive error correction module (400). The bone-air fusion method includes the following steps: Step a: The reference bone-conduction microphone (3) on the outer side of the mask collects the reference bone-conduction environmental noise x N (n); the left bone-conduction microphone (1) and the right bone-conduction microphone (2) on the inner side of the mask collect the nasal bone-conduction speech x L (n), x R (n); the air-conduction microphone (4) on the inner side of the mask collects the air-conduction speech signal d(n); Step b: Use the output reference bone-conducted environmental noise x N (n) to obtain in real time, through a reference frequency tracking module (200), the frequency components of the large-energy noise with a low-frequency period in the environment, and output f i (n); Step c: The frequency component f i (n) passes through the adaptive error correction module (400), using the orthogonal reference signals x Nai and x Nbi as the reference inputs, to track and suppress in real time the low-frequency narrowband periodic component in the air-conducted speech d(n) in step a, and output e N (n); Step d: The frequency component f i (n) from step b is passed through an adaptive depth noise reduction module (300), using orthogonal reference signals x Nai and x Nbi as reference inputs to track and suppress in real time the low-frequency narrowband frequency information after speech fusion, and output a speech signal with optimal intelligibility and signal-to-noise ratio; Step e: Feed the nasal bone-conducted speech x from step a L (n), x R (n) and the output of step e The output e of step c N (n), through the real-time intelligibility enhancement module (100), based on the second-order Volterra model, using the adaptive update algorithm with variable boundaries of filtering error, realizes the online compensation of the high-frequency components of robust bone-conducted speech. To ensure the stability of the real-time intelligibility enhancement module (100), boundary constraint judgment is introduced. When the updated error energy of the corrected prediction is larger than the bone-conducted speech reference input energy where, is the nasal bone-conducted speech x L (n), x R (n) takes the mean value, which is the reference input of the real-time intelligibility enhancement module (100), that is: Introduce a variable boundary r b Step size update μ(n + 1) = r b μ(n) Otherwise μ(n + 1)=μ(n).

6. The bone fusion method according to claim 5, characterized in that, Step c specifically is: The frequency component f i (n) from Step 2 passes through the adaptive error correction module (400), with the orthogonalization reference signals x Nai and x Nbi as the reference inputs, Taking the air conduction voice signal d(n) as the target signal, the environmental periodic component estimation of the adaptive error correction module (400) is: Among them, the i-th group of adaptive coefficients a Ni (n), b Ni (n) are updated as follows: Among them, μ N is the step size of coefficient update, generally set between 0 and 1, and e N (n) is the output of the adaptive error correction module (400).

7. The bone fusion method according to claim 5, characterized in that Step e is specifically: the intelligibility-enhanced bone conduction voice signal y0(n) output by the second-order volterra model is: Among them, the linear training coefficient h 1,i (n + 1) and the non-linear training coefficient h 2,i,j (n + 1) are updated to The self - adaptively updated error e(n) is the difference between the error output e N (n) of the adaptive error correction module (400) and the output error of the adaptive depth noise reduction module (300) difference and e N (n) is expressed as: The noise prediction y of the adaptive depth noise reduction module (300) N (n) is as follows: Its coefficient update formula is:

Citation Information

Patent Citations

  • Respiratory masks, systems and methods

    CN107405508A

  • Respirator mask with integrated bone conduction transducer

    CN110167643A