Communication earphone mixed microphone array body speech enhancement method in high-noise environment
By using a hybrid microphone array and an adaptive beamforming algorithm in communication headphones, the problem of the ontology voice signal being masked or distorted in high-noise environments is solved, and higher speech clarity and noise suppression effects are achieved.
Patent Information
- Application Number
- CN202510236450.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
In high noise environments, traditional communication headphones are difficult to effectively suppress noise, resulting in the body voice signal being masked or distorted, affecting the accuracy and comprehensibility of call information.
Using a hybrid microphone array, including two air-conducting microphones and a bone-conducting microphone, the ontology voice activity detection is performed through the bone-conducting microphone, and adaptive noise suppression and speech enhancement are used in high-noise environments using an adaptive beamforming algorithm (GSC).
It significantly improves the clarity and comprehensibility of ontology speech in high noise environments, overcomes the influence of reverberation effect on speech activity detection, and enhances the robustness and noise suppression effect of the system.
Smart Images

Figure CN120151707A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of adaptive speech enhancement and noise reduction, and particularly relates to a method for enhancing the speech of a hybrid microphone array body of a communication headset in a high-noise environment. Background Art
[0002] In high-noise environments such as industry, transportation, and military, the ambient noise usually exceeds 90 dB and even reaches above 100 dB. When making a voice call in this environment, the body voice signal is easily masked or distorted by background noise, resulting in inaccurate transmission of call information and increasing the risk of misunderstanding and decision-making errors.
[0003] Traditional communication headsets usually use air-conduction microphone arrays for speech enhancement, and beamforming technology is the core. This technology selectively enhances speech signals from specific directions while suppressing noise from other directions by weighting and combining multi-channel signals collected by the microphone array, thereby generating an enhanced single-channel speech signal. The Generalized Sidelobe Canceller (GSC) is an adaptive beamforming algorithm widely used in speech processing and communication systems. Although GSC shows excellent background noise suppression effects in conventional call environments, its noise reduction effect is significantly limited in high-noise environments. To improve the noise reduction effect in high-noise environments, researchers have taken a series of comprehensive measures and technical improvements. The soft-decision adaptive mode controller uses the power ratio of the output signals of the fixed beamformer and the blocking matrix to obtain the weight vector update for controlling the adaptive noise canceller (Choi M.S, Baik C.H., Park Y.C., and Kang H.G., A soft-decision adaptation mode controller for an efficient frequency-domain generalized sidelobe canceller [C] / / Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP). Seattle: IEEE, 2007:893-896.). This method can further improve the noise reduction performance of the GSC structure without particularly considering the noise type and the signal-to-noise ratio of the input signal. However, in high-noise environments, especially in small enclosed spaces, the reverberation effect will significantly affect the noise reduction effect. The noise reflection wave may be close to or even coincide with the direct speech path, resulting in strong noise reflection. This not only increases the intensity of the background noise but also makes the speech activity detection method based on the power ratio of the output signals of the fixed beamformer and the blocking matrix prone to misjudgment. The power ratio distortion problem makes it difficult for the system to accurately identify speech activities, thereby affecting the overall noise reduction effect and causing the original body speech signal to be masked or distorted. Therefore, how to significantly improve the clarity and intelligibility of the original body speech in high-noise environments, overcome the challenges brought by the reverberation effect, and provide a more robust original body speech activity detection mechanism has become a key problem urgently needed to be solved in the development of communication headset technology. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for enhancing the original body speech of a communication headset hybrid microphone array in a high-noise environment to improve the communication speech quality in a high-noise environment.
[0005] Technical solution for achieving the object of the present invention: A method for enhancing the voice of a hybrid microphone array body of a communication headset in a high-noise environment, the steps are as follows:
[0006] Step 1: Configure an air-conduction microphone on the outer side of each of the two earcups, and configure a bone-conduction microphone on the inner side of one of the earcups close to the skin in front of the ear. The above two air-conduction microphones and one bone-conduction microphone form a hybrid microphone array, which picks up sound signals in real time and obtains a digital array signal through a multi-channel synchronous analog-to-digital conversion circuit;
[0007] Step 2: Use one bone-conduction microphone channel obtained in Step 1 for detecting the voice activity of the body, and the two air-conduction microphone channels form a GSC array structure for adaptively enhancing the voice signal of the body;
[0008] Step 3: Before entering a high-noise environment, use the bone-conduction microphone voice activity detection result obtained in Step 2 to control the air-conduction microphone array to update the adaptive beam weight vector during the voice activity of the body, and obtain the static beam weight vector of the GSC main branch for focusing on the voice signal of the body;
[0009] Step 4: After entering a high-noise environment, fix the static beam weight vector of the GSC main branch obtained in Step 3, and use the bone-conduction microphone voice activity detection result obtained in Step 2 to control the adaptive noise cancellation weight vector of the GSC auxiliary branch to be updated during the period without voice activity of the body;
[0010] Step 5: In a high-noise environment, use the main branch static beam weight vector obtained in Step 3 and the auxiliary branch noise cancellation weight vector obtained in Step 4 to obtain the GSC processing result of the microphone array, then suppress the noise residue through a post-filter, and finally obtain the enhanced voice signal of the body through time-domain waveform reconstruction processing.
[0011] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the above method are implemented.
[0012] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0013] A computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1) Hybrid microphone array structure design: A hybrid microphone array structure composed of two air-conduction microphones and one bone-conduction microphone is used to pick up sound signals in real time, and a digital array signal is obtained through a multi-channel synchronous analog-to-digital conversion circuit (ADC). This design effectively solves the deficiencies of traditional microphone arrays in noise suppression and speech clarity, providing a hardware basis for speech enhancement in high-noise environments; 2) Utilizing the characteristic that the bone-conduction microphone is not affected by environmental noise, high-precision detection of the user's own speech activity is realized. This method overcomes the influence of reverberation effects and high noise on the detection of the user's own speech activity, providing reliable data support for noise suppression and speech enhancement in high-noise environments, and significantly improving the robustness of the system in high-noise environments; 3) With the assistance of the bone-conduction microphone to the air-conduction microphone, both the high-frequency components of the speech signal are enhanced, and the clarity of the low-frequency components is improved, while enhancing its low-frequency components. An adaptive speech enhancement processing strategy for heterogeneous microphone fusion is adopted, which not only effectively suppresses strong noise interference but also greatly improves the clarity and naturalness of the user's own speech, providing an innovative technical idea and solution for improving the speech quality of communication headsets in high-noise environments. Description of the Drawings
[0015] Figure 1 It is a layout structure diagram of the hybrid microphone array of the communication headset in the use of the present invention.
[0016] Figure 2 It is a block diagram of the overall system structure of the present invention.
[0017] Figure 3 It is a main beam structure diagram of the array that adaptively optimizes and solidifies the focus on the user's own speech of the present invention.
[0018] Figure 4 It is a GSC structure diagram of the present invention.
[0019] Figure 5 It is a structure diagram of the post-filter of the present invention.
[0020] Figure 6 It is a time-frequency spectrogram and waveform diagram of the air-conduction speech signal collected by the left-ear air-conduction microphone and the bone-conduction speech signal collected by the bone-conduction microphone under the condition that the environmental noise level is 100 dB.
[0021] Figure 7 It is a detection result diagram of the user's own speech based on the air-conduction speech signal and the bone-conduction speech signal in a high-noise environment.
[0022] Figure 8 It is a waveform diagram and time-frequency spectrogram of the speech signal after the noisy speech signal is processed by the GSC structure of the present invention.
[0023] Figure 9The waveform diagram and time-frequency spectrum diagram of the voice signal after the noisy voice signal is processed by the GSC structure of the present invention and post-filtered. Detailed implementation manner
[0024] In view of the urgent need to improve the voice clarity and intelligibility of the communication earphone body in a high-noise environment, the present invention conducts research on the voice enhancement technology of the hybrid microphone array body that combines a bone conduction microphone and an air conduction microphone, and proposes a method for enhancing the voice of the hybrid microphone array body of a communication earphone in a high-noise environment. This method uses a communication earphone equipped with both a bone conduction microphone and an air conduction microphone array, integrates the advantages of the two heterogeneous microphones, overcomes the limitations of traditional air conduction microphone arrays in voice enhancement and noise reduction, and significantly improves the voice communication quality in a high-noise environment. Specifically, this method uses the bone conduction microphone channel to detect the voice activity of the body, and at the same time constructs a GSC array structure that adaptively enhances the voice signal of the body with two air conduction microphone channels. Before entering a high-noise environment, the static beam weight vector of the GSC main branch that focuses on the voice signal of the body is adaptively optimized and solidified; after entering a high-noise environment, the voice activity detection result of the bone conduction microphone is used to control the adaptive noise canceller in the GSC structure to cancel the noise in the main branch. Finally, the GSC processing result is further suppressed of noise residue through a post-filter, and after time-domain waveform reconstruction processing, the enhanced voice signal of the body is output. This method enhances the voice signal through heterogeneous microphone fusion, significantly improving the clarity and intelligibility of the voice signal in a high-noise environment, and providing an effective solution for improving the voice quality of communication earphones.
[0025] Combined with Figure 1 , the layout structure diagram of the hybrid microphone array of the communication earphone used in the present invention. An air conduction microphone is configured on the outside of each of the two ear cups, and a bone conduction microphone is configured on the inside of the left ear cup close to the skin in front of the ear. The aforementioned two air conduction microphones and one bone conduction microphone form a hybrid microphone array, which picks up sound signals in real time and obtains digital array signals through a multi-channel synchronous analog-to-digital conversion circuit (ADC).
[0026] Combined with Figure 2 , the present invention includes a total of three modules, namely, the voice activity detection module of the body based on the bone conduction microphone channel, the GSC module that uses the bone conduction microphone to assist the air conduction microphone array to enhance the voice of the body, and the post-filter module. The specific steps are as follows:
[0027] Step 1: An air conduction microphone is configured on the outside of each of the two ear cups, and a bone conduction microphone is configured on the inside of the left ear cup close to the skin in front of the ear. The aforementioned two air conduction microphones and one bone conduction microphone form a hybrid microphone array, which picks up sound signals in real time and obtains digital array signals through a multi-channel synchronous analog-to-digital conversion circuit;
[0028] Step 2: Use one bone conduction microphone channel obtained in Step 1 to perform detection of the body voice activity, and the two air conduction microphone channels form a GSC array structure for adaptively enhancing the body voice signal, which specifically includes the following steps:
[0029] Step 2-1: In the communication earphone, take the voice signal y 1 (n) collected by the air conduction microphone array as the reference signal, perform a delay process on it for τ 1 sample points, set it to 32 in the measured data, and construct a (2τ BF + 1)-order FIR filter according to the delay amount and the known static beam weight vector w 1 (n) for weighting another air conduction voice signal y 2 (n) to maximize the signal strength of the body voice. By synthesizing the delayed reference signal y 1 (n - τ 1 ) and the filter output signal , obtain the main branch signal y C (n) and the auxiliary branch signal y B (n):
[0030]
[0031] n = n l , n l + 1, …, n l + L 0 - 1, n l = l × R
[0032]
[0033] In the formula, L 0 represents the frame length, R represents the frame shift, n l represents the first sample point of the l-th frame, and w BF (i) is the coefficient of the static beam weight vector w BF (n), representing the weight of each delayed sample point. In the processing of measured data, set L 0 = 256, R = 128.
[0034] Step 2-2: Input the main branch signal y 2 (n - τ C ) and the auxiliary branch signal y 2 (n) with a delay of τ B sample points into a (2τ 2 + 1)-order noise canceller as the target reference signal and the noise reference signal respectively. By synthesizing the environmental noise in the auxiliary branch signal into a real-time estimate of the noise component in the main branch signal And subtract it from the main branch signal to finally obtain the enhanced body voice signal. Then the output of the GSC can be expressed as:
[0035]
[0036] where w NC (n) represents the adaptive noise cancellation weight vector, and w NC (i) is the coefficient of the adaptive noise cancellation weight vector, representing the weight of each delayed sample point. In the processing of measured data, set τ 2 = 64.
[0037] Step 2-3: Frame the bone conduction microphone channel signal. After windowing each frame of data, perform a short-time discrete Fourier transform (Short-Time Fourier Transform, STFT) to calculate the short-time power spectral density p(k, l), where k and l are the frequency point number and frame number respectively; define k L and k U as the frequency point numbers corresponding to the lowest frequency and the highest frequency of the bone conduction microphone signal bandwidth respectively. In the processing of measured data, the lowest frequency is set to 300 Hz and the highest frequency is set to 1500 Hz. Define the normalized power spectral probability density:
[0038]
[0039] Furthermore, obtain the spectral entropy of the l-th frame of the bone conduction microphone signal:
[0040]
[0041] Step 2-4: After entering any new call environment, take the first M-frame leading data segment without body voice to calculate the mean spectral entropy of the ambient noise. The voice activity detection method used in the present invention is a double-threshold voice activity detection algorithm based on spectral entropy (Lai Chunqiang, Wang Qian. A dynamic double-threshold voice activity detection algorithm based on spectral entropy [J]. Ordnance Industry Automation, 2019, 38(01): 39-41.):
[0042]
[0043] Furthermore, obtain two thresholds, a high threshold and a low threshold, for detecting the body voice activity. Define the high threshold T 1 and the low threshold T 2 as:
[0044] T 1 = α 1 T r , T 2 = α 2 T r , α 2 < α1
[0045] In the formula, in the processing of measured data, set α 1 = 0.99, α 2 = 0.97.
[0046] Judge whether the body voice exists according to the set threshold value:
[0047]
[0048] In the formula, f VAD (l) represents the detection result of the body voice of the bone conduction microphone: if the spectral entropy value of the current frame is greater than T 1 then it is determined as a non-voice segment, f VAD (l) = 0; if the current frame is less than T 2 , then it is determined as a voice segment f VAD (l) = 1; if the spectral entropy value is between the two, then it is determined as a pending segment, f VAD (l) = -1.
[0049] Step 3: Combine Figure 3 , before entering the high-noise environment, use the detection result of the body voice activity of the bone conduction microphone obtained in Step 2 to control the adaptive beam weight vector update of the air conduction microphone array during the body voice activity, and obtain the static beam weight vector of the GSC main branch that focuses on the body voice signal, which specifically includes the following steps:
[0050] Adopt the Normalized Least Mean Square (NLMS) adaptive filter algorithm to adaptively optimize the static beam weight vector w BF (n) of the GSC main branch during the body voice activity:
[0051]
[0052] In the formula, μ 0 is an adaptive step size, c is a parameter with a very small value, and α BF (l) is a switch factor for controlling the update of the weight vector.
[0053] Step 4: Combine Figure 4 , after entering the high-noise environment, fix the static beam weight vector of the GSC main branch obtained in Step 3, and use the detection result of the body voice activity of the bone conduction microphone obtained in Step 2 to control the update of the adaptive noise cancellation weight vector of the GSC auxiliary branch during the period without body voice activity, which specifically includes the following steps:
[0054] The NLMS adaptive filter algorithm is adopted to control the update of the adaptive noise cancellation weight vector w of the GSC auxiliary branch during the period without the presence of the body speech activity NC (n):
[0055]
[0056] where δ is an adaptive step size, and α NC (l) is a switching factor for controlling the update of the weight vector
[0057] Step 5. Combine Figure 5 , in a high-noise environment, use the main-branch static beam weight vector obtained in Step 3 and the auxiliary-branch noise cancellation weight vector obtained in Step 4 to obtain the GSC processing result of the microphone array, then further suppress the noise residue through a post-filter, and finally obtain the enhanced body speech signal through time-domain waveform reconstruction processing, which specifically includes the following steps:
[0058] Step 5-1. Perform short-time discrete Fourier transform processing on the output signal y G (n) of the GSC, and use the body speech activity detection result of the bone-conduction microphone and the Minima Controlled Recursive Averaging (MCRA) (Cohen I., Berdugo B. Speech enhancement for non-stationary noise environments[J]. Signal processing, 2001, 81(11): 2403-2418.) algorithm combined with minimum control to obtain the noise power spectrum estimation:
[0059]
[0060] where represents the smoothing factor, and α d represents the fixed smoothing factor, which is set to 0.86 in the processing of measured data. P(k, l) represents the a priori speech presence probability. In the present invention, the a priori speech presence probability is obtained by using the energy-entropy ratio W(k, l) (Feng Qian, Yu Qin, Luo Ruisen, etc. MMSE-LSA speech enhancement algorithm based on improved noise estimation[J]. Computer Applications and Software, 2022, 39(11): 141-147.):
[0061]
[0062] where LE(k, l) represents the improved energy calculation, which is set to a = 5 and b = 0.8 in the processing of measured data
[0063] Step 5-2: Combine with the Optimally Modified Log-Spectral Amplitude Estimator (OMLSA) algorithm (Cohen I., Berdugo B. Microphone array post-filtering for non-stationary noise suppression[C] / / 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing. Orlando: IEEE, 2002, 1: I-901-I-904.), and calculate the log-spectrum gain G of the current frame using the noise power spectrum obtained above H1 , thereby obtaining the post-filtering gain of the current frame:
[0064]
[0065] In the formula, G min represents the gain value when there is no speech, and is set to 0.11 in the processing of measured data. Adjust the spectrum amplitude of the noisy speech to obtain the enhanced spectrum amplitude:
[0066] |Y(k, l)| = G o (k, l)·|Y G (k, l)|
[0067] Combine the enhanced spectrum amplitude with the phase information of the noisy speech signal to obtain the enhanced body speech signal Y(k, l).
[0068] Step 5-3: Replace the low-frequency part of the bone-conduction microphone channel signal with the corresponding low-frequency part of the enhanced body speech signal, perform inverse short-time Fourier transform on the frequency-domain signal after spectrum fusion for each frame to obtain the time-domain signal frame, and finally realize the reconstruction of the time-domain waveform through the overlap-save method to obtain the final body speech output signal:
[0069]
[0070] In the formula, represents the effective signal of each non-overlapping part of the frame, represents the overlapping part between frames, and N represents the total number of frames.
[0071] Figure 6Among them, (a) and (b) respectively represent the time-frequency spectrogram and waveform diagram of the air-conducted speech signal collected by the left-ear air-conducted microphone and the bone-conducted speech signal collected by the bone-conducted microphone under the condition that the environmental noise level is 100 dB. It can be seen from the figure that in a high-noise environment, the air-conducted speech signal is severely contaminated by background noise, resulting in waveform distortion and a significant decrease in speech clarity and intelligibility. In contrast, the bone-conducted speech signal, relying on its natural anti-noise ability, is hardly affected by environmental noise and maintains a high level of clarity. However, the energy of the bone-conducted speech signal is mainly concentrated in the low-frequency region, and there is obvious energy attenuation in the high-frequency part. Therefore, the low-frequency band of the bone-conducted speech signal is used for detecting the body speech activity in the present invention.
[0072] Figure 7 Among them, (a) and (b) respectively represent the diagrams of the body speech detection results based on the air-conducted speech signal and the bone-conducted speech signal in a high-noise environment. It can be seen from the figure that in a high-noise environment, the air-conducted speech signal is severely covered by noise, and the waveform is significantly distorted, resulting in difficulty in accurately distinguishing the speech from the background noise and a low reliability of the detection result. In contrast, the low-frequency part of the bone-conducted speech signal is hardly interfered by noise, and the waveform features are clear, ensuring extremely high detection accuracy and being able to effectively separate the body speech and the noise components. This difference highlights the unique advantages of the bone-conducted microphone in a high-noise environment.
[0073] Figure 8 Among them, (a) and (b) respectively represent the waveform diagram and time-frequency spectrogram of the speech signal after the noisy speech signal is processed by the GSC structure of the present invention. It can be seen from the figure that the processed speech signal shows a significant improvement in quality both in terms of waveform and spectrum: the spectrum is clearer and the waveform features are significantly improved. Although there is still a small amount of residual noise in the processing process, the overall quality of the speech signal has been greatly improved, fully demonstrating the effectiveness of the method of the present invention.
[0074] Figure 9 Among them, (a) and (b) respectively represent the waveform diagram and time-frequency spectrogram of the speech signal after the noisy speech signal is processed by the GSC structure of the present invention and post-filtered. It can be seen from the figure that after these two steps of processing, the residual noise in the speech signal is effectively reduced, the processed spectrum is purer, the waveform features are clearer, and the speech quality has been significantly improved.
Claims
1. A method for enhancing speech of a hybrid microphone array of a communication headset in a high noise environment, characterized in that: The steps include: Step 1, an air conduction microphone is arranged on the outside of each of the two earmuffs, and a bone conduction microphone is arranged on the inside of one of the earmuffs close to the skin in front of the ear. The two air conduction microphones and one bone conduction microphone constitute a hybrid microphone array, pick up sound signals in real time, and obtain digital array signals through a multi-channel synchronous analog-to-digital conversion circuit; Step 2: Use a bone conduction microphone channel obtained in step 1 to detect the body voice activity, and the two air conduction microphone channels form a GSC array structure for adaptively enhancing the body voice signal; Step 3: Before entering a high-noise environment, the bone conduction microphone body voice activity detection result obtained in step 2 is used to control the air conduction microphone array to perform adaptive beam weight vector update during the body voice activity to obtain a GSC main branch static beam weight vector that focuses on the body voice signal; Step 4: After entering a high-noise environment, fix the static beam weight vector of the GSC main branch obtained in step 3, and use the bone conduction microphone body voice activity detection result obtained in step 2 to control the adaptive noise cancellation weight vector of the GSC auxiliary branch to update during the period when there is no body voice activity; Step 5. In a high-noise environment, the main branch static beam weight vector obtained in step 3 and the auxiliary branch noise cancellation weight vector obtained in step 4 are used to obtain the microphone array GSC processing result, and then the noise residue is suppressed through a post-filter. Finally, the enhanced body speech signal is obtained through time domain waveform reconstruction processing.
2. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 1, characterized in that: Step 2 uses a bone conduction microphone channel obtained in step 1 to detect the body voice activity, and the two air conduction microphone channels form a GSC array structure for adaptively enhancing the body voice signal, which specifically includes the following steps: Step 2-1: In the communication headset, the speech signal y1(n) collected by the air conduction microphone array is used as the reference signal, and a delay of τ1 samples is performed on it, and the delay amount and the known static beam weight vector w are used to calculate the delay value. BF (n) Construct a (2τ1+1)th order FIR filter to perform weighted processing on the other air conduction speech signal y2(n) to maximize the signal strength of the main speech; by combining the delayed reference signal y1(n-τ1) and the filter output signal Synthesize to get the main branch signal y C (n) and auxiliary branch signal y B (n): n=n l ,n l +1,…,n l +L0-1,n l =l×R In the formula, L0 represents the frame length, R represents the frame shift, and n l represents the first sample point of the lth frame, w BF (i) is the static beam weight vector w BF The coefficient of (n) represents the weight of each delayed sample point; Step 2-2: delay the main branch signal y by τ2 samples C (n-τ2) and auxiliary branch signal y B (n) are respectively input into the (2τ2+1)-order noise canceller as the target reference signal and the noise reference signal; the environmental noise in the auxiliary branch signal is synthesized into a real-time estimate of the noise component in the main branch signal And subtract it from the main branch signal to finally obtain the enhanced body speech signal; then the output of GSC can be expressed as: In the formula, w NC (n) represents the adaptive noise cancellation weight vector, w NC (i) is the coefficient of the adaptive noise cancellation weight vector, which represents the weight of each delayed sample point; Step 2-3, divide the bone conduction microphone channel signal into frames, add windows to each frame of data, and perform short-time discrete Fourier transform to calculate the short-time power spectrum density p(k,l), where k and l are the frequency point number and frame number respectively; define k L and k U The frequency point numbers corresponding to the lowest and highest frequencies of the bone conduction microphone signal bandwidth are defined as follows: Then the spectral entropy of the bone conduction microphone signal of the first frame is obtained: Step 2-4: After entering any new call environment, take the M-frame leading data segment without the main voice to calculate the spectral entropy mean of the ambient noise: Then we get the high and low thresholds for detecting the voice activity of the subject; the high threshold T1 and the low threshold T2 are defined as: T1=α1T r ,T2=α2T r ,α2<α1 Determine whether the host voice exists based on the set threshold: In the formula, f VAD (l) represents the body speech detection result of the bone conduction microphone: if the spectral entropy value of the current frame is greater than T1, it is judged as a non-speech segment, f VAD (l) = 0; if the current frame is less than T2, it is determined to be a speech segment f VAD (l) = 1; if the spectral entropy value is between the two, it is determined to be a pending segment, f VAD (l) = -1.
3. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 2, characterized in that: Set α1=0.99, α2=0.
97.
4. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 1, characterized in that: Step 3: Before entering the high-noise environment, the bone conduction microphone body voice activity detection result obtained in step 2 is used to control the air conduction microphone array to adaptively update the static beam weight vector during the body voice activity to obtain the GSC main branch static beam weight vector focusing on the body voice signal, which specifically includes the following steps: The normalized minimum mean square error adaptive filter algorithm is used to adaptively optimize the static beam weight vector w of the GSC main branch during the main body speech activity. BF (n): In the formula, μ0 is an adaptive step size, c is a very small parameter, and α BF (l) is the switching factor for updating the control weight vector.
5. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 1, characterized in that: Step 4, after entering the high noise environment, fix the static beam weight vector of the GSC main branch obtained in step 3, and use the bone conduction microphone body voice activity detection result obtained in step 2 to control the adaptive noise cancellation weight vector of the GSC auxiliary branch to update during the period when there is no body voice activity, which specifically includes the following steps: The NLMS adaptive filter algorithm is used to control the adaptive noise cancellation weight vector w of the GSC auxiliary branch during the period without the main speech activity. NC (n) To update: In the formula, δ is an adaptive step size, α NC (l) is the switching factor for updating the control weight vector.
6. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 1, characterized in that: Step 5: In a high-noise environment, the microphone array GSC processing result is obtained by using the main branch static beam weight vector obtained in step 3 and the auxiliary branch noise cancellation weight vector obtained in step 4, and then the noise residue is further suppressed by a post-filter. Finally, the enhanced body speech signal is obtained by time domain waveform reconstruction processing, which specifically includes the following steps: Step 5-1: GSC output signal y G (n) Perform short-time discrete Fourier transform and use the body voice activity detection result of the bone conduction microphone combined with the minimum value controlled recursive averaging algorithm to estimate the noise power spectrum: In the formula, represents the smoothing factor, α d represents a fixed smoothing factor, P(k,l) represents the prior probability of speech existence, and the prior probability of speech existence is obtained using the energy entropy ratio W(k,l): Where LE(k,l) represents the improved energy calculation, a and b are control parameters; Step 5-2: Combine the minimum mean square error log spectrum amplitude algorithm and use the noise power spectrum obtained in step 5-1 to calculate the log spectrum gain G of the current frame. H1 , thus obtaining the post-filter gain of the current frame: In the formula, G min Indicates the gain value when speech does not exist, set to the minimum signal-to-noise ratio of the noisy speech signal; adjust the amplitude of the noisy speech spectrum to obtain the enhanced spectrum amplitude: |Y(k,l)|=G o (k,l)·|Y G (k,l)| The enhanced spectrum amplitude is combined with the phase information of the noisy speech signal to obtain the enhanced body speech signal Y(k,l); Step 5-3, replace the low-frequency part of the bone conduction microphone channel signal with the corresponding low-frequency part of the enhanced body voice signal, and perform inverse short-time Fourier transform on each frame of the frequency domain signal after spectrum fusion to obtain a time domain signal frame, and finally realize time domain waveform reconstruction through overlapping retention method to obtain the final body voice output signal: In the formula, Indicates the valid signal of the non-overlapping part of each frame, represents the overlap between frames, and N represents the total number of frames.
7. The method for enhancing speech of a hybrid microphone array body of a communication headset in a high noise environment according to claim 6, characterized in that: In step 5-1, a=5, b=0.
8.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.