A method for enhancing the voice of an opposite party during face-to-face conversation using high-noise intercom ear muffs
By using microphone arrays and voice signal processing technology, the problem of protective earmuffs hindering communication in high-noise environments has been solved, enabling clear transmission of the other party's voice in high-noise environments, reducing costs and improving wearing comfort.
Patent Information
- Application Number
- CN202411776920.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-05
AI Technical Summary
In high-noise environments, while existing protective earmuffs can effectively block noise, they hinder verbal communication between workers, leading to poor information transmission and increasing safety hazards at work. At the same time, earmuffs with integrated communication modules are expensive and uncomfortable.
A microphone array consisting of a left ear microphone, a right ear microphone, and a mouth-side microphone is used. Noise suppression and speech enhancement are performed through speech signal processing technology, including fixed beamforming, speech endpoint detection, and adaptive filter adjustment, to ensure clear transmission of the other party's speech in high-noise environments.
While reducing production costs, it effectively preserves and enhances the other party's voice signal, ensuring clarity of face-to-face communication in high-noise environments, improving work efficiency and reducing wearing discomfort.
Smart Images

Figure CN119694326B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of speech processing, and particularly relates to a method for enhancing the speech of an opposite party during face-to-face conversation using a high-noise workshop protective earmuff. BACKGROUND
[0002] Noise is generally defined as unwanted sound, interference, or irregular sound in a particular environment. Long-term exposure to noise, especially high-intensity noise, can cause serious harm to human physical and mental health, and even lead to permanent hearing impairment. Studies have shown that long-term exposure to high-intensity noise (such as 85 dB or higher) can lead to noise-induced hearing loss (NIHL). NIHL is a gradually accumulating hearing impairment, which usually manifests as high-frequency hearing loss in the early stage, and the damage range gradually expands to low-frequency hearing over time. In addition, long-term exposure to high-noise environments can also cause psychological problems such as anxiety, depression, and other emotional disorders, as well as cardiovascular health problems such as high blood pressure and heart disease. Therefore, it is particularly important to take effective protective measures in work environments where noise pollution is serious. For example, in some high-noise factory workshops, workers must wear earplugs, earmuffs, or protective earphones to avoid noise-induced hearing damage.
[0003] In recent years, with the continuous progress of noise protection technology, protective earmuffs and various noise-cancelling earphones have been widely used in many fields, especially in industrial sites and high-noise work environments. These protective devices effectively reduce the impact of external noise on workers and reduce the risk of noise-induced hearing impairment. However, with the popularity of these protective measures, some new problems have also arisen. Specifically, although protective earmuffs can significantly shield the noise in the surrounding environment, they also hinder normal language communication between workers. In many industrial production and work environments, language communication is crucial, which not only helps to improve work efficiency, but also relates to the safety of employees. Lack of effective communication can lead to miscommunication, increase the risk of safety in work, and even cause unnecessary accidents. Therefore, how to ensure language communication in work while protecting the hearing health of employees has become a technical challenge that needs to be addressed.
[0004] To solve this problem, the current mainstream solution is to combine protective earmuffs with earphone communication modules. This solution not only effectively isolates external noise and protects hearing safety, but also has the function of real-time voice communication, greatly improving the communication efficiency of users during work or operation. However, this integrated design also brings some challenges, especially in terms of cost and wearing comfort. Due to the addition of the communication module, the price of such protective earmuffs is usually higher, and the weight of the earmuffs is also increased, which may cause discomfort during long-term wear, thereby affecting the comfort of the user. Therefore, how to use microphone array signal processing technology and related speech enhancement algorithms to effectively retain the speech signal of the other party while wearing protective earmuffs to isolate noise, and effectively improve the speech quality through noise suppression and speech enhancement technology, has become an important issue that cannot be ignored in current research and practical application. SUMMARY
[0005] The purpose of the present application is to provide a method for enhancing the speech of the other party when face-to-face talking with protective earmuffs in a high-noise workshop.
[0006] The technical solution to achieve the purpose of the present application is a method for enhancing the speech of the other party when face-to-face talking with protective earmuffs in a high-noise workshop, the steps are as follows:
[0007] Step 1: Synchronously collect the noisy speech signals of the body and the other party during face-to-face talking in a high-noise environment using a microphone array composed of a left ear microphone, a right ear microphone, and a mouth edge microphone, and pre-process the collected signals;
[0008] Step 2: Perform fixed beamforming on the speech signals of the left ear microphone channel and the right ear microphone channel obtained in Step 1 to obtain the main channel and auxiliary channel input signals of the first-level noise interference suppression unit;
[0009] Step 3: Perform speech endpoint detection on the main channel speech signal obtained in Step 2 to determine whether it is a speech period or a non-speech period;
[0010] Step 4: The filter of the first-level noise interference suppression unit adaptively adjusts the filter coefficients using the main channel and auxiliary channel speech signals obtained in Step 2 and the decision information obtained in Step 3 to complete the cancellation of environmental noise interference;
[0011] Step 5: Perform body speech endpoint detection on the speech signal of the mouth edge microphone channel obtained in Step 1 to determine whether it is a body speech period or a non-body speech period;
[0012] Step 6, using the voice signal of the mouth edge microphone channel obtained in step 1, the noise suppressed voice signal obtained in step 4 and the decision information obtained in step 5, taking the noise suppressed voice signal as the main channel input signal of the second level main voice interference suppression unit, taking the voice signal of the mouth edge microphone channel as the auxiliary channel input signal, adaptively adjusting the filter coefficient according to the decision information, completing the cancellation of the main voice interference, obtaining the enhanced signal of the opposite voice, and playing out by the loudspeaker built in the protective earmuff.
[0013] An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method for enhancing the voice of the opposite party during face-to-face conversation of the high-noise intercom protective earmuff.
[0014] A computer readable storage medium having a computer program stored thereon, wherein the program is executable on a processor to implement the method for enhancing the voice of the opposite party during face-to-face conversation of the high-noise intercom protective earmuff.
[0015] A computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the method for enhancing the voice of the opposite party during face-to-face conversation of the high-noise intercom protective earmuff.
[0016] Compared with the prior art, the present application has the following advantages: 1) the present application can be applied to protective earmuffs, and in a high-noise environment, the voice signal of the opposite party during face-to-face conversation can be effectively retained and enhanced, ensuring that the wearer of the earmuff can still communicate face-to-face; 2) the system structure of the present application is composed of low-cost microphones, and compared with the protective earmuff of the prior art which integrates a communication module, the production cost of the present application is significantly reduced, and the present application has strong market competitiveness; 3) the system structure of the present application is simple in design and easy to implement, and can provide a reliable and robust voice communication solution for the wearer of the protective earmuff in a high-noise environment.
[0017] The present application will be described in further detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 The system structure diagram of the method of the present application.
[0019] Figure 2 The schematic diagram of the three-element microphone array for collecting noise-containing voice data in the present application.
[0020] Figure 3 The time-domain waveform diagram of the noise-containing voice signal collected by the three microphone channels and the corresponding time-frequency spectrum diagram.
[0021] Figure 4Time-domain waveform diagram and corresponding time-frequency spectrum diagram of preprocessed noisy speech signal for three microphone channels.
[0022] Figure 5 Time-domain waveform diagram and corresponding time-frequency spectrum diagram of main channel speech signal output by fixed beamforming unit.
[0023] Figure 6 Detection result diagram of speech endpoint detection of main channel speech signal output by beamforming unit by detection control unit.
[0024] Figure 7 Time-domain waveform and corresponding time-frequency spectrum diagram of speech signal output by noise interference suppression unit.
[0025] Figure 8 Detection result diagram of speech endpoint detection of mouth-edge microphone channel speech signal by detection control unit.
[0026] Figure 9 Time-domain waveform and corresponding time-frequency spectrum diagram of speech signal output by body speech interference suppression unit. DETAILED DESCRIPTION
[0027] In combination Figure 1 , the application proposes a method for enhancing the speech of the opposite party when face-to-face talking in a high-noise workshop using protective earmuffs, comprising the following steps:
[0028] Step 1, synchronously collecting the noisy speech signals of the body and the opposite party when face-to-face talking in a high-noise environment using a microphone array composed of a left ear microphone, a right ear microphone and a mouth-edge microphone, and pre-processing the collected signals;
[0029] Further, step 1 is specifically:
[0030] Step 1-1, numbering the left ear microphone, the right ear microphone and the mouth-edge microphone as 1, 2 and 3 respectively, wherein the No. 1 and No. 2 microphones are located at the left and right earmuffs of the protective earmuffs, and the No. 3 microphone is located at the talk rod of the extension mouth edge connected with the protective earmuffs, ensuring that the No. 1 and No. 2 microphones are symmetrically distributed on both sides of the No. 3 microphone.
[0031] Step 1-2, collecting noisy speech signals of each channel using the array setting of step 1-1, in the specific example of the application, the data sampling rate is 16 kHz, and the analysis period length is 16 s. Pre-emphasis processing is performed on the collected noisy speech signals of each channel, and a first-order FIR high-pass digital filter is used to achieve this, and the transfer function of the filter is
[0032] H(z)=1-az -1 0.9<a<1.0
[0033] In the specific embodiment of the present application, the pre-emphasis parameter a is set to 0.95, and z -1 is a unit delay operator in the z transform.
[0034] Step 1-3, short-time frame and window processing are performed on the noise-containing speech signals of each channel obtained by using step 1-2, in the specific embodiment of the present application, the frame length is 20 ms, the frame shift is 10 ms, and the window function is a Hamming window;
[0035] Step 2, fixed beam forming is performed on the speech signals of the left ear microphone channel and the right ear microphone channel obtained by using step 1 to obtain the main channel and the auxiliary channel input signals of the first-stage noise interference suppression unit;
[0036] Further, step 2 is specifically as follows:
[0037] The preprocessed signals of the left ear microphone channel and the right ear microphone channel obtained by using the method of step 1 are subjected to fixed beam forming processing. The output y(n) of the fixed beam former is defined as:
[0038]
[0039] In the formula, M represents the number of elements of the microphone array, x i (n) represents the speech signal collected by the i-th element, where i = 0, 1,..., M-1. w i represents the weight of the i-th filter, τ i represents the time delay estimation of the i-th element relative to the reference element. In the specific embodiment of the present application, the number of elements M of the microphone array is set to 2, and w is 1 / 2. τ i are all 0.
[0040] Step 3, speech endpoint detection is performed on the main channel speech signal obtained by using step 2 to determine whether the current period is a speech period or a non-speech period;
[0041] Further, step 3 is specifically as follows:
[0042] Step 3-1, the main channel signal obtained by using the method of step 2 is subjected to noise reduction processing using the improved spectral subtraction of multi-window spectrum estimation, which is specifically as follows:
[0043] Step 3-1-1, FFT is performed on the current frame signal to calculate its amplitude spectrum |X i (k) and phase spectrum θ i (k) (i represents the i-th frame, and k represents the k-th spectrum line), and smoothing processing is performed between adjacent frames, which is specifically to take M frames before and after the i-th frame to perform averaging, and in the specific embodiment of the present application, M is 3, and the average amplitude spectrum is calculated as:
[0044]
[0045] Step 3-1-2, multi-window spectrum estimation is performed on the current frame signal, and a multi-window spectrum power spectrum density P(k, i) (i represents the i-th frame, and k represents the k-th spectrum line) is calculated. Smoothing is performed between adjacent frames, specifically, 2M+1 frames are taken before and after the i-th frame to average, and in the specific example of the application, M is 3. The smoothed power spectrum density is calculated as:
[0046]
[0047] Step 3-1-3, according to the noise frames of the leading speech-free section, the average power spectrum density value of the noise is calculated as:
[0048]
[0049] In the formula, NIS is the number of frames occupied by the leading speech-free section, and in the specific example of the application, the number of frames NIS of the leading speech-free section is 200; P y (k, i) is the smoothed power spectrum density of the i-th frame obtained by the method described in step 3-1-2.
[0050] Step 3-1-4, the average amplitude spectrum obtained by the method described in step 3-1-1, the smoothed power spectrum density obtained by the method described in step 3-1-2, and the noise average power spectrum density value obtained by the method described in step 3-1-4 are used to calculate a gain factor using the spectral subtraction relationship. The gain factor is defined as:
[0051]
[0052] In the formula, α is an over-subtraction factor, and β is a gain compensation factor. In the specific example of the application, the over-subtraction factor α is 2.8, and the gain compensation factor β is 0.001.
[0053] Step 3-1-5, the average amplitude spectrum obtained by the method described in step 3-1-1 and the gain factor obtained by the method described in step 3-1-4 are used to obtain a noise-reduced speech signal:
[0054]
[0055] Step 3-2, using the noise-reduced speech signal obtained by the method described in step 3-1, the short-time sub-band spectrum entropy of each frame signal is calculated, specifically:
[0056] Step 3-2-1, the current frame signal is divided into N b1 sub-bands, and in the specific example of the application, the number of sub-bands N b1 is 40. The sub-band energy of each sub-band is calculated to obtain the sub-band energy Eb1 (m, i), 1≤m≤N b1 ;
[0057] Step 3-2-2, using the sub-band energy obtained by the method described in step 3-2-1, calculate the sub-band energy probability of the mth sub-band in the ith frame as:
[0058]
[0059] Step 3-2-3, using the sub-band energy probability obtained by the method described in step 3-2-2, calculate the sub-band spectrum entropy of the ith frame as:
[0060]
[0061] Step 3-3, using the short-time sub-band spectrum entropy obtained by the method described in step 3-2, set two threshold limits, and use the double threshold method for voice endpoint detection, specifically:
[0062] Step 3-3-1, set a lower threshold limit T2, and make a rough judgment, and the voice start and end points are located outside the time points corresponding to the intersection of the threshold and the sub-band spectrum entropy envelope if they are below the threshold;
[0063] Step 3-3-2, set a higher threshold limit T1, and search left and right from the voice start and end points obtained by the method described in step 3-3-1, respectively, to find two points where the sub-band spectrum entropy envelope intersects the threshold T1, and obtain the voice start and end point positions determined by the double threshold method.
[0064] Step 4, using the main channel and auxiliary channel voice signals obtained by step 2 and the decision information obtained by step 3, the filter of the first level noise interference suppression unit adaptively adjusts the filter coefficients, and completes the cancellation of environmental noise interference;
[0065] Further, step 4 is specifically:
[0066] Step 4-1, using the voice endpoint detection result obtained by the method described in step 3, determine whether the current frame is a voice frame or a non-voice frame;
[0067] Step 4-2, using the main channel and auxiliary channel signals obtained by the method described in step 2 and the judgment information obtained by the method described in step 4-1, perform first level noise interference suppression processing, specifically:
[0068] Step 4-2-1, filter the auxiliary channel signal, and the filter output is:
[0069] y1(n) NLMS =W1 T (n)X1(n)
[0070] Wherein, W1(n) is the filter coefficient of the current n moment, X1(n) is the auxiliary channel input signal of the current n moment;
[0071] Step 4-2-2, subtract the filter output signal from the main channel signal to obtain the output signal after noise interference suppression processing;
[0072] Step 4-2-3, if the current frame is a non-speech frame, update the filter coefficient, and the update formula is:
[0073]
[0074] Wherein, e1(n) is the signal obtained by subtracting the filter output signal from the main channel signal, c is a very small constant, in the specific embodiment of the application, the modified step constant μ1 is 0.64, and the constant c is 0.0001;
[0075] Step 5, using the speech signal of the mouth edge microphone channel obtained in step 1, the body speech endpoint detection is performed to determine whether the current is a body speech period or a non-body speech period;
[0076] Further, step 5 is specifically:
[0077] Step 5-1, using the speech signal of the mouth edge microphone channel obtained by the method in step 1, the short-time sub-band spectrum entropy of each frame signal is calculated, and the calculation is specifically:
[0078] Step 5-1-1, the current frame signal is divided into N b2 sub-bands, in the specific embodiment of the application, the number of sub-bands N b2 is 40, the sub-band energy of each sub-band is calculated to obtain the sub-band energy E b2 (m, i) of the mth sub-band in the ith frame, 1≤m≤N b2 ;
[0079] Step 5-1-2, using the sub-band energy obtained by the method in step 5-1-1, the sub-band energy probability of the mth sub-band in the ith frame is calculated as:
[0080]
[0081] Step 5-1-3, using the sub-band energy probability obtained by the method in step 5-1-2, the sub-band spectrum entropy of the ith frame is calculated as:
[0082]
[0083] Step 5-2, using the short-time sub-band spectrum entropy obtained by the method in step 5-1, two threshold values are set, and the double-threshold method is used for speech endpoint detection, and the calculation is specifically:
[0084] Step 5-2-1, set a lower threshold T2', make a rough judgment, and the speech frame below the threshold T2' is the main speech frame, and the start and end points of the main speech frame are located outside the time points corresponding to the intersection points of the threshold and the sub-band spectral entropy envelope;
[0085] Step 5-2-2, set a higher threshold T1', and search left and right from the start and end points of the speech frame obtained from the method of step 5-2-1, respectively, to find two points where the sub-band spectral entropy envelope intersects the threshold T1', and obtain the start and end points of the main speech frame determined by the double-threshold method.
[0086] Step 6, use the speech signal of the mouth microphone channel obtained in step 1, the noise-suppressed speech signal obtained in step 4, and the judgment information obtained in step 5, take the noise-suppressed speech signal as the main channel input signal of the second-level main speech interference suppression unit, take the speech signal of the mouth microphone channel as the auxiliary channel input signal, and adaptively adjust the filter coefficients according to the judgment information to complete the cancellation of the main speech interference, obtain the enhanced signal of the opposite speech, and play it out through the loudspeaker built in the protective earmuff.
[0087] Further, step 6 is specifically:
[0088] Step 6-1, use the speech endpoint detection result obtained by the method of step 5 to determine whether the current frame is a main speech frame or a non-main speech frame;
[0089] Step 6-2, use the speech signal of the mouth microphone channel obtained by the method of step 1, the noise-suppressed speech signal obtained by the method of step 4, and the judgment information obtained by the method of step 6-1, take the noise-suppressed speech signal as the main channel input signal of the second-level main speech interference suppression unit, and take the speech signal of the mouth microphone channel as the auxiliary channel input signal to perform the second-level main speech interference suppression processing, which is specifically:
[0090] Step 6-2-1, filter the auxiliary channel signal, and the filter output is:
[0091] y2(n) NLMS =W2 T (n)X2(n)
[0092] In the formula, W2(n) is the filter coefficient at the current time n, and X2(n) is the auxiliary channel input signal at the current time n;
[0093] Step 6-2-2, subtract the main channel signal from the filter output signal to obtain the output signal after noise interference suppression processing;
[0094] Step 6-2-3, if the current frame is a main speech frame, update the filter coefficient, and the update formula is:
[0095]
[0096] where e2(n) is the signal of the main channel signal minus the filter output signal, c is a very small constant, in the specific embodiment of the present application, the revised step constant μ2 is 0.64, and the constant c is 0.0001.
[0097] Figure 2 The schematic diagram of the three-microphone array for collecting noisy speech data is given. First, the left ear microphone, the right ear microphone and the mouth edge microphone are marked as No. 1, No. 2 and No. 3 microphones respectively, the No. 1 and No. 2 microphones are symmetrically placed on the left and right ear covers of the protective ear cover, and the boom is installed on the protective ear cover, and the No. 3 microphone is placed on the boom.
[0098] Figure 3 In the figure, (a), (b) and (c) are respectively the time-domain waveform diagram of the noisy speech signal of the left ear microphone channel, the right ear microphone channel and the mouth edge microphone channel and the corresponding time-frequency spectrum diagram. It can be seen that the speech signal of the left ear microphone channel and the right ear microphone channel is mostly submerged in noise, especially the low-frequency part of the speech signal. The body speech signal collected by the mouth edge microphone channel is relatively pure due to its special response characteristics and the proximity to the mouth edge.
[0099] Figure 4 In the figure, (a), (b) and (c) are respectively the time-domain waveform diagram of the noisy speech signal of the left ear microphone channel, the right ear microphone channel and the mouth edge microphone channel and the corresponding time-frequency spectrum diagram. It can be seen that the noisy speech signal after pre-emphasis processing suppresses most of the low-frequency noise and enhances the speech signal.
[0100] Figure 5 The time-domain waveform diagram of the main channel speech signal output by the fixed beam forming unit and the corresponding time-frequency spectrum diagram are given. It can be seen that after the noisy speech signal of the left ear microphone channel and the right ear microphone channel is subjected to fixed beam forming, the noise signal is further suppressed, and the speech signal is enhanced.
[0101] Figure 6 In the figure, (a) is the time-domain waveform of the noisy speech signal input into the detection control unit, i.e. the waveform of the main channel speech signal output by the beam forming unit, (b) is the speech signal waveform after the spectrum subtraction of the multi-window spectrum estimation improvement of the signal, and (c) is the corresponding spectrum entropy value diagram obtained by calculating the sub-band spectrum entropy value of each frame of the processed speech signal. It can be seen that the detection control unit accurately detects the start and end points of the speech.
[0102] Figure 7The time-domain waveform of the speech signal output by the noise interference suppression unit and the corresponding time-frequency spectrogram are shown in FIG. 6. It can be seen that the noise interference suppression unit effectively suppresses the environmental noise interference and outputs a relatively pure and undistorted speech signal.
[0103] Figure 8 In FIG. 5, (a) is the time-domain waveform of the speech signal with noise input into the detection control unit, i.e., the speech signal waveform of the mouth microphone channel, and (b) is the corresponding spectrogram obtained by calculating the sub-band spectral entropy value of each frame of the speech signal. It can be seen that the detection control unit accurately detects the start and end points and the termination point of the body speech.
[0104] Figure 9 The time-domain waveform of the speech signal output by the body speech interference suppression unit and the corresponding time-frequency spectrogram are shown in FIG. 7. It can be seen that the body speech interference suppression unit effectively suppresses the body speech signal and outputs a relatively pure and undistorted speech signal of the opposite party.
Claims
1. A method for enhancing the voice of another person during face-to-face conversations using protective earmuffs in high-noise workshops, characterized in that, Includes the following steps: Step 1: Use a microphone array consisting of a left ear microphone, a left ear microphone, and a mouth microphone to synchronously collect noisy speech signals of the subject and the person speaking to the subject during face-to-face conversation in a high-noise environment, and preprocess the collected signals. Step 2: Using the speech signals from the left ear microphone channel and the right ear microphone channel obtained in Step 1, perform fixed beamforming to obtain the main channel and auxiliary channel input signals of the first-level noise interference suppression unit; Step 3: Using the main channel speech signal obtained in Step 2, perform speech endpoint detection to determine whether the current period is a speech period or a speechless period; Step 4: Using the main channel and auxiliary channel speech signals obtained in Step 2 and the decision information obtained in Step 3, the filter coefficients of the first-level noise interference suppression unit are adaptively adjusted to complete the cancellation of environmental noise interference. Step 5: Using the speech signal from the mouth-side microphone channel obtained in Step 1, perform body speech endpoint detection to determine whether the current period is a body speech period or a body speech period. Step 6: Using the speech signal from the mouth-side microphone channel obtained in Step 1, the noise-suppressed speech signal obtained in Step 4, and the decision information obtained in Step 5, the noise-suppressed speech signal is used as the main channel input signal of the second-level body speech interference suppression unit, and the speech signal from the mouth-side microphone channel is used as the auxiliary channel input signal. According to the decision information, the filter coefficients are adaptively adjusted to complete the cancellation of body speech interference, obtain the enhanced signal of the other party's speech, and broadcast it through the speaker built into the protective earmuff.
2. The method for enhancing the other party's voice during face-to-face conversations using protective earmuffs in high-noise workshops, as described in claim 1, is characterized in that... The specific process of step 1 is as follows: Step 1-1: Number the left ear microphone, left ear microphone, and mouth microphone as 1, 2, and 3 respectively. Microphones 1 and 2 are located on the left and right earcups of the protective earcups, and microphone 3 is located on the extended mouth microphone pole connected to the protective earcups, ensuring that microphones 1 and 2 are symmetrically distributed on both sides of microphone 3. Step 1-2: Using the array setup from Step 1-1, the noisy speech signals from each channel are pre-emphasized using a first-order FIR high-pass digital filter. The transfer function of this filter is... H(z)=1-z -1 0.9<a<1.0 In the formula, a is the pre-weighting coefficient, and z -1 This is the unit delay operator in the z-transform; Steps 1-3: Perform short-time framing and windowing processing on the noisy speech signals of each channel obtained in Step 1-2, where the frame shift is half the frame length and the window function is a Hamming window.
3. The method for enhancing the other party's voice during face-to-face conversations using protective earmuffs in high-noise workshops, as described in claim 2, is characterized in that... The specific process of step 2 is as follows: Using the preprocessed signals from the left and right ear microphone channels obtained by the method described in step 1, fixed beamforming processing is performed; the output y(n) of the fixed beamformer is defined as: In the formula, M represents the number of elements in the microphone array, and x i (n) represents the speech signal acquired by the i-th array element, where i = 0, 1, ..., M-1, w i τ represents the weight of the i-th filter. i This represents the time delay estimate between the i-th array element and the relative reference array element.
4. The method for enhancing the other party's voice during face-to-face conversations using protective earmuffs in high-noise workshops, as described in claim 3, is characterized in that... The specific process of step 3 is as follows: Step 3-1: Using the main channel signal obtained by the method described in Step 2, perform noise reduction processing using an improved spectral subtraction method based on multi-window spectral estimation, specifically as follows: Step 3-1-1: Perform FFT on the current frame signal and calculate its amplitude spectrum |X i (k)| and phase spectrum θ i (k) Smoothing is performed between adjacent frames. Specifically, M frames are taken before and after frame i, for a total of 2M+1 frames, and the average amplitude spectrum is calculated as follows: Step 3-1-2: Perform multi-window spectrum estimation on the current frame signal, calculate its multi-window power spectral density P(k,i), where i represents the i-th frame and k represents the k-th spectral line. Perform smoothing between adjacent frames, specifically taking M frames before and after the i-th frame as the center, for a total of 2M+1 frames, and averaging the results to calculate the smoothed power spectral density: Step 3-1-3: Based on the noisy frame with no preceding segment, calculate the average power spectral density of the noise as follows: In the formula, NIS represents the number of frames occupied by the preamble without a voice segment, and P... y (k,i) is the smoothed power spectral density of the i-th frame obtained by the method described in step 3-1-2; Step 3-1-4: Using the average amplitude spectrum obtained by the method described in Step 3-1-1, the smoothed power spectral density obtained by the method described in Step 3-1-2, and the noise average power spectral density obtained by the method described in Step 3-1-4, calculate the gain factor using the spectral subtraction relationship. The gain factor is defined as: In the formula, α is the over-subtraction factor, β is the gain compensation factor, and P y (k,i) is the smoothed power spectral density of the i-th frame obtained by the method described in step 3-1-2, P n (k) is the average power spectral density value of the noise frame with no preceding segment obtained by the method described in step 3-1-2. Step 3-1-5: Using the average amplitude spectrum obtained by the method described in Step 3-1-1 and the gain factor obtained by the method described in Step 3-1-4, the noise-reduced speech signal is obtained: Step 3-2: Using the noise-reduced speech signal obtained by the method described in Step 3-1, calculate the short-time subband spectral entropy of each frame of the signal, specifically as follows: Step 3-2-1: Divide the current frame signal into N parts. b1 For each sub-band, calculate the sub-band energy of the m-th sub-band in the i-th frame, and obtain the sub-band energy E of the m-th sub-band. b1 (m,i), 1≤m≤N b1 ; Step 3-2-2: Using the subband energy obtained by the method described in Step 3-2-1, calculate the probability of the subband energy of the m-th subband in the i-th frame as follows: Step 3-2-3: Using the subband energy probability obtained by the method described in Step 3-2-2, calculate the subband spectral entropy of the i-th frame as follows: Step 3-3: Using the short-time subband spectral entropy obtained by the method described in Step 3-2, set two thresholds and use the dual-threshold method for speech endpoint detection, specifically as follows: Step 3-3-1: Set a low threshold T2 and perform a coarse judgment. If the threshold is below T2, it is speech. The start and end points of the speech are outside the time points corresponding to the intersection of the threshold and the sub-band spectral entropy envelope. Step 3-3-2: Set a higher threshold T1, and search to the left and right respectively from the coarsely judged speech start and end points obtained by the method described in step 3-3-1, find two points where the subband spectral entropy envelope intersects with the threshold T1, and obtain the speech start and end point positions determined by the dual threshold method.
5. A method for enhancing the voice of another person during face-to-face conversations using protective earmuffs in high-noise workshops, as described in claim 4, is characterized in that... The specific process of step 4 is as follows: Step 4-1: Using the speech endpoint detection results obtained by the method described in Step 3, determine whether the current frame is a speech frame or a non-speech frame; Step 4-2: Using the main channel and auxiliary channel signals obtained by the method described in Step 2 and the judgment information obtained by the method described in Step 4-1, perform the first-level noise interference suppression processing, specifically as follows: Step 4-2-1: Filter the auxiliary channel signal. The filter output is: y1(n) NLMS =W1 T (n)X1(n) In the formula, W1(n) is the filter coefficient at time n, and X1(n) is the auxiliary channel input signal at time n. Step 4-2-2: Subtract the main channel signal from the filter output signal to obtain the output signal after noise interference suppression processing; Step 4-2-3: If the current frame is a non-speech frame, then update the filter coefficients using the following formula: In the formula, μ1 is the step size factor, e1(n) is the signal obtained by subtracting the main channel signal from the filter output signal, and c is a very small constant.
6. The method for enhancing the other party's voice during face-to-face conversations using protective earmuffs in high-noise workshops, as described in claim 5, is characterized in that: The specific process of step 5 is as follows: Step 5-1: Using the speech signal from the mouth-side microphone channel obtained by the method described in Step 1, calculate the short-time subband spectral entropy of each frame of the signal, specifically as follows: Step 5-1-1: Divide the current frame signal into N parts. b2 For each sub-band, calculate the sub-band energy of the m-th sub-band in the i-th frame, and obtain the sub-band energy E of the m-th sub-band. b2 (m,i), 1≤m≤N b2 ; Step 5-1-2: Using the subband energy obtained by the method described in Step 5-1-1, calculate the probability of the subband energy of the m-th subband in the i-th frame: Step 5-1-3: Using the subband energy probability obtained by the method described in Step 5-1-2, calculate the subband spectral entropy of the i-th frame as follows: Step 5-2: Using the short-time subband spectral entropy obtained by the method described in Step 5-1, set two thresholds and use the dual-threshold method for speech endpoint detection, specifically: Step 5-2-1: Set a low threshold T2′ and perform a coarse judgment. If the value is below the T2′ threshold, it is definitely the main speech. The start and end points of the main speech are outside the time points corresponding to the intersection of the threshold and the sub-band spectral entropy envelope. Step 5-2-2: Set a higher threshold T1′, and search to the left and right respectively from the coarsely judged speech start and end points obtained by the method described in Step 5-2-1, find the two points where the subband spectral entropy envelope intersects with the threshold T1′, and obtain the position of the speech start and end points determined by the dual threshold method.
7. A method for enhancing the voice of another person during face-to-face conversation using protective earmuffs in a high-noise workshop, as described in claim 6, is characterized in that... The specific process of step 6 is as follows: Step 6-1: Using the speech endpoint detection results obtained by the method described in Step 5, determine whether the current frame is a ontological speech frame or a non-ontological speech frame. Step 6-2: Using the speech signal from the mouth-side microphone channel obtained by the method described in Step 1, the noise-suppressed speech signal obtained by the method described in Step 4, and the judgment information obtained by the method described in Step 6-1, the noise-suppressed speech signal is used as the main channel input signal of the second-level body speech interference suppression unit, and the speech signal from the mouth-side microphone channel is used as the auxiliary channel input signal to perform the second-level body speech interference suppression processing, specifically: Step 6-2-1: Filter the auxiliary channel signal. The filter output is: y2(n) NLMS =W2 T (n)X2(n) In the formula, W2(n) are the filter coefficients at time n, and X2(n) are the auxiliary channel input signals at time n. Step 6-2-2: Subtract the main channel signal from the filter output signal to obtain the output signal after noise interference suppression processing; Step 6-2-3: If the current frame is a speech frame, then update the filter coefficients using the following formula: In the formula, e2(n) is the signal obtained by subtracting the main channel signal from the filter output signal, and c is a very small constant.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-6.
Citation Information
Patent Citations
Voice active detection method
CN110047470A
Wind noise suppression method and device for double microphones, equipment and storage medium
CN113270106A