Voice activity detection method and device based on threshold value self-adaption
By using threshold adaptive voice activity detection method in open office scenarios, the problem of background noise affecting user experience is solved, and fast and accurate voice and noise detection is achieved, improving detection accuracy and noise tracking capabilities.
Patent Information
- Application Number
- CN202510093934.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
In open office scenarios, background noise seriously affects the user experience, and existing noise reduction algorithms are difficult to accurately estimate the noise spectrum, resulting in oversuppression of speech or excessive noise residues.
The voice activity detection method based on threshold adaptation is adopted, and the original multi-channel signal picked up by the array microphone is beamformed to obtain the voice beam and noise beam, long-term energy difference and short-term energy difference are calculated, and two voice activity detections are performed to update the judgment threshold in real time.
It realizes rapid and accurate detection of speech and noise in complex environments, can judge transient noise, improves the accuracy of speech activity detection, and enhances noise tracking capabilities.
Smart Images

Figure CN119943079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice activity recognition, and in particular to a voice activity detection method and device based on threshold adaptation. Background Art
[0002] Currently, in open office scenarios, headset users are often surrounded by a mixture of various types and levels of background noise, which seriously affects the user experience. Noise reduction algorithms often need to have an accurate estimate of the noise spectrum, otherwise the voice will be over-suppressed and blurred, or too much noise will remain. Therefore, in order to obtain accurate noise estimation, a fast-response, high-accuracy voice activity detection method is often required. Summary of the invention
[0003] The present invention provides a voice activity detection method and device based on threshold adaptation, which can update the judgment threshold in real time, quickly and accurately detect voice and noise in a complex environment, realize faster noise tracking, and improve the accuracy of voice activity detection.
[0004] In order to solve the above technical problems, the present invention provides a voice activity detection method based on threshold adaptation, comprising:
[0005] Perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain speech beams and noise beams;
[0006] Calculating the long-time energy difference and the short-time energy difference at each moment respectively according to the speech beam and the noise beam;
[0007] Performing a first voice activity detection according to the long-term energy difference at each moment and a preset first threshold;
[0008] If the first voice activity detection result is a signal to be detected, a second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.
[0009] Furthermore, the beamforming process is performed on the original multi-channel signal picked up by the array microphone to obtain the speech beam and the noise beam, specifically:
[0010] Perform super-directional beamforming on the original multi-channel signal picked up by the array microphone to obtain the speech beam;
[0011] The power spectrum of the human voice beam at each moment is obtained in the speech beam, specifically:
[0012] Vbe(t)=abs(Vb(t))^2*(1-ratio)+abs(Vbe(t-1))^2*(ratio)
[0013] Among them, Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Vb(t) is the amplitude of the human voice beam at time t; Vbe(t-1) is the power spectrum of the human voice beam after smoothing in the time domain at time t-1; ratio is the smoothing coefficient;
[0014] Perform null-notch beamforming on the original multi-channel signal picked up by the array microphone to obtain a noise beam;
[0015] The noise beam power spectrum at each moment is obtained in the noise beam, specifically:
[0016] Nbe(t)=abs(Nb(t))^2*(1-ratio)+abs(Nbe(t-1))^2*(ratio)
[0017] Among them, Nbe(t) is the noise beam power spectrum after smoothing in the time domain at time t; Nb(t) is the noise beam amplitude at time t; Nbe(t-1) is the noise beam power spectrum after smoothing in the time domain at time t-1; ratio is the smoothing coefficient.
[0018] Further, the long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the speech beam and the noise beam, specifically:
[0019] According to the human voice beam power spectrum at each moment and the noise beam power spectrum at each moment, the energy difference at each moment is calculated respectively, specifically:
[0020] Ed(t)=10*log10(Vbe(t))-10*log10(Nbe(t))
[0021] Wherein, Ed(t) is the energy difference at time t; Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Nbe(t) is the power spectrum of the noise beam after smoothing in the time domain at time t;
[0022] When the smoothing coefficient is set to a first smoothing coefficient, determining that the energy difference is a long-term energy difference;
[0023] When the smoothing coefficient is set to a second smoothing coefficient, the energy difference is determined to be a short-time energy difference; wherein the first smoothing coefficient is greater than the second smoothing coefficient.
[0024] Further, the first voice activity detection is performed according to the long-term energy difference at each moment and a preset first threshold, specifically:
[0025] Comparing the long-term energy difference at each moment with a preset first threshold;
[0026] Determine the signal corresponding to the moment when the long-term energy difference is less than the first threshold as noise;
[0027] The signal corresponding to the moment when the long-term energy difference is greater than the first threshold is determined as the signal to be detected.
[0028] Further, the performing the second voice activity detection in each scanning window according to the short-time energy difference at each moment is specifically:
[0029] Setting the upper and lower limits of the peak value of the first scanning window and the length of the first scanning window;
[0030] Comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window;
[0031] According to a preset attenuation amount, a peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as a peak initial value of a next scanning window of the first scanning window;
[0032] Adaptively updating the second threshold in the second voice activity detection according to a preset ratio value;
[0033] The second threshold is used to perform a second voice activity detection on the signal at each moment in the first scanning window.
[0034] Further, comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window, is specifically:
[0035] Determining whether the short-time energy difference at each moment is within the range of the upper and lower limits of the peak value of the first scanning window;
[0036] If the short-time energy difference at the first moment is within the range of the upper and lower limits of the peak value of the first scanning window, the peak value of the first scanning window is updated according to the short-time energy difference at the first moment and the peak value of the first scanning window, specifically:
[0037] Ed_fast_peak(t)=max(Ed_fast_peak(t-1),Ed_fast(t))
[0038] Wherein, Ed_fast_peak(t) is the peak value of the first scanning window at time t; Ed_fast_peak(t-1) is the peak value of the first scanning window at time t-1; Ed_fast(t) is the short-time energy difference at time t.
[0039] Further, the peak attenuation value of the first scanning window is obtained according to the preset attenuation amount, and the peak attenuation value of the first scanning window is used as the peak initial value of the next scanning window of the first scanning window, specifically:
[0040] The value obtained by subtracting the preset attenuation from the peak value of the first scanning window is compared with the peak lower limit of the first scanning window to determine the peak attenuation value of the first scanning window, specifically:
[0041] Ed_fast_peak_dec(t)=max(Ed_fast_peak(t)-dec,peak_lower)
[0042] Wherein, Ed_fast_peak_dec(t) is the peak attenuation value of the first scanning window; Ed_fast_peak(t) is the peak value of the first scanning window; dec is the preset attenuation amount; peak_lower is the peak lower limit of the first scanning window.
[0043] Furthermore, the second threshold in the second voice activity detection is adaptively updated according to a preset ratio value, specifically:
[0044] According to the peak value of the first scanning window and the preset ratio value, the second threshold in the second voice activity detection is adaptively updated, specifically:
[0045] threshold2(t)=Ed_fast_peak(t)*prop
[0046] Wherein, threshold2(t) is the second threshold; Ed_fast_peak(t) is the peak value of the first scanning window; prop is a preset ratio value.
[0047] Further, the using the second threshold to perform the second voice activity detection on the signal at each moment in the first scanning window is specifically:
[0048] Comparing the short-time energy difference at each moment with the second threshold;
[0049] Determine the signal corresponding to the moment when the short-time energy difference is greater than the second threshold as a voice activity;
[0050] The signal corresponding to the moment when the short-time energy difference is less than the second threshold is determined as noise.
[0051] The present invention provides a voice activity detection method based on threshold adaptation, which performs beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a voice beam and a noise beam; respectively calculates the long-time energy difference and the short-time energy difference at each moment according to the voice beam and the noise beam; performs a first voice activity detection according to the long-time energy difference at each moment and a first threshold; if the first voice activity detection result is a signal to be detected, performs a second voice activity detection in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment. The present invention can update the judgment threshold in real time, quickly and accurately detect voice and noise in a complex environment, can judge short-term noise, so that noise and voice are separated more thoroughly, realize faster noise tracking, and improve the accuracy of voice activity detection.
[0052] Accordingly, the present invention provides a voice activity detection device based on threshold adaptation, comprising: a beamforming module, a calculation module, a first detection module and a second detection module;
[0053] The beamforming module is used to perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam;
[0054] The calculation module is used to calculate the long-time energy difference and the short-time energy difference at each moment according to the speech beam and the noise beam;
[0055] The first detection module is used to perform a first voice activity detection according to the long-term energy difference at each moment and a preset first threshold;
[0056] The second detection module is used to perform a second voice activity detection in each scanning window according to the short-time energy difference at each moment if the first voice activity detection result is a signal to be detected; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.
[0057] The present invention provides a voice activity detection device based on threshold adaptation. Based on the organic combination of modules, the original multi-channel signal picked up by the array microphone is processed by beamforming to obtain a voice beam and a noise beam; the long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the voice beam and the noise beam; the first voice activity detection is performed according to the long-time energy difference at each moment and the first threshold; if the first voice activity detection result is a signal to be detected, the second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment. The present invention can update the judgment threshold in real time, quickly and accurately detect voice and noise in a complex environment, can judge short-term noise, so that noise and voice are separated more thoroughly, realize faster noise tracking, and improve the accuracy of voice activity detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A schematic flow chart of an embodiment of a voice activity detection method based on threshold adaptation provided by the present invention;
[0059] Figure 2 A schematic flow chart of another embodiment of a voice activity detection method based on threshold adaptation provided by the present invention;
[0060] Figure 3 A schematic structural diagram of an embodiment of a voice activity detection device based on threshold adaptation provided by the present invention. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] Example 1
[0063] See also Figure 1 , is a flow chart of an embodiment of a method for voice activity detection based on threshold adaptation provided by the present invention, the method comprises steps 101 to 104, each step is specifically as follows:
[0064] Step 101: Perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam.
[0065] Furthermore, in the first embodiment of the present invention, beamforming processing is performed on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam, specifically:
[0066] Perform super-directional beamforming on the original multi-channel signal picked up by the array microphone to obtain the speech beam;
[0067] The power spectrum of the human voice beam at each moment is obtained in the speech beam, specifically:
[0068] Vbe(t)=abs(Vb(t))^2*(1-ratio)+abs(Vbe(t-1))^2*(ratio)
[0069] Among them, Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Vb(t) is the amplitude of the human voice beam at time t; Vbe(t-1) is the power spectrum of the human voice beam after smoothing in the time domain at time t-1; ratio is the smoothing coefficient;
[0070] Perform null-notch beamforming on the original multi-channel signal picked up by the array microphone to obtain a noise beam;
[0071] The noise beam power spectrum at each moment is obtained in the noise beam, specifically:
[0072] Nbe(t)=abs(Nb(t))^2*(1-ratio)+abs(Nbe(t-1))^2*(ratio)
[0073] Among them, Nbe(t) is the noise beam power spectrum after smoothing in the time domain at time t; Nb(t) is the noise beam amplitude at time t; Nbe(t-1) is the noise beam power spectrum after smoothing in the time domain at time t-1; ratio is the smoothing coefficient.
[0074] In the first embodiment of the present invention, the array microphone is arranged at the front end of the microphone. In practical applications, the noise source received by the microphone includes environmental noise and surrounding interfering human voices. Beamforming technology refers to a spatial filter based on the direction of sound arrival, wherein the super-directional beamforming technology has a speech enhancement effect on the target direction, and the beam nulling technology has a blocking effect on the target direction. Therefore, the original multi-channel signal picked up by the array microphone is subjected to super-directional beamforming and nulling beamforming in the target direction (the direction of the human mouth relative to the array microphone), respectively, and a speech beam and a noise beam can be obtained respectively. The acquired speech beam mainly contains the speech in the target direction and the residual noise after suppression, and the noise beam mainly contains the surrounding noise and the speech in the target direction after blocking.
[0075] Step 102: Calculate the long-time energy difference and the short-time energy difference at each moment according to the speech beam and the noise beam respectively.
[0076] Further, in the first embodiment of the present invention, the long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the speech beam and the noise beam, specifically:
[0077] According to the human voice beam power spectrum at each moment and the noise beam power spectrum at each moment, the energy difference at each moment is calculated respectively, specifically:
[0078] Ed(t)=10*log10(Vbe(t))-10*log10(Nbe(t))
[0079] Wherein, Ed(t) is the energy difference at time t; Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Nbe(t) is the power spectrum of the noise beam after smoothing in the time domain at time t;
[0080] When the smoothing coefficient is set to a first smoothing coefficient, determining that the energy difference is a long-term energy difference;
[0081] When the smoothing coefficient is set to a second smoothing coefficient, the energy difference is determined to be a short-time energy difference; wherein the first smoothing coefficient is greater than the second smoothing coefficient.
[0082] In the first embodiment of the present invention, the energy difference can be calculated using the speech beam and the noise beam. A preset smoothing coefficient is required in the formula for calculating the energy difference. When the smoothing coefficient is large, the energy difference at this time is the long-term energy. Since the long-term energy difference has a large smoothing coefficient and a slow update speed in time, it can reflect the overall change of the signal energy difference. When the smoothing coefficient is small, the energy difference at this time is the short-term energy difference. Since the short-term energy difference has a small smoothing coefficient and a fast update speed in time, it can capture the detailed changes of the local energy difference of the signal.
[0083] Step 103: Perform first voice activity detection according to the long-term energy difference at each moment and a preset first threshold.
[0084] Further, in the first embodiment of the present invention, the first voice activity detection is performed according to the long-term energy difference at each moment and a preset first threshold, specifically:
[0085] Comparing the long-term energy difference at each moment with a preset first threshold;
[0086] Determine the signal corresponding to the moment when the long-term energy difference is less than the first threshold as noise;
[0087] The signal corresponding to the moment when the long-term energy difference is greater than the first threshold is determined as the signal to be detected.
[0088] In the first embodiment of the present invention, a first voice activity detection can be performed using a long-term energy difference and a fixed first threshold. When the long-term energy difference is less than the first threshold, it can be determined that the signal at this time is a noise segment and there is no long-term voice activity. However, this method cannot accurately judge the short-term noise between voices, so it can only first determine that the signal with a long-term energy difference greater than the first threshold is the signal to be detected, and cannot determine all of them as voice activities.
[0089] Step 104: If the first voice activity detection result is a signal to be detected, a second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.
[0090] Further, in the first embodiment of the present invention, the second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment, specifically:
[0091] Setting the upper and lower limits of the peak value of the first scanning window and the length of the first scanning window;
[0092] Comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window;
[0093] According to a preset attenuation amount, a peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as a peak initial value of a next scanning window of the first scanning window;
[0094] Adaptively updating the second threshold in the second voice activity detection according to a preset ratio value;
[0095] The second threshold is used to perform a second voice activity detection on the signal at each moment in the first scanning window.
[0096] In the first embodiment of the present invention, based on the judgment of the first voice activity detection, the short-time energy difference and the adaptive dynamic threshold can be used to track and judge the short-term noise in the signal to be detected. In the second voice activity detection, each time frame in each scanning window needs to be detected. Specifically, the upper limit value, lower limit value of the scanning window peak value and the length of the scanning window are preset according to the hardware characteristics of the headphone device. In each time window, it is necessary to judge and update the peak value of the current time frame, judge and update the peak attenuation value of the current time frame, judge and update the second threshold value of the current time frame, and use the second threshold value to judge the voice activity of the current time frame. After completing the voice activity judgment of the previous scanning window, move the scanning window, forget the peak value in the previous scanning window in the new scanning window, and use the peak attenuation value in the previous scanning window to initially update the peak value of the new scanning window, and then repeat the above-mentioned steps of the second voice activity detection.
[0097] Further, in the first embodiment of the present invention, the short-time energy difference at each moment in the first scanning window is compared with the peak value of the first scanning window, and the peak value of the first scanning window is updated, specifically:
[0098] Determining whether the short-time energy difference at each moment is within the range of the upper and lower limits of the peak value of the first scanning window;
[0099] If the short-time energy difference at the first moment is within the range of the upper and lower limits of the peak value of the first scanning window, the peak value of the first scanning window is updated according to the short-time energy difference at the first moment and the peak value of the first scanning window, specifically:
[0100] Ed_fast_peak(t)=max(Ed_fast_peak(t-1),Ed_fast(t))
[0101] Wherein, Ed_fast_peak(t) is the peak value of the first scanning window at time t; Ed_fast_peak(t-1) is the peak value of the first scanning window at time t-1; Ed_fast(t) is the short-time energy difference at time t.
[0102] In the first embodiment of the present invention, a peak value that is too large will cause the subsequent threshold value to be updated too slowly, and the judgment of speech will not be too detailed; a peak value that is too small will cause the threshold value of the subsequent update to be too small, and unnecessary noise will be misjudged as speech. Therefore, in order to avoid capturing inappropriate peaks, it is first necessary to confirm whether the short-time energy difference value at the current moment is within the upper and lower limits of the peak value. If the short-time energy difference value at the current moment is within the upper and lower limits of the peak value, it is compared with the current peak value, and the larger value of the two is used as the peak value within the current scanning window moment to complete the peak value update.
[0103] Further, in the first embodiment of the present invention, according to the preset attenuation amount, the peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as the peak initial value of the next scanning window of the first scanning window, specifically:
[0104] The value obtained by subtracting the preset attenuation from the peak value of the first scanning window is compared with the peak lower limit of the first scanning window to determine the peak attenuation value of the first scanning window, specifically:
[0105] Ed_fast_peak_dec(t)=max(Ed_fast_peak(t)-dec,peak_lower)
[0106] Wherein, Ed_fast_peak_dec(t) is the peak attenuation value of the first scanning window; Ed_fast_peak(t) is the peak value of the first scanning window; dec is the preset attenuation amount; peak_lower is the peak lower limit of the first scanning window.
[0107] In the first embodiment of the present invention, the preset attenuation is subtracted from the peak value at the current moment, and the attenuation is compared with the peak lower limit, and the larger value is used as the peak attenuation value of the current scanning window. The attenuation is preset based on experience, and the attenuation is the same at each moment.
[0108] Further, in the first embodiment of the present invention, the second threshold in the second voice activity detection is adaptively updated according to a preset ratio value, specifically:
[0109] According to the peak value of the first scanning window and the preset ratio value, the second threshold in the second voice activity detection is adaptively updated, specifically:
[0110] threshold2(t)=Ed_fast_peak(t)*prop
[0111] Wherein, threshold2(t) is the second threshold; Ed_fast_peak(t) is the peak value of the first scanning window; prop is a preset ratio value.
[0112] In the first embodiment of the present invention, an adaptive update of a second threshold in the second voice activity detection is performed, and the second threshold is obtained by multiplying a peak value of the first scanning window and a proportional value, wherein the proportional value is preset based on experience, and the attenuation at each moment is the same.
[0113] Further, in the first embodiment of the present invention, the second voice activity detection is performed on the signal at each moment in the first scanning window using the second threshold, specifically:
[0114] Comparing the short-time energy difference at each moment with the second threshold;
[0115] Determine the signal corresponding to the moment when the short-time energy difference is greater than the second threshold as a voice activity;
[0116] The signal corresponding to the moment when the short-time energy difference is less than the second threshold is determined as noise.
[0117] In the first embodiment of the present invention, the second voice activity detection can be performed using the short-time energy difference and the second threshold. If the short-time energy difference is greater than the second threshold, the signal at this time is determined to be voice activity, and if the short-time energy difference is less than the second threshold, the signal at this time is determined to be noise. The second voice activity detection is updated more quickly, and can accurately judge the voice word by word in a complex noise background, with higher accuracy.
[0118] In the first embodiment of the present invention, after completing the voice activity judgment of the previous scanning window, the scanning window is moved. When moving to a new scanning window, the peak value in the previous scanning window will be forgotten, and the peak attenuation value of the previous scanning window will be saved in the new scanning window, and a new round of peak value comparison and update will be performed, specifically:
[0119] First, confirm whether the short-term energy difference at the current moment is within the upper and lower limits of the peak value;
[0120] Then the peak value of the new scanning window is updated, the peak attenuation value of the previous scanning window is compared with the short-time energy difference of the new scanning window, and the larger value of the two is taken as the peak value of the new scanning window;
[0121] Then, in the new scanning window, the same peak attenuation, threshold adaptive updating and voice activity judgment process as the previous scanning window are completed.
[0122] As an example of the present invention, see Figure 2 , is a flow chart of another embodiment of the voice activity detection method based on threshold adaptation provided by the present invention. The original signal received by the array microphone is subjected to target direction super-pointing beamforming and target direction nulling beamforming in turn, and a voice beam and a noise beam can be obtained respectively, and then the long-time energy difference and the short-time energy difference at each moment are calculated. The long-time energy difference can be used to make the first voice activity judgment, and the noise segment without long-time voice activity can be identified. However, at this time, it is still impossible to accurately judge the short-term noise between voices. Therefore, it is necessary to make a second voice activity judgment on the basis of the first voice activity judgment to complete the judgment of the short-term noise between voices. The second voice activity judgment includes: performing window scanning, peak capture of the short-time energy difference, attenuation and updating of the peak, and real-time adjustment of the threshold of the second voice activity judgment, so as to make the second voice activity judgment.
[0123] In summary, the first embodiment of the present invention provides a voice activity detection method based on threshold adaptation, which performs beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a voice beam and a noise beam; the long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the voice beam and the noise beam; the first voice activity detection is performed according to the long-time energy difference at each moment and the first threshold; if the first voice activity detection result is a signal to be detected, the second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment. The present invention can update the judgment threshold in real time, quickly and accurately detect voice and noise in a complex environment, can judge short-term noise, so that noise and voice are separated more thoroughly, realize faster noise tracking, and improve the accuracy of voice activity detection.
[0124] Example 2
[0125] See also Figure 3 , is a schematic structural diagram of an embodiment of a voice activity detection device based on threshold adaptation provided by the present invention, the device comprises a beamforming module 201, a calculation module 202, a first detection module 203 and a second detection module 204;
[0126] The beamforming module 201 is used to perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam;
[0127] The calculation module 202 is used to calculate the long-time energy difference and the short-time energy difference at each moment according to the speech beam and the noise beam;
[0128] The first detection module 203 is used to perform a first voice activity detection according to the long-term energy difference at each moment and a preset first threshold;
[0129] The second detection module 204 is used to perform a second voice activity detection in each scanning window according to the short-time energy difference at each moment if the first voice activity detection result is a signal to be detected; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.
[0130] Furthermore, in the second embodiment of the present invention, beamforming processing is performed on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam, specifically:
[0131] Perform super-directional beamforming on the original multi-channel signal picked up by the array microphone to obtain the speech beam;
[0132] The power spectrum of the human voice beam at each moment is obtained in the speech beam, specifically:
[0133] Vbe(t)=abs(Vb(t))^2*(1-ratio)+abs(Vbe(t-1))^2*(ratio)
[0134] Among them, Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Vb(t) is the amplitude of the human voice beam at time t; Vbe(t-1) is the power spectrum of the human voice beam after smoothing in the time domain at time t-1; ratio is the smoothing coefficient;
[0135] Perform null-notch beamforming on the original multi-channel signal picked up by the array microphone to obtain a noise beam;
[0136] The noise beam power spectrum at each moment is obtained in the noise beam, specifically:
[0137] Nbe(t)=abs(Nb(t))^2*(1-ratio)+abs(Nbe(t-1))^2*(ratio)
[0138] Among them, Nbe(t) is the noise beam power spectrum after smoothing in the time domain at time t; Nb(t) is the noise beam amplitude at time t; Nbe(t-1) is the noise beam power spectrum after smoothing in the time domain at time t-1; ratio is the smoothing coefficient.
[0139] Further, in the second embodiment of the present invention, the long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the speech beam and the noise beam, specifically:
[0140] According to the human voice beam power spectrum at each moment and the noise beam power spectrum at each moment, the energy difference at each moment is calculated respectively, specifically:
[0141] Ed(t)=10*log10(Vbe(t))-10*log10(Nbe(t))
[0142] Wherein, Ed(t) is the energy difference at time t; Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Nbe(t) is the power spectrum of the noise beam after smoothing in the time domain at time t;
[0143] When the smoothing coefficient is set to a first smoothing coefficient, determining that the energy difference is a long-term energy difference;
[0144] When the smoothing coefficient is set to a second smoothing coefficient, the energy difference is determined to be a short-time energy difference; wherein the first smoothing coefficient is greater than the second smoothing coefficient.
[0145] Further, in the second embodiment of the present invention, the first voice activity detection is performed according to the long-term energy difference at each moment and a preset first threshold, specifically:
[0146] Comparing the long-term energy difference at each moment with a preset first threshold;
[0147] Determine the signal corresponding to the moment when the long-term energy difference is less than the first threshold as noise;
[0148] The signal corresponding to the moment when the long-term energy difference is greater than the first threshold is determined as the signal to be detected.
[0149] Further, in the second embodiment of the present invention, the second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment, specifically:
[0150] Setting the upper and lower limits of the peak value of the first scanning window and the length of the first scanning window;
[0151] Comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window;
[0152] According to a preset attenuation amount, a peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as a peak initial value of a next scanning window of the first scanning window;
[0153] Adaptively updating the second threshold in the second voice activity detection according to a preset ratio value;
[0154] The second threshold is used to perform a second voice activity detection on the signal at each moment in the first scanning window.
[0155] Further, in the second embodiment of the present invention, the short-time energy difference at each moment in the first scanning window is compared with the peak value of the first scanning window, and the peak value of the first scanning window is updated, specifically:
[0156] Determining whether the short-time energy difference at each moment is within the range of the upper and lower limits of the peak value of the first scanning window;
[0157] If the short-time energy difference at the first moment is within the range of the upper and lower limits of the peak value of the first scanning window, the peak value of the first scanning window is updated according to the short-time energy difference at the first moment and the peak value of the first scanning window, specifically:
[0158] Ed_fast_peak(t)=max(Ed_fast_peak(t-1),Ed_fast(t))
[0159] Wherein, Ed_fast_peak(t) is the peak value of the first scanning window at time t; Ed_fast_peak(t-1) is the peak value of the first scanning window at time t-1; Ed_fast(t) is the short-time energy difference at time t.
[0160] Further, in the second embodiment of the present invention, according to the preset attenuation amount, the peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as the peak initial value of the next scanning window of the first scanning window, specifically:
[0161] The value obtained by subtracting the preset attenuation from the peak value of the first scanning window is compared with the peak lower limit of the first scanning window to determine the peak attenuation value of the first scanning window, specifically:
[0162] Ed_fast_peak_dec(t)=max(Ed_fast_peak(t)-dec,peak_lower)
[0163] Wherein, Ed_fast_peak_dec(t) is the peak attenuation value of the first scanning window; Ed_fast_peak(t) is the peak value of the first scanning window; dec is the preset attenuation amount; peak_lower is the peak lower limit of the first scanning window.
[0164] Further, in the second embodiment of the present invention, the second threshold in the second voice activity detection is adaptively updated according to a preset ratio value, specifically:
[0165] According to the peak value of the first scanning window and the preset ratio value, the second threshold in the second voice activity detection is adaptively updated, specifically:
[0166] threshold2(t)=Ed_fast_peak(t)*prop
[0167] Wherein, threshold2(t) is the second threshold; Ed_fast_peak(t) is the peak value of the first scanning window; prop is a preset ratio value.
[0168] Furthermore, in the second embodiment of the present invention, the second threshold is used to perform second voice activity detection on the signal at each moment in the first scanning window, specifically:
[0169] Comparing the short-time energy difference at each moment with the second threshold;
[0170] Determine the signal corresponding to the moment when the short-time energy difference is greater than the second threshold as a voice activity;
[0171] The signal corresponding to the moment when the short-time energy difference is less than the second threshold is determined as noise.
[0172] In summary, the second embodiment of the present invention provides a voice activity detection device based on threshold adaptation, which performs beamforming processing on the original multi-channel signal picked up by the array microphone based on the organic combination of modules to obtain a voice beam and a noise beam; calculates the long-time energy difference and the short-time energy difference at each moment respectively according to the voice beam and the noise beam; performs a first voice activity detection according to the long-time energy difference at each moment and the first threshold; if the first voice activity detection result is a signal to be detected, performs a second voice activity detection in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment. The present invention can update the judgment threshold in real time, quickly and accurately detect voice and noise in a complex environment, can judge short-term noise, so that noise and voice are separated more thoroughly, realize faster noise tracking, and improve the accuracy of voice activity detection.
[0173] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A voice activity detection method based on threshold adaptation, characterized in that: include: Perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain speech beams and noise beams; Calculating the long-time energy difference and the short-time energy difference at each moment respectively according to the speech beam and the noise beam; Performing a first voice activity detection according to the long-term energy difference at each moment and a preset first threshold; If the first voice activity detection result is a signal to be detected, a second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.
2. The voice activity detection method based on threshold adaptation according to claim 1, characterized in that: The beamforming process is performed on the original multi-channel signal picked up by the array microphone to obtain the speech beam and the noise beam, specifically: Perform super-directional beamforming on the original multi-channel signal picked up by the array microphone to obtain the speech beam; The power spectrum of the human voice beam at each moment is obtained in the speech beam, specifically: Vbe(t)=abs(Vb(t))^2*(1-ratio)+abs(Vbe(t-1))^2*(ratio) Among them, Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Vb(t) is the amplitude of the human voice beam at time t; Vbe(t-1) is the power spectrum of the human voice beam after smoothing in the time domain at time t-1; ratio is the smoothing coefficient; Perform null-notch beamforming on the original multi-channel signal picked up by the array microphone to obtain a noise beam; The noise beam power spectrum at each moment is obtained in the noise beam, specifically: Nbe(t)=abs(Nb(t))^2*(1-ratio)+abs(Nbe(t-1))^2*(ratio) Among them, Nbe(t) is the noise beam power spectrum after smoothing in the time domain at time t; Nb(t) is the noise beam amplitude at time t; Nbe(t-1) is the noise beam power spectrum after smoothing in the time domain at time t-1; ratio is the smoothing coefficient.
3. The voice activity detection method based on threshold adaptation according to claim 2, characterized in that: The long-time energy difference and the short-time energy difference at each moment are calculated respectively according to the speech beam and the noise beam, specifically: According to the human voice beam power spectrum at each moment and the noise beam power spectrum at each moment, the energy difference at each moment is calculated respectively, specifically: Ed(t)=10*log10(Vbe(t))-10*log10(Nbe(t)) Wherein, Ed(t) is the energy difference at time t; Vbe(t) is the power spectrum of the human voice beam after smoothing in the time domain at time t; Nbe(t) is the power spectrum of the noise beam after smoothing in the time domain at time t; When the smoothing coefficient is set to a first smoothing coefficient, determining that the energy difference is a long-term energy difference; When the smoothing coefficient is set to a second smoothing coefficient, the energy difference is determined to be a short-time energy difference; wherein the first smoothing coefficient is greater than the second smoothing coefficient.
4. The voice activity detection method based on threshold adaptation according to claim 1, characterized in that: The first voice activity detection is performed according to the long-term energy difference at each moment and a preset first threshold value, specifically: Comparing the long-term energy difference at each moment with a preset first threshold; Determine the signal corresponding to the moment when the long-term energy difference is less than the first threshold as noise; The signal corresponding to the moment when the long-term energy difference is greater than the first threshold is determined as the signal to be detected.
5. The voice activity detection method based on threshold adaptation according to claim 1, characterized in that: The second voice activity detection is performed in each scanning window according to the short-time energy difference at each moment, specifically: Setting the upper and lower limits of the peak value of the first scanning window and the length of the first scanning window; Comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window; According to a preset attenuation amount, a peak attenuation value of the first scanning window is obtained, and the peak attenuation value of the first scanning window is used as a peak initial value of a next scanning window of the first scanning window; Adaptively updating the second threshold in the second voice activity detection according to a preset ratio value; The second threshold is used to perform a second voice activity detection on the signal at each moment in the first scanning window.
6. The voice activity detection method based on threshold adaptation according to claim 5, characterized in that: The comparing the short-time energy difference at each moment in the first scanning window with the peak value of the first scanning window, and updating the peak value of the first scanning window, is specifically: Determining whether the short-time energy difference at each moment is within the range of the upper and lower limits of the peak value of the first scanning window; If the short-time energy difference at the first moment is within the range of the upper and lower limits of the peak value of the first scanning window, the peak value of the first scanning window is updated according to the short-time energy difference at the first moment and the peak value of the first scanning window, specifically: Ed_fast_peak(t)=max(Ed_fast_peak(t-1),Ed_fast(t)) Wherein, Ed_fast_peak(t) is the peak value of the first scanning window at time t; Ed_fast_peak(t-1) is the peak value of the first scanning window at time t-1; Ed_fast(t) is the short-time energy difference at time t.
7. The voice activity detection method based on threshold adaptation according to claim 6, characterized in that: The step of obtaining the peak attenuation value of the first scanning window according to the preset attenuation amount, and using the peak attenuation value of the first scanning window as the peak initial value of the next scanning window of the first scanning window, is specifically: The value obtained by subtracting the preset attenuation from the peak value of the first scanning window is compared with the peak lower limit of the first scanning window to determine the peak attenuation value of the first scanning window, specifically: Ed_fast_peak_dec(t)=max(Ed_fast_peak(t)-dec,peak_lower) Wherein, Ed_fast_peak_dec(t) is the peak attenuation value of the first scanning window; Ed_fast_peak(t) is the peak value of the first scanning window; dec is the preset attenuation amount; peak_lower is the peak lower limit of the first scanning window.
8. The method for voice activity detection based on threshold adaptation according to claim 7, characterized in that: The step of adaptively updating the second threshold in the second voice activity detection according to the preset ratio value is specifically as follows: According to the peak value of the first scanning window and the preset ratio value, the second threshold in the second voice activity detection is adaptively updated, specifically: threshold2(t)=Ed_fast_peak(t)*prop Wherein, threshold2(t) is the second threshold; Ed_fast_peak(t) is the peak value of the first scanning window; prop is a preset ratio value.
9. The method for voice activity detection based on threshold adaptation according to claim 8, characterized in that: The using the second threshold to perform the second voice activity detection on the signal at each moment in the first scanning window is specifically: Comparing the short-time energy difference at each moment with the second threshold; Determine the signal corresponding to the moment when the short-time energy difference is greater than the second threshold as a voice activity; The signal corresponding to the moment when the short-time energy difference is less than the second threshold is determined as noise.
10. A voice activity detection device based on threshold adaptation, characterized in that: include: A beam forming module, a calculation module, a first detection module and a second detection module; The beamforming module is used to perform beamforming processing on the original multi-channel signal picked up by the array microphone to obtain a speech beam and a noise beam; The calculation module is used to calculate the long-time energy difference and the short-time energy difference at each moment according to the speech beam and the noise beam; The first detection module is used to perform a first voice activity detection according to the long-term energy difference at each moment and a preset first threshold; The second detection module is used to perform a second voice activity detection in each scanning window according to the short-time energy difference at each moment if the first voice activity detection result is a signal to be detected; wherein the second voice activity detection includes peak value update, peak attenuation value update, threshold value update and voice activity judgment.