Background Noise Estimation and Voice Activity Detection System

Through the nonlinear updated background noise estimation method, the problem of the existing technology being difficult to detect voice activity after sudden loud noise is solved, and more efficient voice activity detection is achieved, reducing computing cost and power consumption.

CN114930451BActive Publication Date: 2025-06-13TEXAS INSTRUMENTS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080090845.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-23
Publication Date
2025-06-13
Estimated Expiration
2040-12-23

AI Technical Summary

Technical Problem

Existing voice activity detection systems are difficult to effectively detect voice activity that occurs after sudden loud noise, resulting in high computational costs and increased power consumption.

Method used

By using a nonlinear updated background noise estimation method, the first power spectrum density distribution of the frame is determined by selecting the frame of the audio signal, and an estimate of the background noise in the indicating frame is generated based on the nonlinear weight, thereby determining whether speech activity is detected in the frame.

Benefits of technology

Improves detection accuracy of voice activity, reduces the impact on sudden noise changes, and reduces computational cost and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930451B_ABST
    Figure CN114930451B_ABST
Patent Text Reader

Abstract

A method includes selecting (304) a frame of an audio signal. The method further includes determining (308) a first power spectral density (PSD) distribution of the frame. The method further includes generating (310) a first reference PSD distribution indicative of an estimate of background noise in the frame based on a non-linear weight, a second reference PSD distribution of a previous frame of the audio signal, and a second PSD distribution of the previous frame. The method further includes determining (320) whether speech activity is detected in the frame based on the first PSD distribution of the frame and the first reference PSD distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Speech processing systems are integrated into various electronic devices. For example, many mobile phone devices have virtual assistants that utilize natural language processing systems, which are configured to recognize speech and perform one or more operations based on the recognized speech. Natural language processing is a relatively computationally expensive process. Thus, a mobile phone device (or other device) that performs natural language processing may exhibit increased power consumption and thus have reduced battery life compared to other devices.

[0002] To reduce the computational cost in natural language processing systems, some systems perform a relatively computationally inexpensive voice activity detection process on received sound signals and perform natural language processing on selected portions of the sound signals where voice activity is detected (if any) rather than on the entire sound signal. Some such voice activity detection processes compare samples of the sound signal with an estimate of background noise to determine whether there is voice activity in the sample. The estimate of background noise can be based on historical values associated with the sound signal. However, such systems may fail to detect voice activity that occurs after a sudden loud noise represented in the sound signal. Summary of the Invention

[0003] Systems and methods for detecting voice activity using non-linear updated background noise estimation are described.

[0004] A method includes selecting a frame of an audio signal. The method further includes determining a first power spectral density (PSD) distribution of the frame. The method further includes generating a first reference PSD distribution indicative of an estimate of background noise in the frame based on a non-linear weight, a second reference PSD distribution of a previous frame of the audio signal, and a second PSD distribution of the previous frame. The method further includes determining whether voice activity is detected in the frame based on the first PSD distribution of the frame and the first reference PSD distribution.

[0005] A device includes a processor and a memory storing instructions executable by the processor to select a frame of an audio signal. The instructions are further executable by the processor to determine a first power spectral density (PSD) distribution of the frame. The instructions are further executable by the processor to generate a first reference PSD distribution indicative of an estimate of background noise in the frame based on a non-linear weight, a second reference PSD distribution of a previous frame of the audio signal, and a second PSD distribution of the previous frame. The instructions are further executable by the processor to determine whether voice activity is detected in the frame based on the first PSD distribution of the frame and the first reference PSD distribution.

[0006] A computer-readable storage device stores instructions executable by a processor to select a frame of an audio signal. The instructions are further executable by the processor to determine a first power spectral density (PSD) distribution of the frame. The instructions are further executable by the processor to generate a first reference PSD distribution indicative of an estimate of background noise in the frame based on a non-linear weight, a second reference PSD distribution of a previous frame of the audio signal, and a second PSD distribution of the previous frame. The instructions are further executable by the processor to determine whether voice activity is detected in the frame based on the first PSD distribution of the frame and the first reference PSD distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] For a detailed description of various examples, reference will now be made to the drawings, in which:

[0008] Figure 1 Illustrates an apparatus for performing voice activity detection using non-linear weighted background noise estimation.

[0009] Figure 2 Illustrates an alternative apparatus for performing voice activity detection using non-linear weighted background noise estimation, the apparatus including a filter bank in place of a Fourier transform calculator.

[0010] Figure 3 Illustrates a flowchart of a method for performing voice activity detection using non-linear weighted background noise estimation.

[0011] Figure 4 Is a block diagram of an example computing device that can be used to perform voice activity detection using non-linear weighted background noise estimation. DETAILED DESCRIPTION

[0012] Reference Figure 1 , shows a block diagram of an apparatus 100 for performing voice activity detection. Apparatus 100 includes a microphone 102, an amplifier 104, an analog-to-digital converter (ADC) 106, a window selector 108, a Fourier transform calculator 110, an entropy calculator 112, an energy calculator 114, an entropy difference calculator 116, an energy difference calculator 118, a background entropy calculator 120, a background energy calculator 122, a non-linear reference distribution update calculator 126, a reference distribution storage device 128, an energy entropy feature calculator 130, and a voice activity detector 132. The window selector 108, the Fourier transform calculator 110, the entropy calculator 112, the energy calculator 114, the entropy difference calculator 116, the energy difference calculator 118, the background entropy calculator 120, the background energy calculator 122, the non-linear reference distribution update calculator 126, the energy entropy feature calculator 130, and the voice activity detector 132 correspond to dedicated hardware, software executed by a processor of apparatus 100, or a combination thereof.

[0013] The microphone 102 can correspond to any type of microphone, including a capacitive microphone, a dynamic microphone, a ribbon microphone, a piezoelectric microphone, a microelectromechanical system (MEMS) microphone, etc. The microphone 102 is configured to generate an electrical signal based on sound waves. For example, the microphone 102 can generate an electrical signal based on voice activity, background noise, or a combination thereof.

[0014] The amplifier 104 can correspond to a programmable gain amplifier or other types of amplifiers. The amplifier 104 is configured to receive the electrical signal generated by the microphone and adjust (e.g., increase) the power of the electrical signal.

[0015] The ADC 106 is configured to convert the boosted electrical signal into a digital signal x[n], where n represents a discrete-time sample instance. The ADC 106 can include a delta-sigma modulation ADC or other types of ADCs.

[0016] The window selector 108 is configured to receive the digital signal x[n] generated by the ADC 106 and generate a frame of the digital signal. In some embodiments, the window selector 108 is configured to apply a Hamming window function to select one or more frames (j) from the digital signal. The window selector 108 can correspond to dedicated hardware (e.g., a window selector circuit), correspond to software executed by a processor (not shown) of the device 100, or correspond to a combination thereof. The window selector 108 can generate a frame of the signal x[n] according to the formula w,l x[n] = w[n]x[lB + n], where B is the window length, n identifies a specific sample from the sequence x[n], w[n] is a list of multiplicative scale factors of a given window sequence (e.g., a Hamming window, a rectangular window, etc.), and l is the frame index. In some embodiments, the window selector 108 generates frames at a random rate. Thus, the window selector 108 can generate frames that capture repetitive coherent interval noise that might otherwise fall between frames.

[0017] The Fourier transform calculator 110 is configured to apply a Fourier transform to each frame (l) of the digital signal to transform the frame from the discrete time domain to the discrete frequency domain. In some embodiments, the Fourier transform calculator 110 is configured to transform the frame to the frequency domain according to the formula where k is the frequency band index, N FFT represents the number of samples in the Fourier transform (in some embodiments, for some integer q, N FFT is equal to 2 q . For example, N FFTIt can be 256 or 512), and l is the frame index. The Fourier transform calculator 110 is also configured to generate a power spectral density (PSD) distribution for each frame l based on the conversion of the frame to the frequency domain. In a specific example, the Fourier transform calculator 110 is configured to generate PSDS(k, l) = X(k, l)X * (k, l), where * represents the complex conjugate operation. Thus, for each frame l, the Fourier transform calculator 110 can generate a vector of PSD values. For example, for frame 1, the Fourier transform calculator 110 can generate a PSD distribution [S 1,1 , S 2,1 , … S 11,1 , where S 1,1 is the PSD value of the first frequency band of frame 1, S 2,1 is the PSD value of the second frequency band of frame 1, and so on.

[0018] The entropy calculator 112 is configured to calculate the entropy of each frame (l) based on the PSD distribution S(k, l) of the frame. For example, the entropy calculator 112 can normalize the power spectrum of each frame (l) to produce a probability distribution, where each probability value is The entropy calculator 112 can then calculate the entropy H of frame (l), where H(l) = -∑ k P(k, l) log 2 P(k, l).

[0019] The energy calculator 114 is configured to calculate the energy of each frame (l) by integrating the PSD distribution of the frame. For example, for each frame (l), the energy calculator 114 can determine the energy (E(l)) according to the equation E(l) = ∑ n x w,l 2 [n] = ∑ k S(k, l).

[0020] The background entropy calculator 120 is configured to calculate the entropy value (H noise (l)) attributable to background noise for each frame (l) based on the reference PSD distribution S noise (k, l) of frame (l) stored in the reference distribution storage device 128. As further described below, the non-linear reference distribution update calculator 126 generates the reference PSD distribution of each frame (l) except the first frame based on the PSD distribution of the previous frame (e.g., based on S(k, l - 1)). The reference distribution of the first frame can correspond to a zero vector (e.g., [0, …, 0]). The background entropy calculator 120 is configured to generate Then the background entropy calculator 120 can calculate the background entropy H noise of frame (l), where H noise (l) = -∑k P noise (k, l) log 2 P noise (k, l).

[0021] Similarly, the background energy calculator 122 is configured to calculate the energy value (E noise (l)) attributable to background noise for each frame (l) based on the reference PSD distribution S noise (k, l) of the frame (l) stored in the reference distribution storage device 128. The background energy calculator 122 is configured to calculate the background energy for each frame (l) by integrating the reference PSD distribution of the frame. For example, for each frame (l), the background energy calculator 122 may calculate according to the equation E noise (l) = ∑ n x w,l 2 [n] = ∑ k S noise (k, l) to determine the energy (E noise (l)).

[0022] The non-linear reference distribution update calculator 126 is configured to non-linearly update the reference PSD distribution (E noise (k, l)) for each frame (l) to generate the reference PSD distribution (S noise (k, l)) of the subsequent frame (l + 1) based on the reference PSD distribution (S noise (k, l)) of the frame, the PSD distribution of the frame (S(k, l)), and the non-linear weight term. In a particular embodiment, the non-linear reference distribution update calculator 126 generates the reference PSD distribution (S (k, l + 1)) of the subsequent frame (l + 1) according to the equation noise (k, l + 1)), where D KL (S noise (k, l) || S(k, l)) is the Kullback-Leibler divergence between the reference PSD distribution (S noise (k, l)) of the frame and the PSD distribution (S(k, l)) of the frame and a is a weight term between 0 and 1. The Kullback-Leibler divergence between the probability distributions P(i) and Q(i) is D KL (P|Q) = ∑ i P(i) log(P(i) / Q(i)). For all values of i, this function of the distribution is zero when P(i) = Q(i) and takes positive values, qualitatively measuring the similarity between the distributions. The non-linear reference distribution update calculator 126 applies the Kullback-Leibler divergence weight to any frame used to update the background noise estimate.

[0023] Because the reference PSD distribution corresponding to the background noise estimate is updated non-linearly based on the PSD distribution of the detected sound, the background noise model used by the apparatus 100 may be less susceptible to sudden and increased effects of short-duration sounds (e.g., a door slamming).

[0024] The entropy difference calculator 116 is configured to determine an entropy difference (ΔH(l)) for each frame (l) by subtracting the noise entropy of the frame from the entropy of the frame according to the equation ΔH(l) = |H(l) - H noise (l)|. Similarly, the energy difference calculator 118 is configured to determine an energy difference (ΔE(l)) for each frame (l) by subtracting the noise energy of the frame from the energy of the frame according to the equation ΔE(l) = |E(l) - E noise (l)|.

[0025] The energy entropy feature calculator 130 is configured to calculate an energy entropy feature (F(l)) for each frame (l) based on the entropy difference (ΔH(l)) and energy difference (ΔE(l)) of the frame. For example, the energy entropy feature calculator 130 may calculate the energy entropy feature according to the equation to calculate the energy entropy feature.

[0026] The voice activity detector 132 is configured to compare the energy entropy feature (F(l)) with a threshold for each frame (l) to determine whether the frame (l) includes voice activity. In response to determining that F(l) meets the threshold, the voice activity detector 132 is configured to determine that voice activity is present in the frame (l). In response to determining that the frame (l) does not meet the threshold, the voice activity detector 132 is configured to determine that voice activity is not present in the frame (l). The voice activity detector 132 may be configured to determine that a value greater than the threshold, less than the threshold, greater than or equal to the threshold, or less than or equal to the threshold meets the threshold. The voice activity detector 132 may be configured to initiate one or more actions in response to detecting voice activity in the frame. For example, the voice activity detector 132 may initiate natural language processing of the frame in response to detecting voice activity in the frame.

[0027] Thus, Figure 1 the apparatus can be used to perform voice activity detection. Because the apparatus 100 updates the background noise estimate non-linearly, the apparatus 100 may be less susceptible to sudden changes in the noise level of short duration compared to other voice activity detection apparatuses. Additionally, because the apparatus 100 generates frames for voice activity detection at random intervals, the apparatus 100 can detect evenly spaced noise (e.g., speech) that might otherwise fall between evenly spaced frames. In other embodiments, the apparatus 100 may have an alternative configuration. For example, the above components may be combined or decomposed into different combinations.

[0028] ReferenceFigure 2 , a block diagram of a second apparatus 200 for performing voice activity detection is shown. The second apparatus 200 corresponds to the apparatus 100, except that the second apparatus 200 includes a filter bank 210 instead of the Fourier transform calculator 110 and the second apparatus 200 does not include the window selector 108. Instead, the ADC 106 directly outputs the digital signal x[n] to the filter bank 210. The filter bank 210 corresponds to a plurality of filters configured to output the PSD distribution of frames. The filter bank applies several finite impulse response filters to the digital signal to separate the digital signal into a set of parallel frequency bands. For frequency band i = 1, …, R, the impulse response is represented by h i [n]. The output of each band filter is given by y i [l] = ∑ m x[m]h i [lD - m]. The filter bank output for frame l is a vector [y 1 [l] … y R [l]] T generated by assembling the band filter outputs, where the superscript T represents the transpose operation. The power spectral density elements are generated by squaring the elements of the filter bank output vector. Thus, for frame l, the PSD output is given by . Thus, the apparatus for performing voice activity detection can be implemented with a Fourier transform calculator (e.g., software executable by a processor to perform a Fourier transform or hardware configured to perform a Fourier transform) or a filter bank.

[0029] Reference Figure 3 shows a flowchart depicting a method 300 for performing voice activity detection. The method 300 can be executed by a computing device, such as Figure 1 the apparatus 100 or Figure 2 the second apparatus 200.

[0030] Method 300 includes, at 302, receiving an input audio signal. For example, the microphone 102 can generate an analog audio signal based on the detected sound, the amplifier 104 can amplify the analog audio signal, and the analog - to - digital converter 106 can generate a digital audio signal based on the amplified analog signal. Then the digital audio signal can be received by the window selector 108.

[0031] Method 300 further includes, at 304, selecting a window of the audio signal. For example, the window selector 108 can use a Hamming window function such as x w,l [n] = w[n]x[lB + n] to select a frame of the digital signal output by the ADC 106. In some embodiments, the window selector 108 generates windows at random intervals.

[0032] Method 300 further includes, at 306, determining the distribution of the band power in the frame. For example, the Fourier transform calculator 110 (or the filter bank 210) may output the PSD distribution of the frame according to the equation S(k, l) = X(k, l)X*(k, l), where X[k, l] is the frequency domain mapping of the frame (l), and * represents the complex conjugate operation. The frequency domain mapping of the window may be generated by the Fourier transform calculator (or the filter bank 210).

[0033] Method 300 further includes, at 308, determining a first entropy and a first energy of the distribution of the band power in the frame. For example, the entropy calculator 112 may determine the entropy (H(l)) of the PSD distribution of the frame (l). The entropy calculator 112 may generate the entropy (H(l)) by normalizing the PSD distribution (S(k, l)) of the frame (l) and calculating H(l) = -∑ P(k, l) log k P(k, l). In addition, the energy calculator 114 may determine the energy (E(l)) of the PSD distribution (S(k, l)) of the frame (l) according to the equation E(l) = ∑ 2 x n x w,l 2 [n] = ∑ k S(k, l).

[0034] Method 300 further includes, at 310, retrieving a reference distribution of the band power. For example, the background entropy calculator 120 and the background energy calculator 122 may retrieve the reference PSD distribution (S noise (k, l)) from the reference distribution storage device 128. The reference PSD distribution (S noise (k, l)) may correspond to the estimated PSD distribution of the noise within the frame (l). The reference PSD distribution may be based on the PSD distribution of the previous frame. For the first frame, the reference PSD distribution may correspond to a zero vector.

[0035] Method 300 further includes, at 312, determining a second entropy and a second energy of the reference distribution of the band power. For example, the background entropy calculator 120 may calculate the background entropy (H noise (l)) of the frame (l) based on the reference PSD distribution (S noise (l)), and the background energy calculator 122 may calculate the background entropy (E noise (l)) of the frame (l) based on the reference PSD distribution (S noise (l)).

[0036] Method 300 further includes, at 314, determining a first difference between the first entropy and the second entropy. For example, the entropy difference calculator 116 may, according to the equation ΔH(l) = |H(l) - H noise(l) Determine the difference (ΔH(l)) between the entropy of the frame (H(l)) and the background entropy of the frame (H noise (l)).

[0037] Method 300 further includes, at 316, determining a second difference between a first energy and a second energy. For example, the energy difference calculator 118 may determine the difference (ΔE(l)) between the energy of the frame (E(l)) and the background energy of the frame (H noise (l)) according to the equation ΔE(l) = |E(l) - E noise (l)|.

[0038] Method 300 further includes, at 318, determining an energy entropy feature based on the first difference and the second difference. For example, the energy entropy feature calculator 130 may determine the energy entropy feature (F(l)) according to the equation based on the entropy difference (ΔH(l)) and the energy difference (ΔE(l)) of the frame.

[0039] Method 300 further includes, at 320, determining whether the energy entropy feature meets a threshold. For example, the voice activity detector 132 may compare the energy entropy feature (F(l)) of the frame (l) to determine whether voice activity exists in the frame (l). The voice activity detector 132 may determine that the threshold is met in response to the energy entropy feature (F(l)) exceeding the threshold (or being greater than or equal to the threshold). The threshold may be based on the microphone 412, the gain of the amplifier 410, the number of bits of the ADC 404, or a combination thereof.

[0040] Method 300 further includes, at 302, determining that voice activity exists in the frame in response to the energy entropy feature meeting the threshold, or at 324, determining that no voice activity exists in the frame in response to the energy entropy feature not meeting the threshold. For example, the voice activity detector 132 may determine that voice activity exists in the frame (l) in response to the energy entropy feature (F(l)) being greater than or equal to the threshold, or determine that no voice activity exists in the frame (l) in response to the energy entropy feature (F(l)) being less than the threshold.

[0041] Method 300 further includes, at 326, determining a non - linear weight based on the divergence between the distribution of the band power in the frame and a reference distribution of the band power. For example, the non - linear reference distribution update calculator 126 may determine the non - linear weight based on the Kullback - Leibler divergence (D noise (S KL (k, l)||S(k, l))) between the PSD distribution (S(k, l)) of the frame (l) and the reference PSD distribution (S noise (k, l)) of the frame (l).

[0042] Method 300 further includes, at 328, updating a reference distribution of the band power based on a non-linear weight. For example, the non-linear reference distribution update calculator 126 may calculate a reference PSD distribution (S noise (k, l+1)) for the estimated noise in a subsequent frame (l+1) based on the non-linear weight (F(l)). Specifically, the non-linear reference distribution update calculator 126 may calculate the reference PSD distribution for the subsequent frame according to the formula

[0043] Method 300 further includes, at 304, selecting a subsequent frame and, at 304, continuing with the updated reference distribution. For example, the window selector 108 may select a subsequent frame (l+1), and the subsequent frame (l+1) may be processed as described above for frame (l), except that the updated reference PSD S noise (k, l+1) is used by the background entropy calculator 120 and the background energy calculator 122 to calculate the background entropy and the background energy. Thus, method 300 non-linearly updates the background noise estimate based on the detected sound. Therefore, method 300 for performing voice activity detection may be more accurate in cases where sudden and inconsistent changes occur in the voice activity. Method 300 may be arranged in a different order than that shown in some embodiments.

[0044] Reference Figure 4 , illustrates a block diagram of a computing device 400 that may perform voice activity detection. In some embodiments, the computing device 400 corresponds to device 100 or device 200. The computing device 400 includes a processor 402. The processor 402 may include a digital signal processor, a microprocessor, a microcontroller, another type of processor, or a combination thereof.

[0045] The computing device further includes a memory 406 connected to the processor 402. The memory 406 includes a computer-readable storage device, such as a read-only memory device, a random access memory device, a solid state drive, another type of memory device, or a combination thereof. As used herein, a computer-readable storage device refers to an article of manufacture and not a transient signal.

[0046] The memory 406 stores voice activity detection instructions 408 that are executable to perform one or more of the operations described herein Figures 1-3 . For example, the voice activity detection instructions 408 may be executed by the processor 402 to perform method 300.

[0047] The computing device 400 further includes an ADC 404 connected to the processor 402. The ADC 404 may correspond to a delta-sigma modulated ADC or another type of ADC. The ADC 404 may correspond to Figure 1 and Figure 2 the ADC 106.​

[0048] The computing device 400 further includes an amplifier 410 connected to the ADC 404. The amplifier 410 may include a programmable gain amplifier or another type of amplifier. The amplifier 410 may correspond to Figure 1 and Figure 2 the amplifier 104.

[0049] The computing device 400 further includes a microphone 412 connected to the amplifier 410. The microphone 412 may correspond to Figure 1 and Figure 2 the microphone 102.

[0050] In operation, the microphone 412 generates an analog signal based on the sound detected in the environment, the amplifier 410 amplifies the analog signal, and the ADC 404 generates a digital signal based on the amplified signal. The processor 402 executes the voice activity detection instruction 408 to perform non-linear scaling voice activity detection on the digital signal as described herein. Thus, compared to other devices, the computing device 400 can be used to provide relatively more accurate voice activity detection.

[0051] The computing device 400 may have an alternative configuration in other embodiments. These alternative configurations may include additional and / or fewer components. For example, in some embodiments, one or more of the microphone 412, the amplifier 410, and the ADC 404 are external to the computing device 400 and the computing device 400 includes an interface configured to receive signals or data from the ADC 404, the amplifier 410, or the microphone 412. Additionally, although direct connections between the components of the computing device 400 are illustrated, in some embodiments, these components are connected via a bus or other indirect connections.

[0052] The term "coupled" is used throughout the specification. This term may encompass connections, communications, or signal paths that achieve a functional relationship consistent with this description. For example, if device A generates a signal to control device B to perform an action, then in a first example, device A is coupled to device B, or in a second example, if an intervening component C does not substantially change the functional relationship between device A and device B such that device B is controlled by device A via the control signal generated by device A, then device A is coupled to device B via the intervening component C.

[0053] Modifications to the described embodiments are possible within the scope of the claims, and other embodiments are possible.

Claims

1. A method for voice activity detection, comprising: determining a first power spectral density distribution of a frame of a signal, i.e., a first PSD distribution; determining a first reference PSD distribution indicative of an estimate of noise in the frame, wherein the determination of the first reference PSD distribution is based on a non-linear weight, a second reference PSD distribution of a previous frame of the signal, and a second PSD distribution of the previous frame; determining a first entropy of the first PSD distribution and a second entropy of the first reference PSD distribution; determining a first energy of the first PSD distribution and a second energy of the first reference PSD distribution; and determining whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy.

2. The method according to claim 1, further comprising determining the non-linear weight based on a divergence between the second PSD distribution and the second reference PSD distribution.

3. The method according to claim 2, wherein the divergence corresponds to the Kullback-Leibler divergence.

4. The method according to claim 1, wherein the signal is an audio signal.

5. The method according to claim 1, further comprising: determining an energy difference ΔE between the first energy and the second energy; determining an entropy difference ΔH between the first entropy and the second entropy; and determining an energy entropy feature based on the energy difference and the entropy difference, wherein determining whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy includes determining whether the energy entropy feature meets a threshold.

6. The method according to claim 5, wherein the energy entropy feature is equal to 7. The method according to claim 1, further comprising generating the frame using a Hamming window function.

8. The method according to claim 3, wherein the non-linear weight is equal to according to the equation to be determined, where D KL (S noise (k, l) || S(k, l)) is the Kullback-Leibler divergence between the second reference PSD distribution S noise (k, l) of the frame and the second PSD distribution S(k, l) of the frame and a is a weight term between 0 and 1.

9. A device for voice activity detection, comprising: a processor; and a memory storing instructions executable by the processor to: determine a first power spectral density distribution of a frame of a signal, i.e., a first PSD distribution; generate a first reference PSD distribution indicative of an estimate of noise in the frame, wherein the determination of the first reference PSD distribution is based on a non-linear weight, a second reference PSD distribution of a previous frame of the signal, and a second PSD distribution of the previous frame; determine a first entropy of the first PSD distribution and a second entropy of the first reference PSD distribution; determine a first energy of the first PSD distribution and a second energy of the first reference PSD distribution; and determine whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy.

10. The device according to claim 9, wherein the instructions are further executable by the processor to determine the non-linear weight based on a divergence between the second PSD distribution and the second reference PSD distribution.

11. The device according to claim 10, wherein the divergence corresponds to the Kullback-Leibler divergence.

12. The device according to claim 9, wherein the signal is an audio signal.

13. The device according to claim 9, wherein the instructions are further executable by the processor to: Determine the energy difference ΔE between the first energy and the second energy; Determine the entropy difference ΔH between the first entropy and the second entropy; and Determine an energy-entropy feature based on the energy difference and the entropy difference, wherein determining whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy includes determining whether the energy-entropy feature meets a threshold.

14. The device according to claim 13, wherein the energy entropy feature is equal to 15. The apparatus according to claim 9, wherein the instructions are further executable by the processor to generate the frame using a Hamming window function.

16. The device according to claim 11, wherein the non-linear weight is equal to according to the equation to be determined, where D KL (S noise (k, l) || S(k, l)) is the second reference PSD distribution S of the frame noise (k, l) and the Kullback-Leibler divergence between the second PSD distribution S(k, l) of the frame and a is a weight term between 0 and 1.

17. A computer-readable storage device storing instructions that are executable by a processor to: Determine a first power spectral density distribution, i.e., a first PSD distribution, of a frame of a signal; Generate a first reference PSD distribution indicative of an estimate of noise in the frame, wherein determining the first reference PSD distribution is based on a non-linear weight, a second reference PSD distribution of a previous frame of the signal, and a second PSD distribution of the previous frame; Determine a first entropy of the first PSD distribution and a second entropy of the first reference PSD distribution; Determine a first energy of the first PSD distribution and a second energy of the first reference PSD distribution; And Determine whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy.

18. The computer-readable storage device according to claim 17, wherein the instructions are further executable by the processor to determine the non-linear weight based on a divergence between the second PSD distribution and the second reference PSD distribution.

19. The computer-readable storage device according to claim 18, wherein the divergence corresponds to the Kullback-Leibler divergence.

20. The computer-readable storage device according to claim 17, wherein the signal is an audio signal.

21. The computer-readable storage device according to claim 17, wherein the instructions are further executable by the processor to: Determine the energy difference ΔE between the first energy and the second energy; Determine the entropy difference ΔH between the first entropy and the second entropy; And Determine an energy-entropy feature based on the energy difference and the entropy difference, wherein determining whether voice activity is detected in the frame based on the first entropy, the second entropy, the first energy, and the second energy includes determining whether the energy-entropy feature meets a threshold.

22. The computer-readable storage device according to claim 21, wherein the energy entropy feature is equal to 23. The computer-readable storage device according to claim 19, wherein the non-linear weight is equal to according to the equation to be determined, where D KL (S noise (k, l) || S(k, l)) is the Kullback-Leibler divergence between the second reference PSD distribution S noise (k, l) of the frame and the second PSD distribution S(k, l) of the frame, and a is a weight term between 0 and 1.

24. A method for voice activity detection, Comprising: Determine a first power spectral density distribution, i.e., a first PSD distribution, of a frame of an audio signal; Determine a first reference PSD distribution indicative of an estimate of noise in the frame, wherein determining the first reference PSD distribution is based on a non-linear weight, a second reference PSD distribution of a previous frame of the audio signal, and a second PSD distribution of the previous frame; Determine a first entropy of the first PSD distribution and a second entropy of the first reference PSD distribution; Determine a first energy of the first PSD distribution and a second energy of the first reference PSD distribution; Determine the energy difference between the first energy and the second energy; Determine the entropy difference between the first entropy and the second entropy; Determine the energy entropy feature based on the energy difference and the entropy difference; And Determine whether voice activity is detected in the frame based on whether the energy entropy feature meets a threshold.

25. A device for voice activity detection, Comprising: A processor; And A memory storing instructions executable by the processor to: Determine the first power spectral density distribution, i.e., the first PSD distribution, of a frame of an audio signal; Generate a first reference PSD distribution indicating an estimate of the noise in the frame, and the determination of the first reference PSD distribution is based on non-linear weights, the second reference PSD distribution of the previous frame of the audio signal, and the second PSD distribution of the previous frame; Determine the first entropy of the first PSD distribution and the second entropy of the first reference PSD distribution; Determine the first energy of the first PSD distribution and the second energy of the first reference PSD distribution; Determine the energy difference between the first energy and the second energy; Determine the entropy difference between the first entropy and the second entropy; Determine the energy entropy feature based on the energy difference and the entropy difference; And Determine whether voice activity is detected in the frame based on whether the energy entropy feature meets a threshold.

Citation Information

Patent Citations

  • Speech endpoint detecting method and device

    CN102097095A

  • Voice noise reduction method based on information entropy weighting

    CN110444222A