Voice Activity Detection Device and Method

The environmental detection results are generated through the environment detection circuit and the appropriate speech activity detection algorithm or result are selected based on the results, which solves the problem of poor performance in different environmental conditions in the prior art, and achieves higher speech activity detection accuracy and energy efficiency.

CN114187926BActive Publication Date: 2025-05-30REALTEK SEMICON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010969320.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-15
Publication Date
2025-05-30
Estimated Expiration
2040-09-15

AI Technical Summary

Technical Problem

The existing voice activity detection technology performs poorly under different environmental conditions, especially in non-steady state noise environments, with high error triggering and missing indicator values, which affects user experience and power consumption.

Method used

The sound input signal is processed through the environment detection circuit, an environment detection result is generated, and an appropriate speech activity detection algorithm or result is selected based on the result to optimize the performance of speech activity detection.

Benefits of technology

Under different environmental conditions, by dynamically selecting voice activity detection algorithms or results, the accuracy and efficiency of voice activity detection are significantly improved, false triggering and missing indicator values ​​are reduced, and user experience and equipment energy efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187926B_ABST
    Figure CN114187926B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice activity detection device and method, which can select one of multiple voice activity detection results as the basis for whether there is voice activity according to the environmental detection result. The voice activity detection device includes an environmental detection circuit, a voice activity detection circuit, and a voice activity decision circuit. The environmental detection circuit is used to process the sound input signal to generate an environmental detection result. The voice activity detection circuit is used to analyze the sound input signal according to multiple voice activity detection algorithms to generate multiple voice activity detection results. The voice activity decision circuit is used to select one of the multiple voice activity detection results according to the environmental detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice activity detection device and method, and more particularly to a voice activity detection device and method that can adaptively employ one of different voice activity detection algorithms. Background Art

[0002] Many electronic devices (e.g., mobile devices such as smart phones, smart watches, smart speakers, etc.) can use a speech recognition function to determine commands spoken by a user and perform corresponding operations accordingly. To avoid missing commands spoken by the user, the electronic device can keep the speech recognition function in an always listening state; however, most of the time, the sound signals received by the speech recognition function are not user commands. Therefore, to reduce unnecessary processing and power consumption, the electronic device can use voice activity detection (VAD) to determine whether there is speech and control the operation of the speech recognition function accordingly. More specifically, when there is speech, the electronic device wakes up the speech recognition function to determine whether there is a user command; when there is no speech, the electronic device can turn off the speech recognition function to reduce power consumption. The operation flowchart of a general voice wake-up system is as Figure 1 shown and includes:

[0003] Step S110: Detect voice activity based on an input signal and deactivate the speech recognition function.

[0004] Step S120: Determine whether there is voice activity; if so, go to step S130; if not, return to step S110.

[0005] Step S130: Wake up the speech recognition function and perform speech recognition.

[0006] Step S140: Determine whether there is a user command; if so, go to step S150; if not, return to step S110.

[0007] Step S150: Perform a corresponding operation according to the user command and then return to step S110.

[0008] In practical applications, voice activity detection may run in an environment with many different background noises, which can be classified into stationary noise and non-stationary noise. The energy of stationary noise does not vary much over time, such as the sound of a fan or the noise in a quiet office. This type of noise has less impact on voice activity detection. Non-stationary noise, on the other hand, has a large variation in energy over time, such as the sound of a TV, street traffic, or people talking. The characteristics of many non-stationary noises are similar to those of human voices, which can affect the performance of voice activity detection and reduce the detection accuracy.

[0009] The performance of voice activity detection can be evaluated according to two metric values. One is the metric value of "misjudging speech as noise" (abbreviated as the miss metric value), and the other is the metric value of "misjudging noise as speech" (abbreviated as the false trigger metric value). The relationship between these two metric values is usually a trade-off relationship. When the miss metric value increases, users may have to repeat commands frequently, which will result in a worse user experience. When the false trigger metric value increases, the electronic device will be forced to perform unnecessary signal processing and data transmission, which will increase power consumption.

[0010] General electronic devices use fixed voice activity detection algorithms. A fixed voice activity detection algorithm may perform well in an environment with a certain background noise but poorly in another background noise environment. Therefore, there is a need in the art for a technology that can adopt different voice activity detection algorithms in response to different environmental conditions to achieve good voice activity detection performance under different environmental conditions. Summary of the Invention

[0011] One object of the present disclosure is to provide a voice activity detection device and method to avoid the problems of the prior art.

[0012] An embodiment of the voice activity detection device of the present disclosure can select one of multiple voice activity detection results as the basis for whether there is voice activity according to the environmental detection result. This embodiment includes an environmental detection circuit, a voice activity detection circuit, and a voice activity decision circuit. The environmental detection circuit is used to process the sound input signal to generate an environmental detection result. The voice activity detection circuit is used to analyze the sound input signal according to multiple voice activity detection algorithms to generate multiple voice activity detection results. The voice activity decision circuit is used to select one of the multiple voice activity detection results according to the environmental detection result.

[0013] Another embodiment of the voice activity detection device of the present disclosure can select one of multiple voice activity detection algorithms according to the environmental detection result, and then generate a voice activity detection result based on this to determine whether there is voice activity. The voice activity detection device includes an environmental detection circuit and a voice activity detection and decision circuit. The environmental detection circuit is used to process the sound input signal to generate an environmental detection result. The voice activity detection and decision circuit is used to select one of multiple voice activity detection algorithms as the effective voice activity detection algorithm according to the environmental detection result, and then analyze the sound input signal according to the effective voice activity detection algorithm to generate a voice activity detection result as the basis for determining whether there is voice activity.

[0014] An embodiment of the voice activity detection method of the present disclosure can select one of multiple voice activity detection results / algorithms according to the environmental detection result, including the following steps: receiving and processing the sound input signal to generate the environmental detection result; and selecting one of the multiple voice activity detection results as the final voice activity detection result according to the environmental detection result, or selecting one of the multiple voice activity detection algorithms according to the environmental detection result and generating the final voice activity detection result based on this, where the multiple voice activity detection results are respectively generated according to the multiple voice activity detection algorithms.

[0015] Regarding the features, implementation, and effects of the present invention, the following provides a detailed description of the preferred embodiments in conjunction with the drawings. Description of the Drawings

[0016] Figure 1 Showing the operation flowchart of a general voice wake-up system;

[0017] Figure 2 Showing an embodiment of the voice activity detection device of the present disclosure;

[0018] Figure 3 Showing Figure 2 an embodiment of the environmental detection circuit;

[0019] Figure 4 Showing Figure 3 the steps performed by the energy change detection circuit;

[0020] Figure 5 Showing Figure 2 another embodiment of the environmental detection circuit;

[0021] Figure 6 Showing another embodiment of the voice activity detection device of the present disclosure; and

[0022] Figure 7 Showing an embodiment of the voice activity detection method of the present disclosure. Detailed Description of the Preferred Embodiments

[0023] The present disclosure discloses a voice activity detection (VAD) device and method, which can adopt different voice activity detection results / algorithms respectively in response to different environmental conditions to achieve good voice activity detection performance.

[0024] Figure 2 An embodiment of the voice activity detection device of the present disclosure is shown, which can select one of multiple voice activity detection results as the basis for whether there is voice activity according to the environmental detection result. Figure 2 The voice activity detection device 200 includes an environmental detection circuit 210, a voice activity detection circuit 220, and a voice activity decision circuit 230. The environmental detection circuit 210 is used to process the sound input signal to generate an environmental detection result. The voice activity detection circuit 220 is used to analyze the sound input signal according to multiple voice activity detection algorithms to generate multiple voice activity detection results; the voice activity detection circuit 220 itself can be a known or self-developed circuit, and the multiple voice activity detection algorithms can be known or self-developed algorithms, and the performance (for example: miss value and false trigger value) of different algorithms is usually different. The voice activity decision circuit 230 is used to select one of the multiple voice activity detection results according to the environmental detection result.

[0025] Figure 3 Shown Figure 2 An embodiment of the environmental detection circuit 210 is shown, including a signal analysis circuit 310, an energy change detection circuit 320, and a change information decision circuit 330. These circuits are described as follows.

[0026] Please refer to Figure 3。The signal analysis circuit 310 is used to generate M processed signals based on the voice input signal, where the M processed signals are M band signals or M frequency domain signals, and M is a positive integer. More specifically, during the process of processing the voice input signal, the signal analysis circuit 310 continuously receives the voice input signal and samples the voice input signal; after obtaining J sampling values (e.g., multiple sampling values) of the voice input signal sufficient to form a frame, the signal analysis circuit 310 then generates M processed signals of this frame accordingly. In an implementation example, the signal analysis circuit 310 includes at least one filter circuit, and the at least one filter circuit is used to generate M band signals of each frame based on the voice input signal; for example, the at least one filter circuit includes M filters, and each filter generates a band signal, so that the M filters generate the M band signals. In another implementation example, the signal analysis circuit 310 includes at least one conversion circuit (e.g., a Fast Fourier Transform (FFT) circuit), and the at least one conversion circuit is used to generate M frequency domain signals of each frame based on the voice input signal.

[0027] Please refer to Figure 3 。The energy change detection circuit 320 is used to perform calculations based on the M processed signals of each frame to generate energy change values for each frame, and a total of X energy change values for L frames, where X is equal to M multiplied by L, and L is the number of frames. In an implementation example, the energy change detection circuit 320 performs multiple steps as shown in Figure 4 below, including:

[0028] Step S410: Perform calculations based on the M processed signals of each of the L frames to obtain X signal energy values. For example, step S410 calculates the energy of each band / frequency domain signal in each frame according to the following formula (1) (e.g., the sum of the energies of N sampling points of each band signal in each frame, and the sampling period corresponding to each sampling point is like or ) to obtain M×L = X signal energy values (E m,l ).

[0029]

[0030] In formula (1), l is the frame index between 1 and L, m is the band / frequency domain signal index between 1 and M, M is the number of band / frequency domain signals corresponding to the l-th frame, N is the number of data points of the m-th band / frequency domain signal in the l-th frame, and x m,l (k) is the value of the k-th point of the m-th band / frequency domain signal in the l-th frame.

[0031] Step S420: Calculate X short-term energy values based on the X signal energy values and the number of short-term frames (p st ) and calculate X long-term energy values based on the X signal energy values and the number of long-term frames (p lt ). For example, step S420 calculates the X short-term average energy values (E_st m,l ) and the X long-term average energy values (E_lt m,l ) according to the following formula (2).

[0032]

[0033] Step S430: Obtain X energy relationship values based on the X short-term energy values and the X long-term energy values. For example, step S430 calculates the X energy relationship values according to the following formula (3).

[0034]

[0035] Step S440: Compare each of the X energy relationship values with an energy threshold (thr m ) to generate X energy change values. For example, if the energy relationship value is greater than the energy threshold, step S440 sets the energy change value (fg_E_var m,l ) to 1 to represent a large energy change; if the energy relationship value is not greater than the energy threshold, step S440 sets the energy change value to 0 to represent a small energy change.

[0036] Please refer to Figure 3 . The change information decision circuit 330 is used to process the X energy change values to generate L energy change detection values, then compare each of the L energy change detection values with a change threshold to generate L comparison results, and then generate the environmental detection result based on the L comparison results. In an implementation example, the change information decision circuit 330 adds M energy change values in each frame (for each value of the frame index) among the X energy change values as shown in the following formula (4) to generate L energy change detection values (S_E_var l ); then the change information decision circuit 330 compares each of the L energy change detection values with a change threshold (thr) to generate L comparison results (fg_S l)As shown in the following formula (5); if all / majority of the energy change detection values among the L comparison results (for example: the L energy change detection values) are greater than the change threshold, the change information decision circuit 330 determines that the energy change in the current environment is large; if all / majority of the energy change detection values among the L comparison results are less than the change threshold, the change information decision circuit 330 determines that the energy change in the current environment is small.

[0037]

[0038] fg_S l represents the comparison result between S_E_var l and thr, formula (5)

[0039] Please refer to Figure 2 and Figure 3 . The voice activity decision circuit 230 selects one of the multiple voice activity detection results according to the preset rules and the change of the L comparison results. The preset rule is that when the change of the L comparison results is greater than the preset change degree (i.e., when the energy change in the current environment is large), one of the multiple voice activity detection results is selected; the preset rule is that when the change of the L comparison results is less than the preset change degree (i.e., when the energy change in the current environment is small), another one of the multiple voice activity detection results is selected. For example, the characteristics of pitch-based voice activity detection (pitch-based VAD) and energy-based voice activity detection (energy-based VAD) are shown in Table 1 below; if the voice activity decision circuit 230 first considers the low miss value and then the low false trigger value, in the case of a large energy change in the current environment, the voice activity decision circuit 230 selects the energy-based voice activity detection result, and in the case of a small energy change in the current environment, the voice activity decision circuit 230 selects the pitch-based voice activity detection result.

[0040] Table 1

[0041]

[0042] Figure 5 shows Figure 2Another embodiment of the environmental detection circuit 210 includes a feature extraction circuit 510 and a classification circuit 520. The feature extraction circuit 510 is configured to process the voice input signal according to at least one feature extraction algorithm to generate at least one noise feature. The at least one feature extraction algorithm is a known or self-developed analysis technique, such as Mel-Frequency Cepstral Coefficient (MFCC), Linear Predictive Coding (LPC), Linear Predictive Cepstral Coefficient (LPCC), etc. The classification circuit 520 is configured to determine at least one noise type as the environmental detection result according to the at least one noise feature. For example, the classification circuit 520 obtains the corresponding noise type as the environmental detection result according to the noise feature provided by the feature extraction circuit 510 through a trained statistical model such as a Hidden Markov Model (HMM) and a Gaussian Mixture Model (GMM), or through machine learning methods such as a Support Vector Machine (SVM) and a Neural Network (NN).

[0043] Please refer to Figure 2 and Figure 5 . The voice activity decision circuit 230 selects one of the multiple voice activity detection results according to a preset rule and the at least one noise type. The preset rule is to select one of the multiple voice activity detection results when the noise type is a non-stationary noise type; the preset rule is to select another one of the multiple voice activity detection results when the noise type is a stationary noise type. For example, if the voice activity decision circuit 230 first considers a low miss value and then a low false trigger value, when the noise type is music (non-stationary noise type), the voice activity decision circuit 230 selects the voice activity detection result based on energy; when the noise type is the sound of a fan (stationary noise type), the voice activity decision circuit 230 selects the voice activity detection result based on pitch.

[0044] Figure 6 Another embodiment of the voice activity detection device of the present disclosure is shown, which can select one of multiple voice activity detection algorithms according to the environmental detection result, and thus generate a voice activity detection result as the basis for whether there is voice activity according to the selected voice activity detection algorithm. Figure 6The voice activity detection device 600 includes an environment detection circuit 610 and a voice activity detection and decision-making circuit 620. These circuits are described below.

[0045] An embodiment of the environment detection circuit 610 is Figure 3 or Figure 5 the environment detection circuit 210. The voice activity detection and decision-making circuit 620 is used to select one of the multiple voice activity detection algorithms as the effective voice activity detection algorithm according to the environment detection result of the environment detection circuit 610, and then analyze the voice input signal according to the effective voice activity detection algorithm to generate a voice activity detection result as the basis for whether there is voice activity. For example, when the environment detection circuit 610 is Figure 3 the environment detection circuit 210, the voice activity detection and decision-making circuit 620 selects one of the multiple voice activity detection algorithms as the effective voice activity detection algorithm according to the preset rule and the change of the L comparison results; the preset rule is that when the change of the L comparison results is greater than the preset change degree, select one of the multiple voice activity detection algorithms (for example: the energy-based voice activity detection algorithm), and the preset rule is that when the change of the L comparison results is less than the preset change degree, select another one of the multiple voice activity detection algorithms (for example: the pitch-based voice activity detection algorithm). Another example is that when the environment detection circuit 610 is Figure 5 the environment detection circuit 210, the voice activity detection and decision-making circuit 620 selects one of the multiple voice activity detection algorithms as the effective voice activity detection algorithm according to the preset rule and the at least one noise type; the preset rule is that when the noise type is a non-stationary noise type, select one of the multiple voice activity detection algorithms (for example: the energy-based voice activity detection algorithm), and the preset rule is that when the noise type is a stationary noise type, select another one of the multiple voice activity detection algorithms (for example: the pitch-based voice activity detection algorithm). It should be noted that the technology of analyzing the voice input signal using the effective voice activity detection algorithm can be a known or self-developed technology.

[0046] Since those with ordinary knowledge in the art can refer to Figure 2 the disclosure of the embodiment of Figure 6 to understand Figure 2 the details and variations of the embodiment of Figure 6 that is, the technical features of the embodiment of

[0047] Figure 7 An embodiment showing the voice activity detection method of the present disclosure is from Figure 2 the voice activity detection device 200 ofFigure 6 is performed by the voice activity detection device 600. Figure 7 The voice activity detection method includes the following steps:

[0048] Step S710: Receive and process the sound input signal to generate the environmental detection result; and

[0049] Step S720: Select one of multiple voice activity detection results as the final voice activity detection result according to the environmental detection result, or select one of multiple voice activity detection algorithms according to the environmental detection result and generate the final voice activity detection result accordingly, where the multiple voice activity detection results are respectively generated according to the multiple voice activity detection algorithms.

[0050] Since those with ordinary knowledge in the art can refer to Figure 2 and Figure 6 the disclosure of the embodiments of Figure 7 to understand Figure 2 and Figure 6 the details and variations of the embodiments of Figure 7 That is, the technical features of the embodiments of

[0051] Please note that on the premise that implementation is possible, those with ordinary knowledge in the technical field can selectively implement some or all of the technical features in any of the foregoing embodiments, or selectively implement the combination of some or all of the technical features in multiple of the foregoing embodiments, thereby increasing the flexibility when implementing the present invention.

[0052] In summary, the present invention can respectively adopt different voice activity detection results / algorithms in response to different environmental conditions, so as to achieve good voice activity detection performance under different environmental conditions.

[0053] Although the embodiments of the present invention are as described above, these embodiments are not used to limit the present invention. Those with ordinary knowledge in the technical field can change the technical features of the present invention according to the explicit or implicit content of the present invention. All such changes may fall within the scope of patent protection sought by the present invention. In other words, the scope of patent protection of the present invention shall be determined by the scope of patent application defined in this specification.

[0054] Description of reference numerals:

[0055] S110 - S150: Steps

[0056] 200: Voice activity detection device

[0057] 210: Environmental detection circuit

[0058] 220: Voice activity detection circuit

[0059] 230: Voice Activity Decision Circuit

[0060] 310: Signal Analysis Circuit

[0061] 320: Energy Change Detection Circuit

[0062] 330: Change Information Decision Circuit

[0063] S410~S440: Steps

[0064] 510: Feature Extraction Circuit

[0065] 520: Classification Circuit

[0066] 600: Voice Activity Detection Device

[0067] 610: Environment Detection Circuit

[0068] 620: Voice Activity Detection and Decision Circuit

[0069] S710~S720: Steps

Claims

1. A voice activity detection device capable of selecting one of multiple voice activity detection results as the basis for whether there is voice activity according to the environmental detection result, the voice activity detection device comprises: An environmental detection circuit for processing a sound input signal to generate the environmental detection result, wherein the environmental detection circuit includes: A signal analysis circuit for generating M processing signals for each of the L frames based on the sound input signal, wherein the M processing signals are M band signals or M frequency domain signals, M is a positive integer, and L is the number of frames; An energy change detection circuit for calculating based on the M processing signals for each of the L frames to generate X energy change values for the L frames, wherein X is equal to M multiplied by L; and A change information decision circuit for processing the X energy change values to generate L energy change detection values, then comparing each of the L energy change detection values with a change threshold to generate L comparison results, and then generating the environmental detection result based on the L comparison results; A voice activity detection circuit for analyzing the sound input signal according to multiple voice activity detection algorithms to generate the multiple voice activity detection results, the multiple voice activity detection results including an energy-based voice activity detection result and a pitch-based voice activity detection result; and A voice activity decision circuit for selecting one of the energy-based voice activity detection result and the pitch-based voice activity detection result as the final voice activity detection result according to the environmental detection result, wherein the energy change detection circuit performs multiple steps, including: Calculating based on the M processing signals for each of the L frames to obtain X signal energy values; Calculating X short-term energy values based on the X signal energy values and the number of short-term frames, and calculating X long-term energy values based on the X signal energy values and the number of long-term frames; Obtaining X energy relationship values based on the X short-term energy values and the X long-term energy values; and Comparing each of the X energy relationship values with an energy threshold to generate the X energy change values.

2. The voice activity detection device according to claim 1, wherein the signal analysis circuit includes at least one filtering circuit for generating the M band signals for each of the L frames based on the sound input signal, or the signal analysis circuit includes at least one conversion circuit for generating the M frequency domain signals for each of the L frames based on the sound input signal.

3. The voice activity detection device according to claim 1, wherein the change information decision circuit adds the M energy change values for each of the L frames among the X energy change values to generate the L energy change detection values.

4. The voice activity detection device according to claim 1, wherein the voice activity decision circuit selects one of the multiple voice activity detection results according to a preset rule and the change of the L comparison results; the preset rule is that when the change of the L comparison results is greater than a preset change degree, the voice activity detection result based on energy among the multiple voice activity detection results is selected, and the preset rule is that when the change of the L comparison results is less than the preset change degree, the voice activity detection result based on pitch among the multiple voice activity detection results is selected.

5. The voice activity detection device according to claim 1, wherein the environment detection circuit comprises: a feature extraction circuit for processing the sound input signal according to at least one feature extraction algorithm to generate at least one noise feature; and a classification circuit for determining at least one noise type as the environment detection result according to the at least one noise feature.

6. The voice activity detection device according to claim 5, wherein the voice activity decision circuit selects one of the multiple voice activity detection results according to a preset rule and the at least one noise type; the preset rule is that when the noise type is a non-steady noise type, the voice activity detection result based on energy among the multiple voice activity detection results is selected, and the preset rule is that when the noise type is a steady noise type, the voice activity detection result based on pitch among the multiple voice activity detection results is selected.

7. A voice activity detection device capable of selecting one of multiple voice activity detection algorithms according to the environment detection result, the voice activity detection device comprises: an environment detection circuit for processing the sound input signal to generate the environment detection result, wherein the environment detection circuit comprises: a signal analysis circuit for generating M processing signals for each of the L sound frames based on the sound input signal, wherein the M processing signals are M frequency band signals or M frequency domain signals, M is a positive integer, and L is the number of sound frames; an energy change detection circuit for calculating based on the M processing signals for each of the L sound frames to generate X energy change values for the L sound frames, wherein X is equal to M multiplied by L; and a change information decision circuit for processing the X energy change values to generate L energy change detection values, then comparing each of the L energy change detection values with a change threshold to generate L comparison results, and further generating the environment detection result according to the L comparison results; and a voice activity detection and decision circuit for selecting one of the multiple voice activity detection algorithms including a voice activity detection algorithm based on energy and a voice activity detection algorithm based on pitch as an effective voice activity detection algorithm according to the environment detection result, and then analyzing the sound input signal according to the effective voice activity detection algorithm to generate a voice activity detection result as a basis for whether there is voice activity, wherein the energy change detection circuit performs multiple steps, including: calculating based on the M processing signals for each of the L sound frames to obtain X signal energy values; Calculating X short-term energy values based on the X signal energy values and the number of short-term frames, and calculating X long-term energy values based on the X signal energy values and the number of long-term frames; Obtaining X energy relationship values based on the X short-term energy values and the X long-term energy values; and Comparing each of the X energy relationship values with an energy threshold to generate the X energy change values.

8. A voice activity detection method comprising: Receiving and processing a sound input signal to generate an environmental detection result, wherein generating the environmental detection result includes: Generating M processing signals for each of the L frames based on the sound input signal, where the M processing signals are M frequency band signals or M frequency domain signals, M is a positive integer, and L is the number of frames; Calculating based on the M processing signals for each of the L frames to obtain X signal energy values, where X is equal to M multiplied by L; Calculating X short-term energy values based on the X signal energy values and the number of short-term frames, and calculating X long-term energy values based on the X signal energy values and the number of long-term frames; Obtaining X energy relationship values based on the X short-term energy values and the X long-term energy values; Comparing each of the X energy relationship values with an energy threshold to generate the X energy change values; and Processing the X energy change values to generate L energy change detection values, then comparing each of the L energy change detection values with a change threshold to generate L comparison results, and then generating the environmental detection result based on the L comparison results; and Selecting one of multiple voice activity detection results as the final voice activity detection result based on the environmental detection result, the multiple voice activity detection results including an energy-based voice activity detection result and a pitch-based voice activity detection result, or selecting one of multiple voice activity detection algorithms based on the environmental detection result and generating the final voice activity detection result accordingly, the multiple voice activity detection algorithms including an energy-based voice activity detection algorithm and a pitch-based voice activity detection algorithm, wherein the multiple voice activity detection results are respectively generated based on the multiple voice activity detection algorithms.

Citation Information

Patent Citations

  • Intelligent audio device and method, electronic equipment and computer readable medium

    CN111145752A

  • Voice activity detection method and device, and readable storage medium

    CN111292758A