Voice detection method and apparatus, device, and storage medium
By using the first and second audio features of the audio sequence in speech detection, combining average energy, energy ratio, zero crossing rate and spectrum modulation energy, the problem of traditional technology being difficult to detect speech in music context is solved, and speech detection with high precision, low computing power and no large amount of training data is achieved.
Patent Information
- Application Number
- PCT/CN2023/134703
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
Traditional speech detection technology is difficult to accurately detect speech in audio sequences with music background sounds, and deep learning-based methods require a large amount of training data and a large number of model parameters, which is poor in judgment of unknown data.
By acquiring the audio sequence, the first and second audio features are extracted, and speech detection is performed using the average energy, energy ratio and zero crossing rate, and the spectrum modulation energy, respectively, and the final detection result is determined based on the results of the two.
It realizes speech detection in steady-state noise, transient noise and music, without a large amount of training data, low computing power and high detection accuracy.
Smart Images

Figure CN2023134703_05062025_PF_FP_ABST
Abstract
Description
Voice detection method, device, equipment and storage medium Technical Field
[0001] The present application relates to the field of speech detection technology, and in particular to a speech detection method, apparatus, device and storage medium. Background Art
[0002] Voice Activity Detection (VAD) is widely used in speech processing such as call noise reduction, intelligent speech, voiceprint segmentation and clustering, and speech coding. VAD usually distinguishes between silent segments and speech segments in an audio stream, but cannot distinguish between music segments and speech segments. However, there is also a strong application demand for distinguishing between music and speech. For example, one application is to encode music clips and speech clips differently to achieve a balance between transmission efficiency and audio quality; another application is to detect whether there is speech from the audio stream in real time and respond to the detection results. If traditional VAD is used, some music, instrument sounds, and transient noises may be misjudged and the wrong instructions may be executed.
[0003] Traditional speech detection solutions, based on features such as energy, zero-crossing rate, and spectral entropy, struggle to detect speech fragments from audio sequences with background music. Deep learning-based VADs, due to their ability to automatically learn features, can effectively distinguish between music, speech, silence, and background noise. However, these methods require large amounts of training data and a large number of model parameters. Using small models can be ineffective, and they are not very effective for detecting unknown data.
[0004] Summary of the Invention
[0005] The present application provides a speech detection method, apparatus, device and storage medium, which can realize speech detection in a non-training manner, with low computing power and high detection accuracy.
[0006] To solve the above technical problems, a technical solution adopted by this application is to provide a voice detection method, comprising:
[0007] Get the audio sequence;
[0008] Extracting a first audio feature from the audio sequence, and performing speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result;
[0009] Extracting a second audio feature from the audio sequence, and performing speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result;
[0010] A speech detection result of the audio sequence is determined according to the first speech detection result and the second speech detection result.
[0011] According to one embodiment of the present application, the first audio feature includes average energy, energy ratio, and zero-crossing rate of the audio signal; and extracting the first audio feature from the audio sequence, and performing speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result includes:
[0012] Performing sampling rate conversion and frame processing on the audio sequence to obtain a plurality of frames of audio signals;
[0013] Calculating the average energy and the zero-crossing rate of the audio signal of one frame according to the audio signal of each frame;
[0014] Obtaining an energy spectrum of the audio signal, obtaining low-frequency band energy and high-frequency band energy according to the energy spectrum, and calculating a ratio between an average energy of the low-frequency band energy and an average energy of the high-frequency band energy to obtain the energy ratio;
[0015] Speech detection is performed on the audio sequence according to the average energy, the zero-crossing rate, and the energy ratio to obtain a first speech detection result.
[0016] According to one embodiment of the present application, acquiring the energy spectrum of the audio signal and obtaining the low-frequency band energy and the high-frequency band energy according to the energy spectrum includes:
[0017] Obtaining low-frequency band energy and high-frequency band energy from the frequency domain through Fourier transform, or respectively obtaining a low-frequency signal and a high-frequency signal through a time domain filter and a preset cutoff frequency, and calculating the low-frequency band energy of the low-frequency signal and the high-frequency band energy of the high-frequency signal; wherein obtaining the low-frequency band energy and the high-frequency band energy from the frequency domain through Fourier transform includes:
[0018] Performing windowing processing on the audio signal of each frame;
[0019] Perform fast Fourier transform processing on the windowing processing result;
[0020] Calculate the energy spectrum based on the fast Fourier transform processing results;
[0021] The high-frequency band energy and the low-frequency band energy are counted from the energy spectrum.
[0022] According to one embodiment of the present application, performing speech detection on the audio sequence according to the average energy, the zero-crossing rate, and the energy ratio to obtain a first speech detection result includes:
[0023] comparing the average energy with a first preset threshold;
[0024] comparing the energy ratio with a second preset threshold;
[0025] comparing the zero-crossing rate with a third preset threshold;
[0026] When the average energy is greater than a first preset threshold, the energy ratio is greater than a second preset threshold, and the zero-crossing rate is greater than a third preset threshold, the first speech detection result is that the audio sequence is speech.
[0027] According to one embodiment of the present application, the second feature includes spectrum modulation energy; and extracting the second audio feature from the audio sequence, and performing speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result includes:
[0028] Performing sampling rate conversion and segmentation on the audio sequence to obtain a plurality of audio segments;
[0029] Calculating a mel-spectrogram for each of the audio clips to obtain a mel-spectrogram containing multiple channels;
[0030] Performing Fourier transform processing on each of the channels in the mel-spectrogram, and calculating the normalized modulation energy of each channel;
[0031] Speech detection is performed on the audio sequence according to the normalized modulation energy of each of the channels to obtain a second speech detection result.
[0032] According to one embodiment of the present application, performing speech detection on the audio sequence according to the normalized modulation energy of each of the channels to obtain a second speech detection result includes:
[0033] Calculating the sum of the normalized modulation energies of the channels;
[0034] comparing the calculation result with a fourth preset threshold;
[0035] If the calculation result is greater than the fourth preset threshold, the second speech detection result is that the audio sequence is speech;
[0036] If the calculation result is less than or equal to the fourth preset threshold, the second speech detection result is that the audio sequence is non-speech.
[0037] According to one embodiment of the present application, determining the speech detection result of the audio sequence according to the first speech detection result and the second speech detection result includes:
[0038] Determining whether the first voice detection result and the second voice detection result are both voice;
[0039] If so, determining that the speech detection result is that the audio sequence is speech;
[0040] If not, it is determined that the speech detection result is that the audio sequence is non-speech.
[0041] To solve the above technical problems, another technical solution adopted by the present application is to provide a speech detection device, comprising:
[0042] Acquisition module, used to obtain audio sequence;
[0043] a first audio feature extraction module, configured to extract a first audio feature from the audio sequence, and perform speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result;
[0044] a second audio feature extraction module, configured to extract a second audio feature from the audio sequence, and perform speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result;
[0045] A speech detection module is used to determine a speech detection result of the audio sequence according to the first speech detection result and the second speech detection result.
[0046] To solve the above technical problems, another technical solution adopted in this application is: to provide a computer device, including: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech detection method when executing the computer program.
[0047] In order to solve the above technical problems, another technical solution adopted in the present application is: providing a computer storage medium on which a computer program is stored, and the computer program implements the above speech detection method when executed by a processor.
[0048] The beneficial effects of the present application are: by obtaining an audio sequence; extracting a first audio feature of the audio sequence, and performing speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result; extracting a second audio feature of the audio sequence, and performing speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result; determining the speech detection result of the audio sequence based on the first speech detection result and the second speech detection result, speech detection can be achieved from steady-state noise, transient noise and music in a non-training manner, without the need for a large amount of training data, with low computing power and high detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a flow chart of a voice detection method according to an embodiment of the present application;
[0050] FIG2 is a flow chart of step S20 in the voice detection method according to an embodiment of the present application;
[0051] FIG3 is a flow chart of step S203 in the voice detection method according to an embodiment of the present application;
[0052] FIG4 is a flow chart of step S204 in the voice detection method according to an embodiment of the present application;
[0053] FIG5 is a flow chart of step S30 in the voice detection method according to an embodiment of the present application;
[0054] FIG6 is a schematic structural diagram of a speech detection device according to an embodiment of the present application;
[0055] FIG7 is a schematic diagram of the structure of a computer device according to an embodiment of the present application;
[0056] FIG8 is a schematic diagram of the structure of a computer storage medium according to an embodiment of the present application. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0058] The terms "first," "second," and "third" in this application are used only for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of this application, "multiple" means at least two, for example, two, three, etc., unless otherwise specifically defined. All directional indications in the embodiments of this application (such as up, down, left, right, front, back...) are only used to explain the relative positional relationship, movement, etc. between the components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications also change accordingly. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products, or devices.
[0059] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0060] FIG1 is a flow chart of a speech detection method according to an embodiment of the present application. It should be noted that the method of the present application is not limited to the flow sequence shown in FIG1 if substantially the same results are achieved. As shown in FIG1 , the method includes the following steps:
[0061] Step S10: Acquire an audio sequence.
[0062] In step S10, the audio sequence may include one or more audio signals selected from background noise, music, and speech. Background noise may include steady-state noise and / or transient noise. Exemplarily, the audio sequence includes background noise, music, and speech. Furthermore, the audio sequence includes background noise and speech. Furthermore, the audio sequence includes music and speech. Furthermore, the audio sequence includes background noise and / or music.
[0063] Step S20: extracting a first audio feature from the audio sequence, and performing speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result.
[0064] In step S20, the first audio feature may include the average energy, energy ratio, and zero-crossing rate of the audio signal. This embodiment can detect speech from steady-state noise using the first audio feature. The first speech detection result includes two types: one indicating that the audio sequence is speech and the other indicating that the audio sequence is non-speech.
[0065] In one feasible implementation, referring to FIG2 , step S20 further includes the following steps:
[0066] Step S201: performing sampling rate conversion and frame division processing on the audio sequence to obtain a plurality of frames of audio signals.
[0067] Specifically, the sampling frequency of the audio sequence is converted to 8kHz, and the audio sequence after the sampling frequency conversion is framed, with each frame having 256 sample points and no overlap between frames, to obtain several frames of audio signals.
[0068] Step S202: Calculate the average energy and zero-crossing rate of an audio frame according to each frame of audio signal.
[0069] Specifically, the average energy of a frame of audio signal is calculated according to the following formula:
[0070] Among them, x k is the k-th frame audio signal, with a length of 256, i is the number of sample points, and energy(k) is the average energy of the k-th frame audio signal.
[0071] The zero-crossing rate is calculated according to the following formula: zcr = mean(abs(diff(sign(inputFrame)))), where zcr is the zero-crossing rate, mean, abs, diff, and sign are the average, absolute value, difference, and sign functions of the MATLAB program, respectively, inputFrame is a frame of audio signal, and length is the frame length.
[0072] Step S203: Obtain an energy spectrum of the audio signal, obtain low-frequency band energy and high-frequency band energy according to the energy spectrum, and calculate the ratio between the average energy of the low-frequency band energy and the average energy of the high-frequency band energy to obtain an energy ratio.
[0073] Specifically, the low-frequency band range may be 200 Hz to 1000 Hz, and the high-frequency band range may be 1000 Hz to 4000 Hz. This embodiment may obtain the low-frequency band energy and the high-frequency band energy from the frequency domain through Fourier transform, or obtain the low-frequency signal and the high-frequency signal respectively through a time domain filter and a preset cutoff frequency, and calculate the low-frequency band energy of the low-frequency signal and the high-frequency band energy of the high-frequency signal.
[0074] In one feasible implementation, referring to FIG3 , obtaining low-frequency band energy and high-frequency band energy from the frequency domain through Fourier transform further includes the following steps:
[0075] Step S2031: performing windowing processing on each frame of audio signal.
[0076] Specifically, windowing involves multiplying each frame of audio signal by a Hanning window. This increases the continuity between the left and right ends of the frame. The Hanning window effectively reduces signal leakage during the windowing process. The windowed audio signal is converted into an energy distribution in the frequency domain, and different energy distributions can represent the characteristics of different speech sounds.
[0077] Step S2032: performing fast Fourier transform processing on the windowing processing result.
[0078] Specifically, fast Fourier transform is performed on the windowing result to obtain a frequency spectrum.
[0079] Step S2033: Calculate the energy spectrum according to the fast Fourier transform processing result.
[0080] Specifically, the energy spectrum, or energy spectral density, can characterize the distribution of energy of a signal or time series over frequency. In one embodiment, the energy spectrum is the square of the fast Fourier transform.
[0081] Step S2034: Count the high-frequency band energy and the low-frequency band energy from the energy spectrum.
[0082] Step S2035: Calculate the ratio between the average energy of the low-frequency band energy and the average energy of the high-frequency band energy to obtain an energy ratio.
[0083] Specifically, the average energy of the low-frequency band energy and the average energy of the high-frequency band energy are first calculated, and then the ratio between the average energy of the low-frequency band energy and the average energy of the high-frequency band energy is calculated to obtain the energy ratio.
[0084] Step S204: performing speech detection on the audio sequence according to the average energy, the zero-crossing rate, and the energy ratio to obtain a first speech detection result.
[0085] Specifically, the average energy is compared with a first preset threshold; the energy ratio is compared with a second preset threshold; and the zero-crossing rate is compared with a third preset threshold. When the average energy is greater than the first preset threshold, the energy ratio is greater than the second preset threshold, and the zero-crossing rate is greater than the third preset threshold, the first speech detection result is that the audio sequence is speech; otherwise, the first speech detection result is that the audio sequence is non-speech. The first preset threshold, the second preset threshold, and the third preset threshold of this embodiment can be adjusted according to different application scenarios and can be fixed values or value ranges. For example, referring to FIG4, step S2041 is first performed: whether the average energy is greater than the first preset threshold; if so, step S2042 is performed: whether the energy ratio is greater than the second preset threshold; if so, step S2043 is performed: whether the zero-crossing rate is greater than the third preset threshold; if so, "Decision 1 = 1" is output; after step S2041, if not, "Decision 1 = 0" is output; after step S2042, if not, "Decision 1 = 0" is output; after step S2043, if not, "Decision 1 = 0" is output. In this embodiment, "Decision1=1" indicates that the first speech detection result is that the audio sequence is speech, and "Decision1=0" indicates that the first speech detection result is that the audio sequence is non-speech.
[0086] Step S30: extracting a second audio feature from the audio sequence, and performing speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result.
[0087] In step S30, the second audio feature may include spectral modulation energy, for example, spectral modulation energy between 2 Hz and 9 Hz. This embodiment can detect speech from transient noise and music using the second audio feature. The second speech detection result includes two types: one indicating that the audio sequence is speech and the other indicating that the audio sequence is non-speech.
[0088] In one feasible implementation, referring to FIG5 , step S30 further includes the following steps:
[0089] Step S301: performing sampling rate conversion and segmentation processing on an audio sequence to obtain a number of audio segments.
[0090] Specifically, the sampling frequency of the audio sequence is converted to 8kHz, and the audio signal is divided into several segments, each of which is 1.022s long (i.e., 8176 sample points, 8kHz sampling rate) and has a step size of 10ms (i.e., 80 sample points, 8kHz sampling rate).
[0091] Step S302: Calculate the mel-spectrogram for each audio clip to obtain a mel-spectrogram containing multiple channels.
[0092] Specifically, each audio clip is windowed and Fast Fourier Transformed (FFT) to obtain a mel-spectrogram. The window length is 256 (32ms), and the Hanning window function is used. The Hanning window effectively reduces signal leakage during the windowing process. The FFT length is 256, the overlap length is 256-80 = 176, and the number of channels is 40. The resulting mel-spectrogram is a (40, 100) matrix.
[0093] Step S303: Perform Fourier transform processing on each channel in the mel-spectrogram, and calculate the normalized modulation energy of each channel.
[0094] Specifically, Fourier transform processing is performed on each channel in the mel spectrogram, and the ratio of the spectrum modulation energy of 2 Hz to 9 Hz to the total energy is calculated to obtain the normalized modulation energy of 2 Hz to 9 Hz of each channel.
[0095] Step S304: performing speech detection on the audio sequence according to the normalized modulation energy of each channel to obtain a second speech detection result.
[0096] Specifically, a comprehensive judgment decision is made based on the normalized modulation energy of 2Hz to 9Hz for 40 channels. In one achievable implementation, the sum of the normalized modulation energy of each channel is calculated; the calculated result is compared with a fourth preset threshold; if the calculated result is greater than the fourth preset threshold, the second speech detection result is determined to be speech; if the calculated result is less than or equal to the fourth preset threshold, the second speech detection result is determined to be non-speech.
[0097] Step S40: Determine a speech detection result of the audio sequence according to the first speech detection result and the second speech detection result.
[0098] In step S40, it is determined whether the first voice detection result and the second voice detection result are both voice; if so, the voice detection result is determined to be that the audio sequence is voice; if not, the voice detection result is determined to be that the audio sequence is non-voice. Exemplarily, if the first voice detection result is that the audio sequence is voice and the second voice detection result is that the audio sequence is voice, the voice detection result is determined to be that the audio sequence is voice. Exemplarily, if the first voice detection result is that the audio sequence is voice and the second voice detection result is that the audio sequence is non-voice, the voice detection result is determined to be that the audio sequence is non-voice. Exemplarily, if the first voice detection result is that the audio sequence is non-voice and the second voice detection result is that the audio sequence is voice, the voice detection result is determined to be that the audio sequence is non-voice. If the first voice detection result is that the audio sequence is non-voice and the second voice detection result is that the audio sequence is non-voice, the voice detection result is determined to be that the audio sequence is non-voice.
[0099] The speech detection method of an embodiment of the present application obtains an audio sequence; extracts a first audio feature from the audio sequence, and performs speech detection on the audio sequence based on the first audio feature to obtain a first speech detection result; extracts a second audio feature from the audio sequence, and performs speech detection on the audio sequence based on the second audio feature to obtain a second speech detection result; and determines the speech detection result of the audio sequence based on the first speech detection result and the second speech detection result. This method can realize speech detection from steady-state noise, transient noise, and music in a non-training manner, does not require a large amount of training data, has low computing power, and has high detection accuracy.
[0100] The embodiment of the present application further discloses a speech detection device, as shown in FIG6 , which includes: an acquisition module 61 , a first audio feature extraction module 62 , a second audio feature extraction module 63 and a speech detection module 64 .
[0101] The acquisition module 61 is used to acquire an audio sequence.
[0102] The first audio feature extraction module 62 is coupled to the acquisition module 61 and is configured to extract a first audio feature from the audio sequence and perform speech detection on the audio sequence according to the first audio feature to obtain a first speech detection result.
[0103] The second audio feature extraction module 63 is coupled to the acquisition module 61 and is configured to extract the second audio feature from the audio sequence and perform speech detection on the audio sequence according to the second audio feature to obtain a second speech detection result.
[0104] The speech detection module 64 is coupled to the first audio feature extraction module 62 and the second audio feature extraction module 63 respectively, and is configured to determine a speech detection result of the audio sequence according to the first speech detection result and the second speech detection result.
[0105] Please refer to Figure 7, which is a schematic diagram of the structure of a computer device according to an embodiment of the present application. As shown in Figure 7, the computer device 70 includes a processor 71 and a memory 72 coupled to the processor 71.
[0106] The memory 72 stores program instructions for implementing the speech detection described in any of the above embodiments.
[0107] The processor 71 is configured to execute program instructions stored in the memory 72 to detect speech.
[0108] The processor 71 may also be referred to as a CPU (Central Processing Unit). The processor 71 may be an integrated circuit chip having signal processing capabilities. The processor 71 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.
[0109] Refer to Figure 8, which is a schematic diagram of the structure of the computer storage medium of an embodiment of the present application. The computer storage medium of the embodiment of the present application stores a program file 81 that can implement all the above methods, wherein the program file 81 can be stored in the above-mentioned computer storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned computer storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0110] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0111] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0112] The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A voice detection method, characterized in that, it includes: Obtain an audio sequence; Perform first audio feature extraction on the audio sequence, and perform voice detection on the audio sequence according to the first audio feature to obtain a first voice detection result; Perform second audio feature extraction on the audio sequence, and perform voice detection on the audio sequence according to the second audio feature to obtain a second voice detection result; Determine the voice detection result of the audio sequence according to the first voice detection result and the second voice detection result.
2. The voice detection method according to claim 1, characterized in that, The first audio feature includes the average energy, energy ratio, and zero-crossing rate of the audio signal; the performing first audio feature extraction on the audio sequence and performing voice detection on the audio sequence according to the first audio feature to obtain a first voice detection result includes: Perform sampling rate conversion and frame segmentation processing on the audio sequence to obtain several frames of audio signals; Calculate the average energy and the zero-crossing rate of one frame of the audio signal according to each frame of the audio signal; Obtain the energy spectrum of the audio signal, obtain the low-frequency band energy and the high-frequency band energy according to the energy spectrum, and calculate the ratio between the average energy of the low-frequency band energy and the average energy of the high-frequency band energy to obtain the energy ratio; Perform voice detection on the audio sequence according to the average energy, the zero-crossing rate, and the energy ratio to obtain a first voice detection result.
3. The voice detection method according to claim 2, characterized in that, The obtaining the energy spectrum of the audio signal and obtaining the low-frequency band energy and the high-frequency band energy according to the energy spectrum includes: Obtain the low-frequency band energy and the high-frequency band energy from the frequency domain through Fourier transform, or obtain the low-frequency signal and the high-frequency signal through a time-domain filter and a preset cut-off frequency respectively, and calculate the low-frequency band energy of the low-frequency signal and the high-frequency band energy of the high-frequency signal; wherein, the obtaining the low-frequency band energy and the high-frequency band energy from the frequency domain through Fourier transform includes: Perform windowing processing on each frame of the audio signal respectively; Perform fast Fourier transform processing on the windowing processing result; Calculate the energy spectrum according to the fast Fourier transform processing result; Statistically obtain the high-frequency band energy and the low-frequency band energy from the energy spectrum.
4. The voice detection method according to claim 2, characterized in that, The performing voice detection on the audio sequence according to the average energy, the zero-crossing rate, and the energy ratio to obtain a first voice detection result includes: Compare the average energy with a first preset threshold; Compare the energy ratio with a second preset threshold; Compare the zero-crossing rate with a third preset threshold; When the average energy is greater than the first preset threshold, the energy ratio is greater than the second preset threshold, and the zero-crossing rate is greater than the third preset threshold are satisfied simultaneously, the first voice detection result is that the audio sequence is voice.
5. The voice detection method according to claim 4, characterized in that, The second feature includes spectral modulation energy; the second audio feature extraction of the audio sequence and the voice detection of the audio sequence according to the second audio feature to obtain a second voice detection result include: Performing a sampling rate conversion process and a segmentation process on the audio sequence to obtain a plurality of audio segments; Calculating the Mel spectrum for each of the audio segments to obtain a Mel spectrogram including multiple channels; Performing a Fourier transform process on each of the channels in the Mel spectrogram and calculating the normalized modulation energy of each of the channels; Performing voice detection on the audio sequence according to the normalized modulation energy of each of the channels to obtain a second voice detection result.
6. The voice detection method according to claim 5, wherein, the performing voice detection on the audio sequence according to the normalized modulation energy of each of the channels to obtain a second voice detection result includes: Calculating the sum of the normalized modulation energies of each of the channels; Comparing the calculation result with a fourth preset threshold; If the calculation result is greater than the fourth preset threshold, the second voice detection result is that the audio sequence is voice; If the calculation result is less than or equal to the fourth preset threshold, the second voice detection result is that the audio sequence is non-voice.
7. The voice detection method according to claim 6, wherein, the determining the voice detection result of the audio sequence according to the first voice detection result and the second voice detection result includes: Judging whether both the first voice detection result and the second voice detection result are voice; If so, determining that the voice detection result is that the audio sequence is voice; If not, determining that the voice detection result is that the audio sequence is non-voice.
8. A voice detection device, wherein, it includes: An acquisition module, configured to acquire an audio sequence; A first audio feature extraction module, configured to perform first audio feature extraction on the audio sequence and perform voice detection on the audio sequence according to the first audio feature to obtain a first voice detection result; A second audio feature extraction module, configured to perform second audio feature extraction on the audio sequence and perform voice detection on the audio sequence according to the second audio feature to obtain a second voice detection result; A voice detection module, configured to determine the voice detection result of the audio sequence according to the first voice detection result and the second voice detection result.
9. A computer device, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the voice detection method according to any one of claims 1-7 is implemented.
10. A computer storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the voice detection method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Sound activity detecting method and detector thereof
CN101197130A
Voice endpoint detection method, device and equipment and storage medium
CN110335593A
Breath sound detection method based on time-frequency characteristics, breath sound detection system based on time-frequency characteristics, equipment and medium
CN110473563A
Self-adaptive voice activity detection method and device, equipment and storage medium
CN114495907A
Acoustic event detection method and device, electronic equipment and storage medium
CN114898737A