Voice detection method and related device thereof

By using a multi-microphone VAD and wind noise detection method, the problems of speech distortion and inaccurate recognition in speech detection are solved, achieving high-accuracy speech detection without affecting speech quality. It is applicable to various electronic devices including mobile phones and smart screens.

CN117995225BActive Publication Date: 2025-11-07HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211350590.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-11-07
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

Existing technologies for speech detection suffer from several problems: when noise reduction is effective, speech distortion occurs; the limited training samples for neural network models lead to inaccurate recognition; and small electronic devices cannot effectively reduce wind noise, affecting the accuracy and quality of speech detection.

Method used

By combining audio signals acquired from multiple microphones, VAD and wind noise detection are performed. The voice signal is first distinguished from other signals before wind noise detection is performed to ensure the accuracy of the voice signal.

Benefits of technology

It improves the accuracy of speech detection, avoids affecting speech quality, requires no hardware improvements, and is suitable for small electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117995225B_ABST
    Figure CN117995225B_ABST
Patent Text Reader

Abstract

The application provides a voice detection method and a related device, and relates to the field of audio processing. The voice detection method comprises the following steps: acquiring audio data, wherein the audio data is data collected by a first microphone and a second microphone in the same environment; performing VAD detection on the audio data to determine and filter out voice signals; and performing wind noise detection on the voice signals detected by the VAD to determine and filter out voice signals. The application combines the multi-path audio signals obtained by the multiple microphones to perform VAD detection and wind noise detection, which can avoid affecting the voice quality and improve the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of audio processing, in particular to a voice detection method and a related device thereof. BACKGROUND

[0002] With the popularity and development of electronic devices, electronic devices have become an indispensable part of our daily life and entertainment. Usually, during voice communication or voice operation, the input audio data of the electronic device may be affected due to the interference of external sounds. Therefore, in order to improve the quality of audio, the electronic device needs to process the input audio data.

[0003] In related technologies, noise reduction and voice recognition using a neural network model are usually performed. However, when the noise reduction effect is good, the voice may be distorted. The neural network model needs to be trained in advance, and the sample is usually limited, which may result in inaccurate voice recognition and affect the quality of detection. Therefore, a new voice detection method is needed, which can avoid affecting the quality of voice and improve the accuracy of detection. SUMMARY

[0004] The present application provides a voice detection method and a related device thereof, which combines multiple audio signals obtained by multiple microphones to perform VAD detection and wind noise detection, thereby avoiding affecting the quality of voice and improving the accuracy of detection.

[0005] In a first aspect, a voice detection method is provided, applied to an electronic device including a first microphone and a second microphone, and the method includes:

[0006] obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment;

[0007] performing VAD detection on the audio data to determine and select voice signals;

[0008] performing wind noise detection on the voice signals detected by the VAD to determine and select voice signals.

[0009] In the embodiments of the present application, in the process that a user uses an electronic device including multiple microphones to make a voice call or voice operation, the electronic device can first perform VAD detection on audio data received by the multiple microphones to distinguish voice signals and other signals; then, wind noise detection is performed again on the filtered voice signals, which is equivalent to filtering the voice signals again, so that the real voice signals and wind noise signals mistaken as voice signals can be distinguished, and the voice signals detected by the wind noise detection are the final detection results. Therefore, in combination with the to-be-detected signals generated by the multiple microphones, the real voice signals, wind noise signals and other signals can be distinguished through the VAD and wind noise detection in two stages. Such a simple detection method does not involve hardware changes, can avoid affecting the voice quality, and can improve the detection accuracy.

[0010] In the present application, the other signals refer to signals other than voice signals and wind noise signals.

[0011] In combination with the first aspect, in an implementation manner of the first aspect, when the audio data is data located in the time domain, the method further includes:

[0012] The audio data is preprocessed, and the preprocessing at least includes framing and time-frequency conversion.

[0013] Optionally, the preprocessing at least includes framing and time-frequency conversion.

[0014] It should be understood that after the multiple to-be-detected signal streams are framed with the same length, the number of the multiple frames of first time domain signals and the multiple frames of second time domain signals is the same, and they have a one-to-one corresponding relationship in sequence. Therefore, after the multiple frames of first time domain signals and the multiple frames of second time domain signals are converted in the frequency domain, the number of the multiple frames of first frequency domain signals and the multiple frames of second frequency domain signals is also the same, and they also have a one-to-one corresponding relationship in sequence.

[0015] In the embodiments of the present application, preprocessing can make the audio data convenient for subsequent detection.

[0016] In combination with the first aspect, in an implementation manner of the first aspect, the audio data includes a first to-be-detected signal stream collected by the first microphone and a second to-be-detected signal stream collected by the second microphone;

[0017] The preprocessing of the audio data includes:

[0018] The first to-be-detected signal stream is framed to obtain multiple frames of first time domain signals;

[0019] The multiple frames of first time domain signals are converted in the time-frequency domain to obtain multiple frames of first frequency domain signals;

[0020] frame the second to-be-tested signal stream to obtain a plurality of frames of second time-domain signals;

[0021] perform the time-frequency transformation on the plurality of frames of second time-domain signals to obtain a plurality of frames of second frequency-domain signals;

[0022] wherein the plurality of frames of first time-domain signals and the plurality of frames of first frequency-domain signals correspond to each other, and the plurality of frames of second time-domain signals and the plurality of frames of second frequency-domain signals correspond to each other.

[0023] In the embodiments of the present application, the plurality of frames of first time-domain signals and the plurality of frames of first frequency-domain signals can be obtained according to the first to-be-tested signal stream, and the plurality of frames of second time-domain signals and the plurality of frames of second frequency-domain signals can be obtained according to the second to-be-tested signal stream, so that the same order of multiple signals can be combined for subsequent voice detection.

[0024] With reference to the first aspect, in an implementation form of the first aspect, the VAD detection is performed on the audio data to determine and filter out the voice signal, including:

[0025] For the first time-domain signal, the first data corresponding to the first time-domain signal is determined according to the first time-domain signal and the first frequency-domain signal corresponding to the first time-domain signal, and the first data at least includes a zero-crossing rate, a spectral entropy and a flatness;

[0026] The VAD detection is performed on the first time-domain signal based on the first data to determine and filter out the voice signal.

[0027] In the embodiments of the present application, the voice signal and other signals can be distinguished based on the different performances of the voice signal and other signals in the first data, and then the first time-domain signal can be distinguished as the voice signal or other signals.

[0028] With reference to the first aspect, in an implementation form of the first aspect, the VAD detection is performed on the first time-domain signal based on the first data to determine and filter out the voice signal, including:

[0029] When the first data meets a first condition, it is determined that the tentative state of the first time-domain signal is the voice signal;

[0030] When the first data does not meet the first condition, it is determined that the tentative state of the first time-domain signal is other signals, and the other signals are used to indicate signals other than the voice signal and wind noise signal;

[0031] For the first time-domain signal, it is determined whether the tentative state is the same as the current state;

[0032] When the temporary state is different from the current state and the temporary state is the speech signal, the value of the first frame number flag bit is increased by 1, and it is determined whether the value of the first frame number flag bit is greater than a first preset frame number threshold value;

[0033] When the value of the first frame number flag bit is greater than the first preset frame number threshold value, the current state is modified, when the current state is the speech signal, the current state is modified to be the other signal, and when the current state is the other signal, the current state is modified to be the speech signal;

[0034] When the temporary state is different from the current state and the temporary state is the other signal, the value of the second frame number flag bit is increased by 1, and it is determined whether the value of the second frame number flag bit is greater than a second preset frame number threshold value;

[0035] When the value of the second frame number flag bit is greater than the second preset frame number threshold value, the current state is modified;

[0036] The first time domain signal whose modified current state is the speech signal is determined and screened out.

[0037] Since a speech word usually lasts for several frames and there is an interval between words, in order to completely determine the start and end of a sentence and prevent the sentence from being interrupted, the first time domain signal is provided with a temporary state and a current state. The temporary state and the current state can be divided into three states: a speech signal, a wind noise signal and an other signal.

[0038] In the embodiment of the application, when the temporary state is different from the current state, it indicates that the two times of determination are inconsistent, at this time, at least one time of determination is wrong, therefore, frame number accumulation can be performed. When the frame number accumulation is greater than a frame number threshold value, the corresponding current state is modified, which is equivalent to determining the state of the first time domain signal according to the continuity of the multiple frames of to-be-detected signals before the first time domain signal.

[0039] With reference to the first aspect, in an implementation form of the first aspect, the method further includes:

[0040] When the temporary state is the same as the current state, the first time domain signal whose current state is the speech signal is determined and screened out; or,

[0041] When the temporary state is different from the current state and the value of the first frame number flag bit is less than or equal to the first preset frame number threshold value, the first time domain signal whose current state is the speech signal is determined and screened out; or,

[0042] When the temporary state is different from the current state and the value of the second frame number flag bit is less than or equal to the second preset frame number threshold value, the first time domain signal whose current state is the speech signal is determined and screened out.

[0043] In the embodiments of the present application, when the provisional state is the same as the current state, or although different, when the frame number accumulation is less than the frame number threshold, the corresponding current state is not modified, which is equivalent to preventing the interruption of the statement in the middle, ignoring the abnormality of the short frames and regarding it as a voice signal. Or, it is equivalent to avoiding the error of identifying a small amount of other signals as a voice signal, and regarding it as other signals.

[0044] With reference to the first aspect, in an implementation form of the first aspect, before the first data satisfies the first condition, the method further includes: performing first initialization processing, the first initialization processing at least including zeroing the value of the first frame number flag and the value of the second frame number flag.

[0045] In the embodiments of the present application, by performing the first initialization processing, data errors or interference of some detection results in other stages can be avoided.

[0046] With reference to the first aspect, in an implementation form of the first aspect, when the first data includes the zero-crossing rate, the spectral entropy and the flatness, the first condition includes:

[0047] The zero-crossing rate is greater than a zero-crossing rate threshold, the spectral entropy is less than a spectral entropy threshold, and the flatness is less than a flatness threshold.

[0048] With reference to the first aspect, in an implementation form of the first aspect, the wind noise detection on the voice signal detected by the VAD includes:

[0049] For the first time domain signal detected by the VAD as a voice signal, the second data corresponding to the first time domain signal is determined according to the first time domain signal, the first frequency domain signal corresponding to the first time domain signal, and the second frequency domain signal with the same sequence as the first frequency domain signal, the second data at least including the spectral centroid, the low frequency energy and the correlation;

[0050] The second data is determined, and the wind noise detection is performed on the first time domain signal to determine and select the voice signal.

[0051] In the embodiments of the present application, since the wind noise signal and the speech signal have similar characteristics, the wind noise signal and the speech signal cannot be distinguished accurately after the first-stage VAD detection, and the wind noise signal may be mistaken as the speech signal, that is, the speech signal in the first detection result obtained after the VAD detection is only a suspected speech signal, and may include the wind noise signal. Then, the wind noise detection is continued, and the real speech signal and the false speech signal (i.e., the wind noise signal) can be further distinguished. Therefore, the detection accuracy can be greatly improved after the continuous VAD detection and wind noise detection.

[0052] With reference to the first aspect, in an implementation form of the first aspect, the wind noise detection is performed on the first time-domain signal based on the second data, and the speech signal is determined and filtered, including:

[0053] When the second data satisfies a second condition, the tentative state of the first time-domain signal is determined as the wind noise signal;

[0054] When the second data does not satisfy the second condition, the tentative state of the first time-domain signal is determined as the speech signal;

[0055] For the first time-domain signal, it is determined whether the tentative state is the same as a current state;

[0056] When the tentative state is different from the current state, and the tentative state is the wind noise signal, a value of a third frame number flag is increased by 1, and it is determined whether the value of the third frame number flag is greater than a third preset frame number threshold;

[0057] When the value of the third frame number flag is greater than the third preset frame number threshold, the current state is modified, when the current state is the speech signal, the current state is modified as the wind noise signal, and when the current state is the wind noise signal, the current state is modified as the speech signal;

[0058] When the tentative state is different from the current state, and the tentative state is the speech signal, a value of a first frame number flag is increased by 1, and it is determined whether the value of the first frame number flag is greater than a fourth preset frame number threshold;

[0059] When the value of the first frame number flag is greater than the fourth preset frame number threshold, the current state is modified;

[0060] The first time-domain signal with the modified current state being the speech signal is determined and filtered.

[0061] In the embodiments of the present application, when the temporary state is different from the current state, it indicates that the two times of judgment are inconsistent, at this time, it is possible that at least one time is wrong, or the interval between words when the user speaks, therefore, frame number accumulation can be performed. When the frame number accumulation is greater than the frame number threshold, the corresponding current state is modified, which is equivalent to relying on the continuity between the multiple frames of to-be-detected signals in front of the first time domain signal of the frame determined by the algorithm to predict and determine the state corresponding to the first time domain signal of the frame.

[0062] With reference to the first aspect, in an implementation form of the first aspect, the method further includes:

[0063] when the same, determining and screening the first time domain signal of the current state as the speech signal; or,

[0064] when different, and the value of the third frame number flag is less than or equal to the third preset frame number threshold, determining and screening the first time domain signal of the current state as the speech signal; or,

[0065] when different, and the value of the first frame number flag is less than or equal to the fourth preset frame number threshold, determining and screening the first time domain signal of the current state as the speech signal.

[0066] In the embodiments of the present application, when the temporary state is the same as the current state, or although different, when the frame number accumulation is less than the frame number threshold, the corresponding current state is not modified, which is equivalent to in order to ensure the integrity of the statement, prevent the statement from being interrupted in the middle, the short-time abnormality of several frames can be ignored, and the abnormality is still regarded as the speech signal. Or, it is equivalent to in order to avoid the error of identifying a small amount of wind noise signal as the speech signal, the wind noise signal is still regarded as the wind noise signal.

[0067] With reference to the first aspect, in an implementation form of the first aspect, before the second data satisfies the second condition, the method further includes: performing second initialization processing, the second initialization processing at least including zeroing the value of the first frame number flag and the value of the third frame number flag.

[0068] In the embodiments of the present application, by performing the second initialization processing, data errors or interference of some detection results in other stages can be avoided.

[0069] With reference to the first aspect, in an implementation form of the first aspect, when the second data includes the spectral centroid, the low-frequency energy and the correlation, the second condition includes:

[0070] the spectral centroid is less than a spectral centroid threshold, the low-frequency energy is greater than a low-frequency energy threshold, and the correlation is less than the correlation threshold.

[0071] With reference to the first aspect, in an implementation form of the first aspect, the first microphone includes one or more first microphones, and / or the second microphone includes one or more second microphones.

[0072] With reference to the first aspect, in an implementation form of the first aspect, the first microphone is a microphone arranged at a bottom of the electronic device, and the second microphone is a microphone arranged at a top or a back of the electronic device.

[0073] In a second aspect, an electronic device is provided, including one or more processors, a memory, and a display screen; the memory is coupled to the one or more processors, and the memory is configured to store computer program codes including computer instructions, and the one or more processors are configured to invoke the computer instructions to cause the electronic device to perform any of the voice detection methods in the first aspect.

[0074] In a third aspect, a voice detection apparatus is provided, including a unit configured to perform any of the voice detection methods in the first aspect.

[0075] In a possible implementation form, when the voice detection apparatus is an electronic device, the processing unit can be a processor, and the input unit can be a communication interface; the electronic device can further include a memory configured to store computer program codes, and when the processor executes the computer program codes stored in the memory, the electronic device is caused to perform any of the methods in the first aspect.

[0076] In a fourth aspect, a chip system is provided, the chip is applied to an electronic device, and the chip includes one or more processors configured to invoke computer instructions to cause the electronic device to perform any of the voice detection methods in the first aspect.

[0077] In a fifth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores computer program codes, and when the computer program codes are run by an electronic device, the electronic device is caused to perform any of the voice detection methods in the first aspect.

[0078] In a sixth aspect, a computer program product is provided, and the computer program product includes computer program codes, and when the computer program codes are run by an electronic device, the electronic device is caused to perform any of the voice detection methods in the first aspect.

[0079] The embodiment of the present application provides a voice detection method and a related device, in the process that a user uses an electronic device including at least two microphones to make a voice call or a voice operation, the electronic device can perform pre-processing such as frame division and time-frequency conversion on multiple paths of the to-be-detected signals received by the multiple microphones, and then performs VAD detection to distinguish voice signals and other signals; then, wind noise detection is performed on the screened voice signals, so that the voice signals can be screened again, and the real voice signals and the wind noise signals mistaken as the voice signals can be distinguished. After the continuous VAD detection and wind noise detection on the to-be-detected signals generated by the multiple microphones, the detection accuracy can be greatly improved, the real voice signals, the wind noise signals and other signals can be distinguished, the method is simple, the influence on the voice quality can be avoided, and the detection accuracy can be improved.

[0080] In addition, the voice detection method provided by the present application only involves a method, does not involve improvement on hardware, and does not need to add a complex acoustic structure, so that compared with the related art, the voice detection method provided by the present application is more friendly to small electronic devices and has stronger applicability. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1 is a layout schematic diagram of a microphone provided by the embodiment of the present application;

[0082] Figure 2 is a schematic diagram of an application scenario suitable for the present application;

[0083] Figure 3 is a schematic diagram of another application scenario suitable for the present application;

[0084] Figure 4 is a flow schematic diagram of a voice detection method provided by the embodiment of the present application;

[0085] Figure 5 is a flow schematic diagram of another voice detection method provided by the embodiment of the present application;

[0086] Figure 6 is a flow schematic diagram of VAD detection provided by the embodiment of the present application;

[0087] Figure 7 is a flow schematic diagram of wind noise detection provided by the embodiment of the present application;

[0088] Figure 8 is an example of VAD detection provided by the embodiment of the present application;

[0089] Figure 9 is an example of data for wind noise detection provided by the embodiment of the present application;

[0090] Figure 10 is an example of wind noise detection provided by an embodiment of the present application;

[0091] Figure 11 is an example of a related interface provided by an embodiment of the present application;

[0092] Figure 12 is an example of a hardware system of an electronic device suitable for the present application;

[0093] Figure 13 is an example of a software system of an electronic device suitable for the present application;

[0094] Figure 14 is an example of a structure of a voice detection device provided by the present application;

[0095] Figure 15 is an example of a structure of an electronic device provided by the present application. DETAILED DESCRIPTION

[0096] The technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0097] First, some terms in the embodiments of the present application will be explained to facilitate understanding by those skilled in the art.

[0098] 1. Noise, generally refers to the sound produced by other sound sources in the background of the sound source.

[0099] 2. Noise reduction, refers to the process of reducing noise in audio data.

[0100] 3. Wind noise, is the sound produced by air turbulence near the microphone, including the sound produced by air turbulence caused by the wind; it should be understood that the sound source of wind noise is near the microphone.

[0101] 4. Speech recognition, refers to a technology in which an electronic device processes a collected voice signal according to a pre-configured speech recognition algorithm to obtain a recognition result representing the meaning of the voice signal.

[0102] 5. Framing, is to segment the audio data according to a specified length (time period or number of samples) for subsequent batch processing, and to structure the entire segment of audio data into a certain data structure. It should be understood that the signal after framing processing is a time domain signal.

[0103] 6. Time-frequency transform, that is, converting audio data from time domain (relationship between time and amplitude) to frequency domain (relationship between frequency and amplitude). For example, time-frequency transform can be performed using Fourier transform, fast Fourier transform, etc.

[0104] 7、Fourier transform, Fourier transform is a linear integral transform used to represent the transformation of a signal between the time domain (or, spatial domain) and the frequency domain.

[0105] 8、Fast Fourier transform (FFT), FFT refers to a fast algorithm for the discrete Fourier transform, which can transform a signal from the time domain to the frequency domain.

[0106] 9、Voice activity detection (VAD), voice activity detection is a technology used for speech processing, which aims to detect whether a speech signal exists.

[0107] The above is a brief introduction to the terms involved in the embodiments of the present application, which will not be described hereinafter.

[0108] The voice detection method provided by the embodiments of the present application can be applied to various electronic devices.

[0109] In some embodiments of the present application, the electronic device can be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, a smart dictionary pen, a smart recording pen, a smart translator, a smart sound box, a headset, a hearing aid, a conference telephone device, and the like, including at least two microphones. The embodiments of the present application do not make any limitation on the specific type of electronic device.

[0110] Taking the mobile phone as an example, Figure 1 a layout diagram of the microphones arranged on the mobile phone is shown.

[0111] Exemplarily, as Figure 1 shown, the electronic device 10 has two microphones (microphone, MIC). The microphone, also known as "microphone", "microphone" or "sound pickup device", is used to convert a sound signal into an electrical signal. In the embodiments of the present application, the electronic device can receive a sound signal based on multiple microphones, and convert the sound signal into an electrical signal that can be processed subsequently.

[0112] Generally, the electronic device 10 includes two microphones, one of which is arranged at the bottom of the phone, and the other is arranged at the top of the phone. When the user holds the phone to make a call, the microphone arranged at the bottom of the phone is close to the user's mouth, which can be referred to as a main microphone, and the other can be referred to as an auxiliary microphone. The main microphone can also be referred to as a bottom microphone, and the auxiliary microphone can also be referred to as a top microphone. In the case of only one bottom microphone and one top microphone, the voice detection method provided by the present application executed by the electronic device can also be referred to as a dual-microphone voice detection method.

[0113] Figure 1 This is only an example of a microphone layout, and when the electronic device 10 includes two microphones, the arrangement positions of the two microphones can also be adjusted as needed. For example, one microphone can be arranged at the bottom of the phone, and the other can be arranged on the back of the phone.

[0114] Of course, the electronic device 10 can also include three or more microphones, and the embodiments of the present application do not make any limitation in this regard. For example, when the electronic device is a phone with two foldable display screens, the electronic device can arrange one bottom microphone and one top microphone on one display screen, and arrange one bottom microphone on the other display screen; or arrange one bottom microphone and one top microphone on each display screen; or also arrange multiple bottom microphones and multiple top microphones on each display screen, which can be arranged and adjusted as needed, and the embodiments of the present application do not make any limitation.

[0115] In combination with the above-mentioned electronic device 10, Figure 2 and Figure 3 schematics of two application scenarios provided by the embodiments of the present application.

[0116] As Figure 2 shown, when the user uses the electronic device to make a voice call, due to the reason of pronunciation and exhalation, the user may exhale against the microphone in the electronic device during speaking, resulting in that the audio data received by the electronic device not only includes voice content, but also may include wind noise caused by blowing.

[0117] As Figure 3As shown, when a user uses an electronic device to perform a voice operation (for example, wakes up a voice assistant to open a map application on the electronic device) while running, the electronic device carried by the user also moves quickly due to the user's fast running; at this time, a relatively fast wind speed is formed around the electronic device, causing the audio data received by the electronic device to include not only voice content but also wind noise generated by the relatively fast airflow near the microphone. Because wind noise is similar to voice in some characteristics, such as both being low-frequency and non-stationary signals, it is possible for the voice assistant in the electronic device to mistakenly recognize wind noise as voice, thereby causing false wake-up, false operation, and the like.

[0118] In addition, in addition to receiving voice generated by a user, a microphone generally also receives other sounds in the surrounding environment. For example, the sound of a car horn, the sound of metal striking, the sound of walking on the ground, and the like.

[0119] Currently, the processing of audio data received by an electronic device in the related art generally includes noise reduction, voice recognition using a trained neural network model, and the like.

[0120] However, when noise reduction is performed on audio data, the voice content may also be reduced to some extent when the noise reduction effect is good, causing voice distortion in the later stage. When a trained neural network model is used to recognize voice from audio data, because the samples used for training the neural network model are generally limited and the learning is imperfect, the trained neural network model cannot accurately recognize voice when used. In addition, the cost of deploying a neural network model on an electronic device is also relatively high.

[0121] In addition, for small electronic devices such as mobile phones and earphones, because of the size limitation of the electronic device, a complex acoustic structure cannot be used to weaken or eliminate wind noise.

[0122] To address these problems, there is an urgent need for a new voice detection method.

[0123] Therefore, the voice detection method provided in the embodiments of the present application can distinguish the real voice signal, the wind noise signal and other signals by combining the signals to be detected generated by the multiple microphones and by detecting the signals to be detected in the VAD and wind noise stages. The simple detection method does not involve hardware change, can avoid affecting the voice quality and can improve the detection accuracy.

[0124] In the embodiments of the present application, the other signal refers to the signal other than the voice signal and the wind noise signal.

[0125] The voice detection method provided in the embodiments of the present application will be described below. Figures 4 to 10 The voice detection method provided in the embodiments of the present application will be described below.

[0126] Figure 4 FIG. 1 is a flowchart of a voice detection method provided in the embodiments of the present application. The voice detection method 100 can be executed by the electronic device 10 shown in FIG. 2, and the two microphones are used to collect the sound in the same environment. The voice detection method includes the following S110 to S150, which will be described in detail below. Figure 1 The voice detection method 100 can be executed by the electronic device 10 shown in FIG. 2, and the two microphones are used to collect the sound in the same environment. The voice detection method includes the following S110 to S150, which will be described in detail below.

[0127] For example, the microphones used to collect the sound in the same environment can refer to that when a user makes a call by using a mobile phone outdoors, the two microphones of the mobile phone collect the call sound, wind noise and other sound in the surrounding environment of the user.

[0128] For example, the microphones used to collect the sound in the same environment can refer to that when a user makes a call by using a mobile phone outdoors, the two microphones of the mobile phone collect the call sound, wind noise and other sound in the surrounding environment of the user.

[0129] S110, acquiring audio data. The audio data includes multiple signal streams to be detected.

[0130] The signal stream to be detected refers to a signal sequence including voice, wind noise and other sound and having a certain time sequence.

[0131] For example, one microphone is used to acquire one to-be-tested signal stream, and two microphones can be used to acquire two to-be-tested signal streams, for example, a first microphone is used to acquire a first to-be-tested signal, and a second microphone is used to acquire a second to-be-tested signal. It should be understood that the multiple to-be-tested signal streams should have the same starting time and ending time. One channel can also be understood as one path.

[0132] For example, taking an earphone as an example, in response to a user operation, an electronic device activates a voice call application; in a process of running the voice call application to make a voice call, the electronic device can acquire audio data such as a user's call content.

[0133] For example, taking a smart recording pen as an example, in response to a user operation, an electronic device activates a recording application; in a process of running the recording application to record, the electronic device can acquire audio data such as a user's singing voice.

[0134] For example, taking a smart speaker as an example, in response to a user operation, an electronic device activates a voice assistant application; in a process of running the voice assistant application to interact with a machine, the electronic device acquires audio data such as a user's keyword instruction.

[0135] For example, taking a tablet computer as an example, the audio data can also be audio data such as a voice of another person received by the electronic device when the electronic device runs a third-party application (for example, WeChat).

[0136] S120, pre-processing the multiple to-be-tested signal streams.

[0137] Optionally, the pre-processing at least includes framing and time-frequency transformation, and in terms of execution sequence, the framing is prior to the time-frequency transformation. Of course, the pre-processing can also include other steps, and the embodiments of the present application do not make any limitation in this regard.

[0138] For example, the framing can be performed with a length of 20 ms as one frame.

[0139] For example, the first to-be-tested signal stream acquired by the first microphone can be framed, divided into multiple frames of first time-domain signals, and the time-frequency transformation is performed on the multiple frames of first time-domain signals, to obtain multiple frames of first frequency-domain signals. The first time-domain signal is located in the time domain, the first frequency-domain signal is located in the frequency domain, and the first time-domain signal and the first frequency-domain signal have a one-to-one correspondence.

[0140] Similarly, the second to-be-tested signal stream acquired by the second microphone can be framed, divided into multiple frames of second time-domain signals, and the time-frequency transformation is performed on the multiple frames of second time-domain signals, to obtain multiple frames of second frequency-domain signals. The second time-domain signal is located in the time domain, the second frequency-domain signal is located in the frequency domain, and the second time-domain signal and the second frequency-domain signal have a one-to-one correspondence.

[0141] It should be understood that after the same length is used to frame the multiple to-be-tested signal streams, the number of the multiple first time-domain signals and the multiple second time-domain signals obtained is the same, and they have a one-to-one corresponding relationship in order. Therefore, after the frequency domain conversion is performed on the multiple first time-domain signals and the multiple second time-domain signals after framing, the number of the multiple first frequency-domain signals and the multiple second frequency-domain signals obtained is also the same, and they also have a one-to-one corresponding relationship in order.

[0142] It should also be understood that the multiple first time-domain signals and the multiple second time-domain signals generated after framing, and the multiple first frequency-domain signals and the multiple second frequency-domain signals generated after time-frequency conversion, can be stored in order, so as to improve the efficiency of subsequent processing.

[0143] S130, performing VAD detection on at least one to-be-tested signal stream of the multiple to-be-tested signal streams after preprocessing, to obtain a first detection result.

[0144] The VAD detection is used to detect whether the to-be-tested signal stream includes a speech signal, and the first detection result includes multiple speech signals and / or other signals.

[0145] Optionally, the VAD detection can be repeatedly performed multiple times, and the speech signals and the other signals are distinguished from the intersection of multiple detection results, to serve as the first detection result.

[0146] For example, two times of VAD detection can be performed on one to-be-tested signal stream after preprocessing, and the signals determined as speech signals both times, and the signals determined as speech signals once and as other signals the other time, are all taken as the speech signals in the first detection result; and the signals determined as other signals both times are taken as the other signals in the first detection result.

[0147] Alternatively, the signals determined as speech signals both times can be taken as the speech signals in the first detection result, and the signals determined as speech signals once and as other signals the other time, and the signals determined as other signals both times, are all taken as the other signals in the first detection result.

[0148] For example, two times of VAD detection can be performed on one to-be-tested signal stream after preprocessing, and the signals determined as speech signals both times, and the signals determined as speech signals once and as other signals the other time, are all taken as the speech signals in the first detection result; and the signals determined as other signals both times are taken as the other signals in the first detection result.

[0149] S140, in combination with the pre-processed multi-path to-be-detected signal stream, wind noise detection is performed on the speech signal in the first detection result to obtain a second detection result.

[0150] The wind noise detection is used to distinguish the speech signal and the wind noise signal, and the second detection result includes multiple frames of speech signals and / or wind noise signals.

[0151] It should be understood that the VAD detection performed on the pre-processed multi-path to-be-detected signal can determine whether the to-be-detected signal includes a speech signal, and further distinguish the speech signal and other signals therefrom; and since the characteristics of the wind noise signal are similar to those of the speech signal, after the first-stage VAD detection, the wind noise signal and the speech signal cannot be accurately distinguished, and there may be a case that the wind noise signal is mistaken for the speech signal, that is, the speech signal in the first detection result obtained after the VAD detection is only a suspected speech signal, which may include a wind noise signal. Then, the wind noise detection is continued to further distinguish the true speech signal and the false speech signal (i.e., the wind noise signal). Thus, after the continuous VAD detection and wind noise detection, the accuracy of the detection can be greatly improved. Since the VAD detection and the wind noise detection provided in the present application do not affect the quality of the signal itself, there is no problem of loss of to-be-detected signal quality.

[0152] Optionally, when the first detection result does not include a speech signal, the step S140 can not be performed.

[0153] Optionally, the wind noise detection can be repeatedly performed multiple times, and the speech signal and the wind noise signal are distinguished from the intersection of multiple second detection results.

[0154] For example, the wind noise detection is performed three times on the speech signal in the first detection result, and the signals determined as the speech signal in any two of the three times are taken as the speech signal in the second detection result.

[0155] It should be understood that during the execution of the entire method, the number of times of performing the VAD detection and the wind noise detection can not be the same, and the specific number of repetitions can be set and modified as needed, and the embodiments of the present application do not make any limitation thereon.

[0156] Optionally, the VAD detection and the wind noise detection can be performed on multiple frames of to-be-detected signals in a time period in the pre-processed to-be-detected signal stream, and then the VAD detection and the wind noise detection are repeatedly performed on multiple frames of to-be-detected signals in a next time period, and the subsequent is sequentially similar.

[0157] It should be understood that this way has relatively lower requirements on the hardware performance for executing the method, and is easier to implement.

[0158] Optionally, after VAD detection and wind noise detection are performed on a frame of the preprocessed to-be-detected signal, VAD detection and wind noise detection are repeatedly performed on the next frame of to-be-detected signal, and the subsequent frames are sequentially detected.

[0159] Optionally, VAD detection and wind noise detection are performed on a frame of to-be-detected signal, and while the wind noise detection is performed on the frame of to-be-detected signal, VAD detection is performed on the next frame of to-be-detected signal.

[0160] It should be understood that this method has a faster response speed and processing speed, and can detect the voice signal, wind noise signal, and other signals in the signal in real time while collecting the signal.

[0161] The embodiments of the present application provide a voice detection method. During a voice call or voice operation of a user using an electronic device including at least two microphones, the electronic device can first perform frame division and time-frequency conversion on the multiple to-be-detected signals received by the multiple microphones, and then perform VAD detection to distinguish the voice signal and other signals in the to-be-detected signals. Then, wind noise detection is performed on the filtered voice signal, so that the voice signal is filtered again to distinguish the real voice signal and the wind noise signal that is misjudged as a voice signal. After the continuous VAD detection and wind noise detection on the to-be-detected signals generated by the multiple microphones, the detection accuracy can be greatly improved, and the real voice signal, wind noise signal, and other signals can be distinguished. The method is simple, can avoid affecting the voice quality, and can improve the detection accuracy.

[0162] In addition, the voice detection method provided by the present application only involves a method and does not involve improvement of hardware, and does not need to add a complex acoustic structure. Therefore, compared with the related art, the voice detection method provided by the present application is more friendly to small electronic devices and has stronger applicability.

[0163] For example, VAD detection can be performed on a first to-be-detected signal stream in the preprocessed multiple to-be-detected signal streams to obtain a first detection result. VAD detection is not performed on the other to-be-detected signal streams.

[0164] Then, wind noise detection is performed on the voice signal in the first detection result in combination with the to-be-detected signals of the other to-be-detected signal streams preprocessed in the corresponding order, to determine whether the voice signal in the first detection result remains as a voice signal or is changed to a wind noise signal.

[0165] It should be understood that in this way, the first to-be-detected signal is the main detected signal, and the other to-be-detected signals are used to assist in detecting the voice signal in the first to-be-detected signal.

[0166] The following will be described in detail in combination with Figure 5 The example will be described in detail.Figure 5 Another flow diagram of the voice detection method provided by the embodiment of the present application is shown, which can include the following S210 to S250. The steps S210 to S250 are described as follows.

[0167] S210, a first to-be-detected signal stream and a second to-be-detected signal stream are acquired.

[0168] It should be understood that the first to-be-detected signal stream and the second to-be-detected signal stream are audio data, and the present application is used for processing audio data in a period of time. For example, the time length of the first time domain signal stream and the second time domain signal stream is 600 ms.

[0169] S220, the first to-be-detected signal stream and the second to-be-detected signal stream are preprocessed to obtain a plurality of first time domain signals corresponding to the first to-be-detected signal stream, a plurality of first frequency domain signals, a plurality of second time domain signals corresponding to the second to-be-detected signal stream, and a plurality of second frequency domain signals. The preprocessing includes framing and time-frequency transformation.

[0170] Optionally, as shown in the above S220 can include: Figure 5

[0171] S221, the first to-be-detected signal is framed to obtain a plurality of first time domain signals; and the second to-be-detected signal stream is framed to obtain a plurality of second time domain signals.

[0172] For example, the 600 ms first to-be-detected signal is framed to obtain 30 first time domain signals; and the 600 ms second to-be-detected signal stream is framed to obtain 30 second time domain signals.

[0173] It should be understood that the plurality of first time domain signals and the plurality of second time domain signals are time domain signals.

[0174] S222, the plurality of first time domain signals obtained by S221 are subjected to time-frequency transformation to obtain a plurality of first frequency domain signals corresponding to the frame number; and the plurality of second time domain signals are subjected to time-frequency transformation to obtain a plurality of second frequency domain signals corresponding to the frame number.

[0175] For example, the 30 first time domain signals are subjected to time-frequency transformation to obtain 30 first frequency domain signals; and the 30 second time domain signals are subjected to time-frequency transformation to obtain 30 second frequency domain signals.

[0176] S230, VAD detection is performed on the preprocessed first to-be-detected signal stream.

[0177] ​The S230 can also be expressed as: performing VAD detection on the first plurality of time-domain signals and the first plurality of frequency-domain signals corresponding to the first to-be-detected signal stream. The first plurality of time-domain signals and the first plurality of frequency-domain signals have a one-to-one correspondence.

[0178] Here, no VAD detection is performed on the preprocessed second to-be-detected signal stream.

[0179] Optionally, as shown in the S230 can include: Figure 5

[0180] S231, for the first time-domain signal, determining the corresponding zero crossing rate (ZCR).

[0181] The zero crossing rate refers to the ratio of the zero points (from positive to negative or from negative to positive) of the speech signal in each frame of the first time-domain signal. Generally, the zero crossing rate of noise or other sounds is small, and the zero crossing rate of speech signals is relatively large.

[0182] For example, the value of the zero crossing rate of the first time-domain signal can be determined by the following formula (1). Formula (1) is:

[0183]

[0184] Where t is a time point within a frame, T is the length of each frame, S represents the amplitude of the signal (S has positive and negative values); if the amplitudes of two adjacent time points are both positive or both negative, then π{A} is 0; if one is positive and the other is negative, then π{A} is 1; the π values of T-1 adjacent points within a frame are summed and then divided by T-1, which is the ratio of the zero crossing points within a frame, simply referred to as the zero crossing rate.

[0185] S232, for the first frequency-domain signal corresponding to the first time-domain signal, determining the corresponding spectral entropy and flatness, respectively.

[0186] It should be understood that the spectral entropy describes the relationship between the power spectrum and the entropy rate. In this application, the dispersion degree of the signal can be described. If the signal is noise, the signal is relatively dispersed, and the spectral entropy is relatively high; if the signal is speech, the signal is relatively concentrated, and the spectral entropy is relatively low. Flatness is used to describe the flatness of the signal. The flatness of the noise is large, and the flatness of the speech signal is relatively small.

[0187] For example, the value of the spectral entropy of the first time-domain signal can be determined by the following set of formulas (2). Formula (2) is:

[0188]

[0189] N power (k, m) = X(k, m), 1 ≤ k ≤ N / 2​

[0190]

[0191]

[0192] wherein r(n) represents a short-time autocorrelation function of each frame signal, L is a window length, N is an FFT transform length, and X(k, m) represents a power spectrum amplitude of a kth frequency point of an mth frame; for an actual signal, X(k, m) is symmetrical about N / 2+1, so that X(k, m) = X(N-k+1, m) power (k, m) is equal to X(k, m), and X power (k, m) represents a power spectrum energy; P(i, m) represents a probability that a power spectrum energy of each frequency component accounts for a power spectrum energy of the entire frame; a power spectrum entropy size corresponding to each frame can be represented as H(m).

[0193] For example, a value of the flatness of the first time-domain signal can be determined by the following formula (3). The formula (3) is:

[0194]

[0195] wherein L is an Lth frequency point after FFT transform, N is an Nth frequency point after FFT transform, Y(L) is an energy of the Lth frequency point, and a calculation formula is the same as that of X power (k); exp(x) is e raised to the xth power.

[0196] S233, in combination with the value of the zero-crossing rate, the spectrum entropy and the flatness corresponding to each frame of the first time-domain signal, determine whether the frame of the first time-domain signal is a speech signal or other signal.

[0197] It should be understood that in addition to the zero-crossing rate, the spectrum entropy and the flatness, other related data can also be determined to distinguish whether the first time-domain signal is a speech signal or other signal, and the related data can be set and modified as needed, and the present application does not make any limitation in this regard.

[0198] S234, screen out the first time-domain signal determined as a speech signal.

[0199] If the first time-domain signal is a speech signal, the first time-domain signal can be cut off; at the same time, the first frequency-domain signal corresponding to the first time-domain signal after time-frequency transform can also be cut off, which is convenient for subsequent continuous detection.

[0200] S240, in combination with the preprocessed second road signal stream to be detected, perform wind noise detection on the speech signal determined in S230.

[0201] The S240 can also be expressed as: in combination with the multiple frames of the second frequency domain signals corresponding to the second to-be-tested signal stream, performing wind noise detection on the first time domain signals determined as speech signals from the preprocessed first to-be-tested signal stream.

[0202] Optionally, as shown in FIG. 2B, the S240 can include: Figure 5

[0203] The S241 can determine the spectral centroid and the low frequency energy corresponding to each frame of the first frequency domain signals based on the multiple frames of the first frequency domain signals corresponding to the multiple frames of the first time domain signals determined as speech signals in the VAD detection.

[0204] It should be understood that the spectral centroid is used to describe the position of the center of gravity of a signal. The spectral centroid of a wind noise signal is low, and the spectral centroid of a speech signal is high. The low frequency energy is used to describe the size of the low frequency energy in a signal. The low frequency energy of a wind noise signal is high, and the low frequency energy of a speech signal is small.

[0205] For example, the value of the spectral centroid of the first time domain signal can be determined by the following formula (4).

[0206] The formula (4) is:

[0207]

[0208] wherein r is the spectral centroid, i is the coordinate value of each point on the spectrum, f ndata (i) is the amplitude of each point on the spectrum.

[0209] For example, the value of the low frequency energy of the first time domain signal can be determined by the following formula (5). The formula (5) is:

[0210]

[0211] wherein E is the low frequency energy, X(f) is the FFT result corresponding to the frequency f, and the energy is calculated by taking the absolute value and squaring. f1 and f2 represent the start and end frequencies of the selected low frequency range; for example, if the low frequency range is selected as 100-500 Hz, then f1=100 and f2=500.

[0212] The S242 can determine the correlation corresponding to a group of the first frequency domain signals and the second frequency domain signals in the same order based on the multiple frames of the first frequency domain signals corresponding to the multiple frames of the first time domain signals determined as speech signals in the VAD detection, and the multiple frames of the second frequency domain signals selected in the corresponding order from the preprocessed second to-be-tested signal stream.

[0213] ​It should be understood that the correlation is used to describe the similarity between two signals. The correlation of wind noise is relatively low, and the correlation of speech signal is relatively high.

[0214] For example, the value of the correlation of the first time-domain signal can be determined by the following formula (6). The formula (6) is:

[0215]

[0216] Wherein, X is the first frequency-domain signal, Y is the second frequency-domain signal, r(X, Y) is the correlation between the two; Cov(X, Y) is the covariance of X and Y, D(X), D(Y) are the variances of X and Y, respectively.

[0217] S243, at least in combination with the value of the correlation, the spectral centroid and the low frequency energy corresponding to each frame of the first time-domain signal, determine whether the frame of the first time-domain signal is a speech signal or a wind noise signal.

[0218] It should be understood that in addition to the correlation, the spectral centroid and the low frequency energy, other relevant data can also be determined to distinguish whether the first time-domain signal is a speech signal or a wind noise signal. The relevant data can be set and modified as needed, and the present application does not make any limitation thereto.

[0219] S244, screening out the first time-domain signal determined as a speech signal again.

[0220] If the first time-domain signal is a speech signal, the first time-domain signal can be cut out as the final detected speech signal.

[0221] S250, obtaining the detection result.

[0222] When the above detection is performed on a frame of the first time-domain signal, the detection result obtained is that the frame of the first time-domain signal is determined as a speech signal, other signal or wind noise signal. When the above detection is performed on multiple frames of the first time-domain signal, the detection result obtained includes information that each frame of the first time-domain signal in the multiple frames of the first time-domain signal is a speech signal, other signal or wind noise signal, and the signal determined as a speech signal is cut out.

[0223] Exemplarily, the first stream of the signal to be detected is the signal obtained by the bottom microphone of the mobile phone, and the second stream of the signal to be detected is the signal obtained by the top microphone of the mobile phone. In combination with the above process, the signal received by the bottom microphone is equivalent to the main signal to be detected, and the signal received by the top microphone is used to assist in detecting the speech signal in the signal received by the bottom microphone. In combination with the signal received by the top microphone, it can be determined that all signals in the bottom microphone are speech signals, wind noise signals or other signals, and the speech signal can be cut out.

[0224] It should be understood that the determined multiple frames of voice signals can be reordered in sequence for storage or other processing such as recognition, and the embodiments of the present application do not make any limitation thereon.

[0225] In the voice detection method provided by the embodiments of the present application, in the process of a user using an electronic device including two microphones for voice call or voice operation, the electronic device can first perform pre-processing such as framing and time-frequency transformation on two channels of to-be-detected signals received by the two microphones; then, in combination with multiple frames of first time-domain signals and multiple frames of first frequency-domain signals generated during pre-processing of the first channel of to-be-detected signal stream, determine the zero-crossing rate, spectral entropy and flatness; then, in combination with the zero-crossing rate, spectral entropy and flatness, determine whether the first time-domain signal is a voice signal or other signal, and screen out the first time-domain signal determined as a voice signal and the first frequency-domain signal corresponding thereto; then, for the first frequency-domain signal corresponding to the screened voice signal and the second frequency-domain signal corresponding to the same sequence after pre-processing of the second channel of to-be-detected signal stream, determine the correlation, spectral centroid and low-frequency energy; and then, in combination with the correlation, spectral centroid and low-frequency energy, determine whether the voice signal determined in the VAD detection stage is a real voice signal or a wind noise signal misjudged as a voice signal. Thus, through the cooperation of the two channels of to-be-detected signals and the continuous detection of signal characteristics in the VAD detection and wind noise detection stages, the real voice signal, wind noise signal and other signals can be distinguished. The method is simple, can improve the accuracy of detection while avoiding the influence on the voice quality, and can improve the accuracy of detection while avoiding the influence on the voice quality.

[0226] Optionally, Figure 6 A flowchart of a method for determining whether each frame of first time-domain signal is a voice signal or other signal (i.e., S233) in combination with the values of the zero-crossing rate, spectral entropy and flatness corresponding to the frame of first time-domain signal is shown. As shown in Figure 6 The determination method 300 can include the following S301 to S310.

[0227] S301, first initialization processing is performed.

[0228] It should be understood that in addition to the signal data itself, the multiple frames of first time-domain signals can also include three frame number flag bits (i, j and k) and two signal flag bits (int and SF) corresponding to each frame of first time-domain signal.

[0229] For example, the signal flag bit int is used to represent the tentative state of the first time-domain signal; int equal to 1 indicates that the frame of first time-domain signal is tentatively determined as a voice signal; int equal to 0 indicates that the frame of first time-domain signal is tentatively determined as other signal; and int equal to -1 indicates that the frame of first time-domain signal is tentatively determined as a wind noise signal.

[0230] The signal flag bit SF is used to represent the current state of the first time domain signal; when SF equals 1, it indicates that the first time domain signal in the frame is currently determined to be a speech signal; when SF equals 0, it indicates that the first time domain signal in the frame is currently determined to be other signals; and when SF equals -1, it indicates that the first time domain signal in the frame is currently determined to be wind noise signal.

[0231] The frame number flag i is used to represent the number of accumulated frames when the tentative state is a speech signal, for example, i equals 1 indicates that the number of accumulated signals in the tentative state of speech signal is 1 frame. The second frame number flag bit j is used to represent the number of accumulated frames when the tentative state is other state, for example, j equals 2 indicates that the number of accumulated signals in the tentative state of other signal is 2 frames. The third frame number flag bit k is used to represent the number of accumulated frames when the tentative state is wind noise signal, for example, k equals 3 indicates that the number of accumulated signals in the tentative state of wind noise signal is 3 frames.

[0232] Based on this, for multiple frames of first time domain signals, the first initialization processing is equivalent to zero processing of the three frame number flag bits and the two signal flag bits corresponding to each first time domain signal, avoiding interference, so that they are all 0.

[0233] S302, determine whether the spectral entropy, flatness and zero-crossing rate corresponding to the first time domain signal meet the first condition?

[0234] The first condition includes that the zero-crossing rate is greater than the zero-crossing rate threshold, the spectral entropy is less than the spectral entropy threshold, and the flatness is less than the flatness threshold.

[0235] The above S302 can also be described as: determining whether the zero-crossing rate corresponding to the first time domain signal is greater than the zero-crossing rate threshold? Determining whether the spectral entropy determined by the first frequency domain signal converted from the first time domain signal is less than the spectral entropy threshold? And whether the flatness is less than the flatness threshold?

[0236] It should be understood that the zero-crossing rate threshold, the spectral entropy threshold and the flatness threshold can be set and modified as needed, and the embodiments of the present application do not make any limitation on this.

[0237] S303, when the spectral entropy, flatness and zero-crossing rate corresponding to the first time domain signal meet the first condition, determining that the tentative state of the first time domain signal is a speech signal, and modifying the value of the first signal flag bit to X.

[0238] It should be understood that since a speech word usually lasts for several frames and there is an interval between words, in order to completely judge the beginning and end of a sentence and prevent the sentence from being interrupted in the middle, each frame of first time domain signal is provided with a tentative state and a current state. Among them, the tentative state and the current state can be divided into three states: speech signal, wind noise signal and other signals.

[0239] S304, when the spectral entropy, flatness and zero-crossing rate corresponding to the first time-domain signal do not meet the first condition, determining that the tentative state of the first time-domain signal is other signal, and modifying the first signal flag bit to Y.

[0240] That is, when the zero-crossing rate corresponding to the first time-domain signal is greater than the zero-crossing rate threshold value, the spectral entropy determined by the converted first frequency-domain signal is less than the spectral entropy threshold value, and the flatness is also less than the flatness threshold value, it can be considered that the first time-domain signal meets the characteristics of the speech signal, and it can be determined that the tentative state of the first time-domain signal is the speech signal, and the first time-domain signal corresponds to the signal flag bit int representing the tentative state, that is, X equals 1.

[0241] In addition, when any one of the zero-crossing rate, spectral entropy and flatness corresponding to the first time-domain signal does not meet the corresponding condition, it can be considered that the first time-domain signal does not meet the characteristics of the speech signal, and it can be determined that the tentative state of the first time-domain signal is other signal, and the first time-domain signal corresponds to the signal flag bit int representing the tentative state, that is, Y equals 0.

[0242] S305, after determining the tentative state corresponding to the first time-domain signal, whether the tentative state determined by the first time-domain signal is the same as the current state corresponding thereto, no matter whether the tentative state of the first time-domain signal is the speech signal or other signal.

[0243] The signal flag bit for representing the current state is SF, therefore, whether the tentative state determined by the first time-domain signal is the same as the current state corresponding thereto can be determined by comparing the value of the signal flag bit int and the value of the signal flag bit SF.

[0244] S306, when the tentative state is different from the current state, frame number accumulation is performed. If the tentative state is the speech signal, the first frame number flag i is accumulated by 1; if the tentative state is other signal, the second frame number flag j is accumulated by 1.

[0245] S307, when the frame number accumulated by the first frame number flag i is greater than the first preset frame number threshold value, the current state is modified, that is, the corresponding current state is modified from the speech signal to other signal, or from other signal to the speech signal.

[0246] Similarly, when the frame number accumulated by the second frame number flag j is greater than the second preset frame number threshold value, the current state is modified, that is, the corresponding current state is modified from the speech signal to other signal, or from other signal to the speech signal.

[0247] It should be understood that when the provisional state differs from the current state, it indicates that the two judgments are inconsistent. In this case, it is possible that at least one judgment is incorrect. Therefore, frame count accumulation can be performed. When the accumulated frame count exceeds the frame count threshold, the corresponding current state is modified. This is equivalent to relying on the continuity between the test signals of multiple frames preceding the first time-domain signal of this frame, as determined by the algorithm, to predict and determine the state corresponding to the first time-domain signal of this frame.

[0248] For example, the provisional state of the first time domain signal in the 6th frame is a speech signal, and the current state is other signals. After counting the frames, the number of frames with the provisional state of speech signal is already 6, indicating that the first time domain signals in the previous 5 frames were all speech signals. At this time, it is more likely that the first time domain signal in the 6th frame is still a speech signal. The original current state is no longer trusted, and the current state is changed from other signals to speech signal.

[0249] It should be understood that the first preset frame rate threshold and the second preset frame rate threshold can be set and modified as needed, and the embodiments of this application do not impose any restrictions on this.

[0250] S308. In S305 above, when the provisional state is the same as the current state, continue to determine whether the current state is a voice signal; or, after S306, when the first frame count flag i is less than or equal to the first preset frame count threshold, or the second frame count flag j is less than or equal to the second preset frame count threshold, continue to determine whether the current state is a voice signal; or, in S307, after modifying the current state, continue to determine whether the current state is a voice signal.

[0251] It should be understood that consistent results from two judgments are more accurate than a single judgment. Therefore, when the provisional state is the same as the current state, the state result corresponding to the first time-domain signal is more accurate, and there is no need to modify the current state.

[0252] Alternatively, although the provisional state is different from the current state, the cumulative number of corresponding frames does not exceed the preset frame threshold. In this case, it can be considered that the number of first time-domain signals in the same continuous provisional state is too small to be ignored, so no modification is needed, and the current state can continue to be a voice signal or other signal.

[0253] S309. If the current state corresponds to other signals, remove the first time domain signal whose corresponding signal flag bit SF is equal to 0. SF equal to 0 indicates that the determined first time domain signal is other signals.

[0254] S310. If the current state corresponds to a speech signal, filter the first time-domain signal whose corresponding signal flag bit SF is equal to 1 as the first detection result. SF equal to 1 indicates that the determined first time-domain signal is a speech signal.

[0255] Here, if the tentative state is different from the current state, the current state here refers to the modified current state. If the tentative state is the same as the current state, the current state here refers to the original current state.

[0256] Optionally, Figure 7 A flowchart of a method for determining whether a first time-domain signal is a speech signal or a wind-noise signal (i.e., S242) according to an embodiment of the present application is shown. As shown in FIG. 4, the method 400 can include the following steps S401-S410. Figure 7

[0257] S401, for the multiple frames of the first time-domain signals determined as speech signals in S310, a second initialization processing is performed.

[0258] It should be understood that since the signal flag SF used to represent the current state has been determined in the method shown in FIG. 3 to be equal to 1, the second initialization processing can not be performed on the signal flag SF, and the second frame number flag j corresponding to the tentative state as other signals is not processed. Only the signal flag int, the first frame number flag i used to represent the tentative state as a speech signal, and the third frame number flag k used to represent the tentative state as a wind-noise signal are reset to 0. Figure 6

[0259] Of course, since the third frame number flag k has been reset to 0 in the first initialization processing in the VAD detection stage and is not used, the third frame number flag k can not be reset to 0 in the wind-noise detection stage. If the third frame number flag k is not reset to 0 in the first initialization processing, the third frame number flag k can be reset to 0 before the wind-noise detection to avoid calculation errors.

[0260] S402, determine whether the correlation, spectral centroid, and low-frequency energy corresponding to the first time-domain signal meet a second condition?

[0261] The second condition includes that the correlation is less than a correlation threshold, the spectral centroid is less than a spectral centroid threshold, and the low-frequency energy is greater than a low-frequency energy threshold.

[0262] The above S402 can also be described as: combining the first frequency-domain signal corresponding to the time-frequency transform of the first time-domain signal, and the second frequency-domain signal determined in sequence from the multiple frames of the second frequency-domain signals included in the preprocessed second path of the to-be-detected signal stream, to determine the correlation, spectral centroid, and low-frequency energy of the two first frequency-domain signals and the second frequency-domain signal as the values of the correlation, spectral centroid, and low-frequency energy corresponding to the first time-domain signal.​​

[0263] It should be understood that the correlation threshold, the spectral centroid threshold and the low frequency energy threshold can be set and modified as needed, and the embodiments of the present application do not make any limitation thereto.

[0264] S403, when the correlation, the spectral centroid and the low frequency energy corresponding to the first time domain signal meet the second condition, determining that the provisional state of the first time domain signal is a wind noise signal, and modifying the value of the first signal flag bit to Z.

[0265] S404, when the correlation, the spectral centroid and the low frequency energy corresponding to the first time domain signal do not meet the second condition, determining that the provisional state of the first time domain signal is a speech signal, and modifying the value of the first signal flag bit to X.

[0266] That is, when the correlation determined by the first frequency domain signal corresponding to the first time domain signal and the second frequency domain signal of the same order as the first frequency domain signal is less than the correlation threshold, the spectral centroid is less than the spectral centroid threshold, and the low frequency energy is greater than the low frequency energy threshold, it can be considered that the first time domain signal meets the characteristics of the wind noise signal, and the provisional state of the first time domain signal can be determined as the wind noise signal, and the signal flag bit int of the first time domain signal is equal to -1, that is, Z is equal to -1.

[0267] In addition, when any one of the correlation, the spectral centroid and the low frequency energy corresponding to the first time domain signal does not meet the corresponding condition, it can be considered that the first time domain signal does not meet the characteristics of the wind noise signal, and the provisional state of the first time domain signal can be determined as the speech signal, and the signal flag bit int of the first time domain signal is equal to 1, that is, X is equal to 1.

[0268] S405, after determining the provisional state corresponding to the output first time domain signal, whether the provisional state of the first time domain signal is a speech signal or a wind noise signal, it is determined whether the provisional state determined by the first time domain signal is the same as the current state corresponding thereto.

[0269] The signal flag bit for indicating the current state is SF, therefore, whether the provisional state determined by the first time domain signal is the same as the current state corresponding thereto can be determined by comparing the value of the signal flag bit int and the value of the signal flag bit SF.

[0270] S406, when the provisional state is different from the current state, frame number accumulation is performed. If the provisional state is a speech signal, the first frame number flag i is accumulated by 1; if the provisional state is a wind noise signal, the third frame number flag k is accumulated by 1.

[0271] S407, when the frame number accumulated by the third frame number flag k is greater than the third preset frame number threshold, the current state is modified, that is, the corresponding current state is modified from the speech signal to the wind noise signal, or from the wind noise signal to the speech signal.

[0272] When the accumulated frame number of the first frame number flag i is greater than the fourth preset frame number threshold, the current state is modified, that is, the corresponding current state is modified from the speech signal to the wind noise signal, or from the wind noise signal to the speech signal.

[0273] It should be understood that when the provisional state is different from the current state, it means that the two judgments are inconsistent, at this time, it is possible that at least one time is wrong, or the interval between words when the user speaks, therefore, the frame number can be accumulated. When the frame number accumulation is less than the frame number threshold, the corresponding current state is not modified, which is equivalent to in order to ensure the integrity of the statement, preventing the statement from being interrupted in the middle, the short-term abnormality of a few frames can be ignored, and the speech signal is still regarded.

[0274] For example, the provisional state of the 7th frame of the first time domain signal is the wind noise signal, the current state is the speech signal, and after the frame number is counted, the frame number of the provisional state of the speech signal is 6 frames, and the frame number of the provisional state of the wind noise signal is 1 frame. The quantity is small, which indicates that the first 6 frames of the first time domain signal are speech signals, at this time, the 7th frame of the first time domain signal is more likely to be a speech signal, or the 7th frame of the first time domain signal may be a wind noise signal, but in order to ensure the integrity of the statement, preventing the statement from being interrupted in the middle, the current state can continue to be maintained as the speech signal without modification.

[0275] When the frame number accumulation is greater than the frame number threshold, the corresponding current state is modified, which is equivalent to relying on the continuity between the multiple frames of signals before the frame of the first time domain signal determined by the algorithm to predict and determine the state corresponding to the frame of the first time domain signal.

[0276] It should be understood that the third preset frame number threshold and the fourth preset frame number threshold can be set and modified as needed, and the embodiments of the present application do not make any limitation in this regard.

[0277] S408, in the above S405, when the provisional state is the same as the current state, it is determined whether the current state is the wind noise signal; or, after S406, when the third frame number flag k is less than or equal to the third preset frame number threshold, or the first frame number flag i is less than or equal to the fourth preset frame number threshold, it is determined whether the current state is the wind noise signal; or, in S407, after the current state is modified, it can be determined whether the current state is the wind noise signal.

[0278] It should be understood that the two determination results are more accurate than the one determination result. Therefore, when the tentative state is the same as the current state, the state result corresponding to the determined first time domain signal is more accurate, and the current state does not need to be modified.

[0279] Alternatively, the tentative state is different from the current state, but the cumulative number of corresponding frame numbers does not exceed the preset frame number threshold. At this time, it can be considered that the number of first time domain signals of the same tentative state in succession is too small to be ignored, so the current state does not need to be modified and continues to be maintained as the speech signal or the wind noise signal.

[0280] S409, if the current state corresponds to the wind noise signal, the first time domain signal corresponding to the signal flag bit SF equal to -1 is removed, and SF equal to -1 indicates that the determined first time domain signal is the wind noise signal.

[0281] S410, if the current state corresponds to the speech signal, the first time domain signal corresponding to the signal flag bit SF equal to 1 is selected as the second detection result; SF equal to 1 indicates that the determined first time domain signal is the speech signal.

[0282] Here, if the tentative state is different from the current state and the current state is modified, the current state here refers to the modified current state. If the tentative state is the same as the current state, the current state here refers to the current state determined by the VAD detection.

[0283] In combination Figures 5 to 7 , for example, Figures 8 to 10 An example of a speech detection method provided by the embodiment of the application.

[0284] As shown in (a) of Figure 8 , 30 first time domain signals can be obtained after the first road to be detected signal stream is framed. The first initialization processing is performed on the three frame number flag bits involved in the 30 first time domain signals and the two signal flag bits corresponding to each first time domain signal, so that they are all 0.

[0285] Then, as shown in (b) of Figure 8 , the VAD detection is performed on the first frame first time domain signal, the zero crossing rate corresponding to the first frame first time domain signal is determined, and the spectrum entropy and flatness corresponding to the first frequency domain signal obtained by time-frequency transformation of the first frame first time domain signal are determined. Whether the values of the zero crossing rate, the spectrum entropy and the flatness meet the first condition is determined. When the values of the zero crossing rate, the spectrum entropy and the flatness determined by the first frame first time domain signal do not meet the first condition, the tentative state of the first frame first time domain signal is determined to be other signals, int=0; the frame number flag bit for indicating the cumulative frame number of the tentative state being other signals is updated to 1, j=1.

[0286] At this time, since the signal flag bit SF corresponding to the first time domain signal of the first frame is 0; the tentative state and the current state are the same, it is continued to determine whether the current state is a speech signal, which is not a speech signal at this time. Thus, the signal flag bit SF corresponding to the current state of the first time domain signal of the first frame remains 0, i.e. SF = 0.

[0287] Then, the VAD detection is performed on the first time domain signal of the second frame, and the tentative state of the first time domain signal of the second frame is determined to be other signals by using the above method, i.e. int = 0; the tentative state and the current state are the same, it is continued to determine whether the current state is a speech signal, which is not a speech signal at this time. Thus, the signal flag bit SF corresponding to the current state of the first time domain signal of the second frame remains 0, i.e. SF = 0.

[0288] The VAD detection is performed on the first time domain signal of the third frame, and the zero-crossing rate corresponding to the first time domain signal of the third frame is determined by using the above method, and the spectrum entropy and the flatness corresponding to the first frequency domain signal after the time-frequency transformation of the first time domain signal of the third frame are determined. It is determined whether the values of the zero-crossing rate, the spectrum entropy and the flatness meet the first condition. When the values of the zero-crossing rate, the spectrum entropy and the flatness determined by the first time domain signal of the third frame meet the first condition, it is determined that the tentative state of the first time domain signal of the third frame is a speech signal, i.e. int = 1; since the signal flag bit SF corresponding to the current state after the initialization is 0, it can be judged that the tentative state and the current state are different, and the frame number flag bit used to represent the cumulative frame number of the tentative state being a speech signal is updated to 1, i.e. i = 1; the value of i is less than the first preset frame number threshold (for example, 2 frames), at this time, it can be considered that the number of the tentative state being a speech signal is too small, and the judgment is unreliable. Then, it is determined that the current state corresponds to other signals, and the value of the signal flag bit SF corresponding to the current state is maintained, i.e. SF = 0.

[0289] The VAD detection is performed on the first time domain signal of the fourth frame, and the zero-crossing rate corresponding to the first time domain signal of the fourth frame is determined by using the above method, and the spectrum entropy and the flatness corresponding to the first frequency domain signal after the time-frequency transformation of the first time domain signal of the fourth frame are determined. It is determined whether the values of the zero-crossing rate, the spectrum entropy and the flatness meet the first condition. When the values of the zero-crossing rate, the spectrum entropy and the flatness determined by the first time domain signal of the fourth frame meet the first condition, it is determined that the tentative state of the first time domain signal of the fourth frame is a speech signal, i.e. int = 1; since the signal flag bit SF corresponding to the current state after the initialization is 0, it can be judged that the tentative state and the current state are still different, and the frame number flag bit used to represent the cumulative frame number of the tentative state being a speech signal is updated to 2, i.e. i = 2; the value of i is less than the first preset frame number threshold, at this time, it can be continued to consider that the number of the tentative state being a speech signal does not meet the requirement, and the judgment is unreliable. Then, it is determined that the current state corresponds to other signals, and the value of the signal flag bit SF corresponding to the current state is maintained, i.e. SF = 0.

[0290] Similarly, after VAD detection is performed on the first time-domain signal of the 5th frame to the first time-domain signal of the 8th frame, it can be determined that the first time-domain signal of the 5th frame to the first time-domain signal of the 8th frame remains the current state as other signals, and the signal flag bit SF is 0, i.e., SF = 0.

[0291] Next, VAD detection is performed on the first time-domain signal of the 9th frame, and the above method is used to determine that the tentative state of the first time-domain signal of the 9th frame is other signals, and int = 0; the tentative state and the current state are the same, and it is continued to determine whether the current state is a speech signal, and here the speech signal is a speech signal, and thus the signal flag bit SF corresponding to the current state of the first time-domain signal of the 9th frame remains 0, i.e., SF = 0.

[0292] The subsequent frame numbers are sequentially deduced, which will not be described here.

[0293] Optionally, the second VAD detection can also be continued in combination with the speech signal detected in the first VAD detection. It should be noted that when the first initialization of the second VAD detection is started, the flag signal bit of the current state does not need to be zeroed, and the current state result of the first VAD detection should be retained as the initial current state data of the second VAD detection.

[0294] On this basis, as shown in (a) of Figure 9 , taking the speech signal detected in the previous 9 frames of first time-domain signals as an example, although the current state of the first time-domain signal of the 5th frame to the first time-domain signal of the 8th frame included in the first to-be-detected signal stream is a speech signal, it may include wind noise signals that are misjudged as speech signals. Thus, as shown in (b) of Figure 9 , the first frequency-domain signals corresponding to the first time-domain signals of the 5th frame to the 8th frame in the first to-be-detected signal stream can be screened out. At the same time, the second frequency-domain signals corresponding to the second time-domain signals of the 5th frame to the 8th frame in the second to-be-detected signal stream, which have the same sequence as the first time-domain signals of the 5th frame to the 8th frame, also need to be determined. Then, the wind noise detection is continued in combination with the first frequency-domain signals and the second frequency-domain signals to distinguish the true speech signal and the wind noise signal.

[0295] As shown in (a) of Figure 10 , the current state signal flag bit SF involved in the first time-domain signal of the 5th frame to the first time-domain signal of the 8th frame determined for the first to-be-detected signal stream is not processed, and only the signal flag bit int corresponding to the tentative state is zeroed. At the same time, the second frame number flag j corresponding to the tentative state as other signals can not be processed, and only the frame number flag i for indicating that the tentative state is a speech signal and the third frame number flag k for indicating that the tentative state is a wind noise signal are subjected to the second initialization processing, so that they are both 0.

[0296] As Figure 10 (b) shown, from the 5th frame of the first time domain signal to carry out wind noise detection, according to the first frequency domain signal, the second frequency domain signal associated with the 5th frame of the first time domain signal, to determine the correlation, the spectral center of gravity and the low frequency energy value corresponding to the 5th frame of the first time domain signal. And determine whether the correlation, the spectral center of gravity and the low frequency energy value meet the second condition? When the correlation, the spectral center of gravity and the low frequency energy value determined by the 5th frame of the first time domain signal do not meet the second condition, the provisional state of the 5th frame of the first time domain signal is determined to be a speech signal, int = 1.

[0297] At this time, since the signal flag SF corresponding to the 5th frame of the first time domain signal is 1; the provisional state and the current state are the same, continue to determine whether the current state is a speech signal, which is a speech signal here. Therefore, the signal flag SF corresponding to the current state of the 5th frame of the first time domain signal remains 1, that is, SF = 1.

[0298] Next, the 6th frame of the first time domain signal is detected for wind noise, and the correlation, the spectral center of gravity and the low frequency energy value corresponding to the 6th frame of the first time domain signal are determined according to the first frequency domain signal, the second frequency domain signal associated with the 6th frame of the first time domain signal. And determine whether the correlation, the spectral center of gravity and the low frequency energy value meet the second condition? When the correlation, the spectral center of gravity and the low frequency energy value determined by the 6th frame of the first time domain signal meet the second condition, the provisional state of the 6th frame of the first time domain signal is determined to be a wind noise signal, int = -1; the current state is a speech signal, SF = 1, the provisional state and the speech state are different, the third frame number flag k used to represent the cumulative frame number of the provisional state as a wind noise signal is updated to 1, that is, k = 1; the value of k is less than the third preset frame number threshold (for example, 4 frames), at this time it can be considered that the number of the provisional state as a wind noise signal is too small, the judgment is unreliable, or it is considered that the wind noise belongs to the interval between words when the user speaks. Then, it is determined that the current state corresponds to a speech signal, and the value of the signal flag SF corresponding to the current state can be maintained, that is, SF = 1.

[0299] The 7th frame of the first time domain signal is detected for wind noise, and the provisional state of the 7th frame of the first time domain signal is determined to be a wind noise signal by using the above method, int = -1; since the provisional state is different from the current state, the third frame number flag k used to represent the cumulative frame number of the provisional state as a wind noise signal is updated to 2, that is, k = 2; the value of k is still less than the third preset frame number threshold, at this time the value of the signal flag SF corresponding to the current state is maintained, that is, SF = 1.

[0300] The wind noise detection is performed on the 8th first time domain signal, the correlation, the spectral center of gravity and the low frequency energy corresponding to the 8th first time domain signal are determined according to the first frequency domain signal and the second frequency domain signal having the correlation relationship with the 8th first time domain signal, and it is determined whether the values of the correlation, the spectral center of gravity and the low frequency energy meet the second condition. When the values of the correlation, the spectral center of gravity and the low frequency energy determined by the 8th first time domain signal do not meet the second condition, it is determined that the tentative state of the 8th first time domain signal is a voice signal, and int=1. Since the tentative state is the same as the current state, it is determined whether the current state is a voice signal, which is a voice signal here. Therefore, the signal flag SF corresponding to the current state of the 8th first time domain signal remains 1, that is, SF=1.

[0301] The following will be described in combination with Figure 11 The interface schematic diagram of the electronic device is described by way of example.

[0302] In a possible implementation, the function of starting the voice detection can be set in the setting interface of the electronic device, and the function of starting the voice detection can be automatically started to execute the voice detection method of the embodiment of the application after the application program for calling in the electronic device is run.

[0303] In another possible implementation, the function of starting the voice detection can be set in the recording application program of the electronic device, and the function of starting the voice detection can be started to execute the voice detection method of the embodiment of the application according to the setting when the audio is recorded.

[0304] In still another possible implementation, the function of starting the voice detection can be automatically started to execute the voice detection method of the embodiment of the application.

[0305] In combination with the third implementation, the function of starting the voice detection of the electronic device is taken as an example, Figure 6 which is an interface schematic diagram of an electronic device provided by the embodiment of the application.

[0306] For example, as Figure 11 shown, the electronic device is taken as a mobile phone as an example, the electronic device displays a lock screen interface 501, as shown in (a) of Figure 11 When the electronic device receives audio data of the user, such as "Hello, YoYo!", the intelligent assistant application program is run, and the voice detection method of the application is automatically executed, and then the key word can be further determined according to the detection result, and the appropriate content is screened from the text library according to the key word to be broadcasted and replied, such as "I am here"; meanwhile, the interface 502 shown in (b) of Figure 11 is displayed.

[0307] When the electronic device receives audio data of the user again, such as "open the map", the interface 503 shown in (c) of

[0307] can be displayed.Figure 11 interface 503 shown in (c) of FIG. 6B; meanwhile, the voice detection method of the present application is automatically performed, and the keyword is further determined according to the detection result, and then the map application is run in response to the keyword, and the homepage 504 in the map application as shown in (d) of FIG. 6B is loaded and displayed. Figure 11

[0308] It should be understood that the above examples are used to help those skilled in the art to understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific values or specific scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.

[0309] The voice detection method and the related display interface of the embodiments of the present application are described above in combination with Figures 1 to 11 The software system, hardware system, device and chip of the electronic device to which the present application is applicable are described in detail below in combination with Figures 12 to 15 It should be understood that the software system, hardware system, device and chip system in the embodiments of the present application can perform the various methods of the foregoing embodiments of the present application, i.e., the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.

[0310] Figure 12 A hardware system of an electronic device suitable for the present application is shown. The electronic device 600 can be used to implement the voice detection method described in the foregoing method embodiments.

[0311] The electronic device 600 can include a processor 610, an external memory interface 620, an internal memory 621, a universal serial bus (USB) interface 630, a charge management module 640, a power management module 641, a battery 642, an antenna 1, an antenna 2, a mobile communication module 650, a wireless communication module 660, an audio module 670, a speaker 670A, a receiver 670B, a microphone 670C, a headset interface 670D, a sensor module 680, a key 690, a motor 691, an indicator 692, a camera 693, a display screen 694, and a subscriber identification module (SIM) card interface 695, etc. The sensor module 680 can include a pressure sensor 680A, a gyroscope sensor 680B, a barometric pressure sensor 680C, a magnetic sensor 680D, an acceleration sensor 680E, a distance sensor 680F, a proximity light sensor 680G, a fingerprint sensor 680H, a temperature sensor 680J, a touch sensor 680K, an ambient light sensor 680L, a bone conduction sensor 680M, etc. ​

[0312] Exemplarily, the audio module 670 is configured to convert digital audio information into an analog audio signal output, and can also be configured to convert an analog audio input into a digital audio signal. The audio module 670 can also be configured to encode and decode audio signals. In some embodiments, the audio module 670 or part of the functions of the audio module 670 can be arranged in the processor 610.

[0313] For example, in the embodiments of the present application, the audio module 670 can send the audio data collected by the microphone to the processor 610.

[0314] It should be noted that, Figure 12 The structures shown do not constitute a specific limitation on the electronic device 600. In some other embodiments of the present application, the electronic device 600 can include more or fewer components than those shown, or the electronic device 600 can include a combination of some of the components shown, or the electronic device 600 can include sub-components of some of the components shown. Figure 12 The structures shown do not constitute a specific limitation on the electronic device 600. In some other embodiments of the present application, the electronic device 600 can include more or fewer components than those shown, or the electronic device 600 can include a combination of some of the components shown, or the electronic device 600 can include sub-components of some of the components shown. Figure 12 The structures shown do not constitute a specific limitation on the electronic device 600. In some other embodiments of the present application, the electronic device 600 can include more or fewer components than those shown, or the electronic device 600 can include a combination of some of the components shown, or the electronic device 600 can include sub-components of some of the components shown. Figure 12 The structures shown do not constitute a specific limitation on the electronic device 600. In some other embodiments of the present application, the electronic device 600 can include more or fewer components than those shown, or the electronic device 600 can include a combination of some of the components shown, or the electronic device 600 can include sub-components of some of the components shown. Figure 12 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0315] The processor 610 can include one or more processing units. For example, the processor 610 can include at least one of the following processing units: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU). Different processing units can be independent devices or integrated devices.

[0316] The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0317] The processor 610 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 610 is a cache memory. The memory can store instructions or data that have just been used or are repeatedly used by the processor 610. If the processor 610 needs to use the instructions or data again, it can directly call them from the memory. This avoids repeated access and reduces the waiting time of the processor 610, thereby improving the efficiency of the system.

[0318] In some embodiments, the processor 610 can include one or more interfaces. For example, the processor 610 can include at least one of an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and a USB interface.

[0319] Exemplarily, the processor 610 can be configured to perform the video processing method of the embodiments of the present application. For example, the processor 610 can be configured to acquire audio data, the audio data being data collected by a first microphone and a second microphone in the same environment; perform VAD detection on the audio data to determine and filter out a voice signal; and perform wind noise detection on the voice signal determined by the VAD detection to determine and output the voice signal.

[0320] Figure 12 The connection relationship between the modules shown is only illustrative and does not constitute a limitation on the connection relationship between the modules of the electronic device 600. Alternatively, the modules of the electronic device 600 can also use a combination of the above-mentioned various connection modes.

[0321] The wireless communication function of the electronic device 600 can be implemented by the antenna 1, the antenna 2, the mobile communication module 650, the wireless communication module 660, the modem processor, and the baseband processor, and the like. The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 600 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.

[0322] In some embodiments, the antenna 1 of the electronic device 600 and the mobile communication module 650 are coupled, and the antenna 2 of the electronic device 600 and the wireless communication module 660 are coupled, so that the electronic device 600 can communicate with the network and other electronic devices through wireless communication technology.

[0323] The electronic device 600 can implement a display function through a GPU, a display screen 694, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 694 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 610 can include one or more GPUs that execute program instructions to generate or change display information.

[0324] The display screen 694 can be used to display images or videos.

[0325] The electronic device 600 can implement a shooting function through an ISP, a camera 693, a video codec, a GPU, a display screen 694, and an application processor, etc.

[0326] The ISP is used to process data fed back by the camera 693. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, and the optical signal is converted into an electrical signal. The camera photosensitive element transmits the electrical signal to the ISP for processing, and converts it into an image visible to the naked eye. The ISP can optimize the noise, brightness, and color of the image through algorithms, and can also optimize the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be arranged in the camera 693.

[0327] The camera 693 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard red green blue (RGB), YUV, etc. format image signal. In some embodiments, the electronic device 600 can include one or N cameras 693, where N is a positive integer greater than 1.

[0328] Exemplarily, in the embodiments of the present application, the voice detection method can be executed in the processor 110.

[0329] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 600 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0330] A video codec is used to compress or decompress digital video. The electronic device 600 can support one or more video codecs. In this way, the electronic device 600 can play or record videos in a variety of formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, and MPEG 4.

[0331] The external memory interface 620 can be used to connect an external memory card, such as a secure digital (SD) card, to expand the memory capacity of the electronic device 600. The external memory card communicates with the processor 610 via the external memory interface 620 to perform data storage functions. For example, files such as music, videos, and the like can be stored in the external memory card.

[0332] The internal memory 621 can be used to store computer-executable program code, which includes instructions. The internal memory 621 can include a program area and a data area.

[0333] The electronic device 600 can implement audio functions, such as music playback and voice recording, via the audio module 670, the speaker 670A, the receiver 670B, the microphone 670C, the earphone interface 670D, and the application processor, and the like.

[0334] The speaker 670A, also known as a loudspeaker, is used to convert audio electrical signals into sound signals. The electronic device 600 can listen to music or engage in hands-free calls via the speaker 670A. The receiver 670B, also known as an earpiece, is used to convert audio electrical signals into sound signals.

[0335] The fingerprint sensor 680H is used to collect a fingerprint. The electronic device 600 can use the collected fingerprint characteristics to implement functions such as unlocking, accessing an application lock, taking a photo, and answering an incoming call, and the like.

[0336] The touch sensor 680K, also known as a touch device. The touch sensor 680K can be disposed on the display screen 694, and the touch sensor 680K and the display screen 694 together form a touch screen, also known as a touch screen. The touch sensor 680K is used to detect a touch operation acting on or near it. The touch sensor 680K can pass the detected touch operation to the application processor to determine the touch event type. Visual output related to the touch operation can be provided via the display screen 694. In other embodiments, the touch sensor 680K can also be disposed on the surface of the electronic device 600 and disposed in a different position from the display screen 694.

[0337] The hardware system of the electronic device 600 is described in detail above, and the software system of the electronic device 600 is introduced below. The software system can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. Embodiments of the present application take the layered architecture as an example to describe the software system of the electronic device 600.

[0338] As shown in Figure 13 , the software system adopting the layered architecture is divided into several layers, each layer has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into four layers, from top to bottom, the application program layer, the application program framework layer, the Android runtime and the system library, and the kernel layer.

[0339] The application program layer can include call, navigation, recording, voice assistant, and the like.

[0340] Exemplarily, the voice detection method provided by the embodiments of the present application can be applied to the call application program; for example, running the call application program, obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment; performing VAD detection on the audio data to determine and filter out the voice signal; performing wind noise detection on the voice signal detected by the VAD to determine and output the voice signal.

[0341] Exemplarily, the voice detection method provided by the embodiments of the present application can be applied to the recording application program; for example, running the recording application program, obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment; performing VAD detection on the audio data to determine and filter out the voice signal; performing wind noise detection on the voice signal detected by the VAD to determine and output the voice signal.

[0342] Exemplarily, the voice detection method provided by the embodiments of the present application can be applied to the navigation assistant application program; for example, running the navigation assistant application program, obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment; performing VAD detection on the audio data to determine and filter out the voice signal; performing wind noise detection on the voice signal detected by the VAD to determine and output the voice signal.

[0343] Exemplarily, the voice detection method provided by the embodiments of the present application can be applied to the voice assistant application program; for example, running the voice assistant application program, obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment; performing VAD detection on the audio data to determine and filter out the voice signal; performing wind noise detection on the voice signal detected by the VAD to determine and output the voice signal.

[0344] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer can include some predefined functions.

[0345] For example, the application framework layer includes a window manager, a content provider, a view system, a phone manager, a resource manager, and a notification manager.

[0346] The window manager is used to manage window programs. The window manager can obtain the size of a display screen, determine whether there is a status bar, a lock screen, and a screenshot.

[0347] The content provider is used to store and obtain data and make the data accessible to applications. The data can include videos, images, audio, dialed and received calls, browsing history and bookmarks, and a phone book.

[0348] The view system includes visual controls, such as a control for displaying text and a control for displaying pictures. The view system can be used to build an application. A display interface can be composed of one or more views, for example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.

[0349] The phone manager is used to provide communication functions of the electronic device, such as management of a call state (on or off).

[0350] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, and video files.

[0351] The notification manager enables an application to display notification information in a status bar, which can be used to convey a message of the notification type and can automatically disappear after a short stay without user interaction.

[0352] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0353] The core library includes two parts: one part is a function function that the java language needs to call, and the other part is the core library of Android.

[0354] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java files of the application layer and the application framework layer into binary files. The virtual machine is used to perform functions such as management of the life cycle of an object, management of a stack, management of a thread, management of security and exceptions, and garbage collection.

[0355] The system library can include a plurality of functional modules, such as a surface manager, media libraries, a three-dimensional graphics processing library (for example, an open graphics library for embedded systems (OpenGL ES) and a 2D graphics engine (for example, a skia graphics library (SGL)).

[0356] The surface manager is used to manage a display subsystem and provides a plurality of applications with fusion of 2D layers and 3D layers.

[0357] The media libraries support playback and recording of a plurality of audio formats, playback and recording of a plurality of video formats, and static image files. The media libraries can support a plurality of audio and video coding formats, such as MPEG4, H.264, moving picture experts group audio layer III (MP3), advanced audio coding (AAC), adaptive multi-rate (AMR), joint photographic experts group (JPG), and portable network graphics (PNG).

[0358] The three-dimensional graphics processing library can be used to implement three-dimensional graphics drawing, image rendering, synthesis, and layer processing.

[0359] The 2D graphics engine is a drawing engine for 2D drawing.

[0360] The kernel layer is a layer between hardware and software. The kernel layer can include driving modules such as an audio driver and a display driver.

[0361] Figure 14 FIG. 7 is a structural schematic diagram of a voice detection device provided by an embodiment of the present application. The voice detection device 700 includes a display unit 710 and a processing unit 720.

[0362] The acquisition unit 710 is configured to acquire audio data, the audio data being data collected by a first microphone and a second microphone in the same environment.

[0363] The processing unit 720 is configured to perform VAD detection on the audio data, determine and filter out a voice signal; perform wind noise detection on the voice signal detected by the VAD, and determine and output the voice signal.

[0364] It should be noted that the voice detection apparatus 700 is embodied in the form of functional units. The term "unit" herein can be implemented in the form of software and / or hardware, and is not specifically limited.

[0365] For example, the "unit" can be a software program, a hardware circuit or a combination of both, which implements the above functions. The hardware circuit can include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combination logic circuit and / or other suitable components supporting the described functions.

[0366] Therefore, the units of each example described in the embodiments of the present application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0367] Figure 15 A structural schematic diagram of an electronic device provided by the present application is shown. Figure 15 The dashed line in the electronic device 800 indicates that the unit or the module is optional, and the electronic device 800 can be used to implement the voice detection method described in the method embodiment.

[0368] The electronic device 800 includes one or more processors 801, which can support the electronic device 800 to implement the method in the method embodiment. The processor 801 can be a general-purpose processor or a dedicated processor. For example, the processor 801 can be a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, such as discrete gates, transistor logic devices or discrete hardware components.

[0369] The processor 801 can be used to control the electronic device 800, execute software programs and process data of the software programs. The electronic device 800 can further include a communication unit 805 to implement input (reception) and output (transmission) of signals.

[0370] For example, the electronic device 800 can be a chip, the communication unit 805 can be an input and / or output circuit of the chip, or the communication unit 805 can be a communication interface of the chip, and the chip can be a component of a terminal device or other electronic device.

[0371] For another example, the electronic device 800 can be a terminal device, the communication unit 805 can be a transceiver of the terminal device, or the communication unit 805 can be a transceiver circuit of the terminal device.

[0372] The electronic device 800 can include one or more memories 802, which store programs 804 that can be executed by the processor 801 to generate instructions 803, so that the processor 801 executes the voice detection method described in the above method embodiments according to the instructions 803.

[0373] Optionally, the memory 802 can also store data. Optionally, the processor 801 can also read the data stored in the memory 802, which can be stored in the same storage address as the program 804, or can be stored in a different storage address from the program 804.

[0374] The processor 801 and the memory 802 can be separately arranged or integrated together, for example, integrated on a system on chip (SOC) of the terminal device.

[0375] For example, the memory 802 can be used to store the related programs 804 of the voice detection method provided in the embodiments of the present application, and the processor 801 can be used to call the related programs 804 of the voice detection method stored in the memory 802 during the transition processing, and execute the voice detection method of the embodiments of the present application. For example: obtaining audio data, the audio data being data collected by the first microphone and the second microphone in the same environment. Performing VAD detection on the audio data to determine and filter out the voice signal; performing wind noise detection on the voice signal detected by the VAD to determine and output the voice signal.

[0376] The present application also provides a computer program product, which, when executed by the processor 801, implements the voice detection method described in any of the method embodiments of the present application.

[0377] The computer program product can be stored in the memory 802, for example, the program 804, which is finally converted into an executable object file that can be executed by the processor 801 through preprocessing, compiling, assembling and linking and other processing processes.

[0378] The application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a computer to implement the voice detection method described in any method embodiment of the application. The computer program can be a high-level language program or an executable target program.

[0379] Optionally, the computer readable storage medium is, for example, the memory 802. The memory 802 can be a volatile memory or a non-volatile memory, or the memory 802 can include both volatile memory and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a synchlink DRAM (SLDRAM), and a direct rambus RAM (DR RAM).

[0380] Those skilled in the art can clearly understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.

[0381] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0382] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the embodiments of the electronic device described above are merely illustrative. For example, the division of the modules is merely a logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0383] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0384] In addition, the functional units in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0385] It should be understood that in various embodiments of the present application, the size of the sequence of each process does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0386] In addition, the term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.

[0387] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0388] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. In summary, the above is only a preferred embodiment of the technical solutions of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A voice detection method characterized by, The method is applied to an electronic device including a first microphone and a second microphone, and the method includes: obtaining audio data, the audio data including a first to-be-tested signal stream collected by the first microphone and data collected by the second microphone in the same environment; preprocessing the audio data includes: framing the first to-be-tested signal stream to obtain a plurality of first time-domain signals; performing time-frequency transformation on the plurality of first time-domain signals to obtain a plurality of first frequency-domain signals; the plurality of first time-domain signals and the plurality of first frequency-domain signals correspond to each other in a one-to-one manner; performing VAD detection on the audio data to determine and filter out a voice signal, including: for the first time-domain signal, determining first data corresponding to the first time-domain signal according to the first time-domain signal and the first frequency-domain signal corresponding to the first time-domain signal, the first data including at least a zero-crossing rate, a spectral entropy, and a flatness; performing VAD detection on the first time-domain signal based on the first data to determine and filter out a voice signal, including: when the first data satisfies a first condition, determining that a tentative state of the first time-domain signal is a voice signal; when the first data does not satisfy the first condition, determining that the tentative state of the first time-domain signal is other signals, the other signals being used to indicate signals other than voice signals and wind noise signals; for the first time-domain signal, determining whether the tentative state is the same as a current state; when different, and the tentative state is a voice signal, a value of a first frame number flag is incremented by 1, and it is determined whether the value of the first frame number flag is greater than a first preset frame number threshold; when the value of the first frame number flag is greater than the first preset frame number threshold, the current state is modified, when the current state is a voice signal, it is modified to other signals, and when the current state is other signals, it is modified to a voice signal; when different, and the tentative state is other signals, a value of a second frame number flag is incremented by 1, and it is determined whether the value of the second frame number flag is greater than a second preset frame number threshold; when the value of the second frame number flag is greater than the second preset frame number threshold, the current state is modified; and first time-domain signals with a modified current state being a voice signal are determined and filtered out; performing wind noise detection on the voice signal detected by the VAD to determine and filter out a voice signal.

2. The voice detection method of claim 1, wherein, The data collected by the second microphone in the same environment includes a second to-be-tested signal stream, and preprocessing the audio data further includes: framing the second to-be-tested signal stream to obtain a plurality of second time-domain signals; performing time-frequency transformation on the plurality of second time-domain signals to obtain a plurality of second frequency-domain signals; the plurality of second time-domain signals and the plurality of second frequency-domain signals correspond to each other in a one-to-one manner.

3. The voice detection method of claim 1, wherein, The VAD detection on the first time-domain signal based on the first data to determine and filter out a voice signal further includes: when the same, first time-domain signals with the current state being a voice signal are determined and filtered out; or, determining and screening the current state as the first time-domain signal of the voice signal when the first frame number flag value is less than or equal to the first preset frame number threshold; or determining and screening the current state as the first time-domain signal of the voice signal when the second frame number flag value is less than or equal to the second preset frame number threshold.

4. The voice detection method of claim 1, wherein, Before the first data satisfies the first condition, the method further comprises: performing a first initialization processing, the first initialization processing at least comprising zeroing the value of the first frame number flag and the value of the second frame number flag.

5. The voice detection method of claim 1, wherein, When the first data comprises the zero-crossing rate, the spectral entropy and the flatness, the first condition comprises: the zero-crossing rate is greater than a zero-crossing rate threshold, the spectral entropy is less than a spectral entropy threshold, and the flatness is less than a flatness threshold.

6. The voice detection method of any of claims 1 to 5, wherein, The wind noise detection on the voice signal detected by the VAD comprises: for the first time-domain signal of the voice signal detected by the VAD, determining second data corresponding to the first time-domain signal according to the first time-domain signal, a first frequency-domain signal corresponding to the first time-domain signal, and a second frequency-domain signal having the same sequence as the first frequency-domain signal, the second data at least comprising a spectral centroid, a low-frequency energy and a correlation; determining the second data, and performing the wind noise detection on the first time-domain signal to determine and screen the voice signal.

7. The voice detection method of claim 6, wherein, The wind noise detection on the first time-domain signal based on the second data to determine and screen the voice signal comprises: when the second data satisfies a second condition, determining that the tentative state of the first time-domain signal is wind noise signal; when the second data does not satisfy the second condition, determining that the tentative state of the first time-domain signal is voice signal; determining whether the tentative state is the same as the current state for the first time-domain signal; when different, and the tentative state is wind noise signal, the value of a third frame number flag is incremented by 1, and it is determined whether the value of the third frame number flag is greater than a third preset frame number threshold; when the value of the third frame number flag is greater than the third preset frame number threshold, modifying the current state, when the current state is voice signal, modifying it to wind noise signal, and when the current state is wind noise signal, modifying it to voice signal; when different, and the tentative state is voice signal, the value of a first frame number flag is incremented by 1, and it is determined whether the value of the first frame number flag is greater than a fourth preset frame number threshold; when the value of the first frame number flag is greater than the fourth preset frame number threshold, modifying the current state; determining and screening the modified current state as the first time-domain signal of the voice signal.

8. The voice detection method of claim 7, wherein, The wind noise detection on the first time-domain signal based on the second data to determine and screen the voice signal further comprises: when the same, determining and screening the current state as the first time-domain signal of the voice signal; or when different, and the value of the third frame number flag is less than or equal to the third preset frame number threshold, determining and screening the current state as the first time-domain signal of the voice signal; or When the difference is greater than zero and the value of the first frame number flag is less than or equal to the fourth preset frame number threshold, it is determined that the current state is a first time-domain signal of a voice signal.

9. The voice detection method of claim 7, wherein, Before the second data satisfies the second condition, the method further includes performing a second initialization process, and the second initialization process at least includes resetting the value of the first frame number flag and the value of the third frame number flag.

10. The voice detection method of claim 7, wherein, When the second data includes a spectral centroid, a low-frequency energy, and a correlation, the second condition includes: the spectral centroid is less than a spectral centroid threshold, the low-frequency energy is greater than a low-frequency energy threshold, and the correlation is less than a correlation threshold.

11. The voice detection method of claim 1, wherein, The first microphone includes one or more first microphones, and / or the second microphone includes one or more second microphones.

12. The voice detection method of claim 1, wherein, The first microphone is a microphone arranged at the bottom of the electronic device, and the second microphone is a microphone arranged at the top or back of the electronic device.

13. An electronic device, comprising: The chip system includes one or more processors, and the processor is configured to invoke a computer instruction to cause the electronic device to execute the voice detection method. The computer readable storage medium stores a computer program, and the computer program includes program instructions, and the program instructions, when executed by a processor, cause the processor to execute the voice detection method. The computer readable storage medium stores a computer program, and the computer program includes program instructions, and the program instructions, when executed by a processor, cause the processor to execute the voice detection method.

14. A chip system, characterized by ​ 15. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Method for processing double-microphone signal

    CN101763858A

  • Noise detection method and device, noise suppression method and device, terminal equipment, system and chip

    CN112669877A