Voice Detection Method, Device, Electronic Device, Storage Medium and Product

By using the target frame spectrum and frequency doubling information of the audio signal in the speech detection technology, the problem of inaccurate speech detection in the low signal-to-noise ratio audio signal is solved, and higher applicability of speech detection is achieved.

CN114724588BActive Publication Date: 2025-06-17SOUNDAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210325910.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-06-17
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

The existing voice detection technology cannot accurately detect voice signals in audio signals with low signal-to-noise ratio, resulting in reduced applicability.

Method used

By acquiring the target frame spectrum of the audio signal, the sampling point with the amplitude maximum value and its frequency multiplication information are determined, and speech detection is performed based on this information.

Benefits of technology

Effectively detecting whether the audio signal is a voice signal avoids the accuracy problems based on short-term energy and zero-crossing rate, and improves the applicability of voice detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724588B_ABST
    Figure CN114724588B_ABST
Patent Text Reader

Abstract

The present application provides a voice detection method, apparatus, electronic device, storage medium and product, belonging to the technical field of voice signal processing. Among them, the voice detection method includes: obtaining a target frame spectrum of an audio signal, where the target frame spectrum includes amplitude values of multiple sampling points of a target frame; determining a first sampling point among the multiple sampling points based on the target frame spectrum, and the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum; obtaining the frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes amplitude values of at least one frequency doubling sampling point of the first sampling point; and performing voice detection on the audio signal based on the frequency doubling information of the first sampling point. This method improves the applicability of voice detection for audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of voice signal processing, and in particular, to a voice detection method, device, electronic device, storage medium, and product. Background Art

[0002] Voice Activity Detection (VAD) technology plays a very important role in voice signal processing fields such as voice enhancement and speech recognition. The VAD technology mainly detects voice signals and non-voice signals from an audio signal.

[0003] In related technologies, the VAD technology generally compares the short-time energy and zero-crossing rate of an audio signal with their corresponding thresholds respectively when the signal-to-noise ratio of the audio signal is relatively high, so as to detect voice signals and non-voice signals. However, for an audio signal with a relatively low signal-to-noise ratio, the noise is large, which makes the interference of the noise in the audio signal on the voice signal large. Furthermore, voice signals and non-voice signals cannot be accurately detected by this method, thus reducing the applicability of the VAD technology. Summary of the Invention

[0004] Embodiments of this application provide a voice detection method, device, electronic device, storage medium, and product, which can improve the applicability of voice detection for audio signals. The technical solution is as follows:

[0005] On the one hand, a voice detection method is provided. The method includes:

[0006] Obtain the target frame spectrum of an audio signal, where the target frame spectrum includes the amplitude values of multiple sampling points of a target frame;

[0007] Based on the target frame spectrum, determine a first sampling point among the multiple sampling points, where the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum;

[0008] Obtain the frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes the amplitude values of at least one frequency doubling sampling point of the first sampling point;

[0009] Based on the frequency doubling information of the first sampling point, perform voice detection on the audio signal.

[0010] In some embodiments, the voice detection of the audio signal based on the frequency doubling information of the first sampling point includes: obtaining the frequency doubling information of at least one second sampling point, where the frequency doubling information of the second sampling point includes the amplitude values of at least one frequency doubling sampling point of the second sampling point, the amplitude value of the second sampling point is the maximum amplitude value in the spectra of other frames, and the spectra of other frames are the spectra adjacent to the target frame spectrum in the audio signal; determining a first target frequency doubling sampling point based on the frequency doubling information of the first sampling point and the frequency doubling information of the at least one second sampling point, where the amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in the multi-frame spectrum composed of the spectra of other frames and the target frame spectrum; determining a target spectrum in the multi-frame spectrum based on the first target frequency doubling sampling point, where the target spectrum includes the target sampling points among the first sampling point and the at least one second sampling point, and the first number of the first target frequency doubling sampling points corresponding to the target sampling points is greater than or equal to a first number threshold; if the second number is greater than or equal to a second number threshold, determining that the audio signal is a voice signal, where the second number is the number of target spectra included in the multi-frame spectrum.

[0011] In some embodiments, the voice detection method further includes: if the second number is less than the second number threshold, obtaining the energy parameter and the zero-crossing rate parameter of the multi-frame spectrum; performing voice detection on the audio signal based on the second number, the energy parameter, and the zero-crossing rate parameter.

[0012] In some embodiments, the energy parameter includes the energy corresponding to each of the multi-frame spectra, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to each of the multi-frame spectra; the performing voice detection on the audio signal based on the second number, the energy parameter, and the zero-crossing rate parameter includes: determining a first difference and a second difference between every two spectra among the multi-frame spectra except the first frame and the last frame based on the energy corresponding to each of the multi-frame spectra and the zero-crossing rate corresponding to each of the multi-frame spectra, where the first difference is the difference in energy between the two spectra, and the second difference is the difference in zero-crossing rate between the two spectra; performing voice detection on the audio signal based on the second number and the multiple first differences and multiple second differences corresponding to the multi-frame spectra.

[0013] In some embodiments, the performing voice detection on the audio signal based on the second number and the multiple first differences and multiple second differences corresponding to the multi-frame spectra includes: performing linear fitting on the multiple first differences to obtain a first fitting parameter; performing linear fitting on the multiple second differences to obtain a second fitting parameter; performing voice detection on the audio signal based on the second number, the first fitting parameter, and the second fitting parameter.

[0014] In some embodiments, the number of non-target spectra in the multi-frame spectrum is a third number. The speech detection of the audio signal based on the second number, the first fitting parameter, and the second fitting parameter includes: if the second number is greater than or equal to a third number threshold, and the first fitting parameter is greater than a first fitting threshold or the second fitting parameter is less than a second fitting threshold, determining that the audio signal is a speech signal; if the third number is less than the second number threshold and greater than or equal to the third number threshold, and the first fitting parameter is less than the first fitting threshold or the second fitting parameter is greater than the second fitting threshold, determining that the audio signal is a non-speech signal.

[0015] In some embodiments, the speech detection of the audio signal based on the second number and the multiple first differences and multiple second differences corresponding to the multi-frame spectrum includes: determining the energy change probabilities corresponding to the multiple first differences respectively and determining the zero-crossing rate change probabilities corresponding to the multiple second differences respectively; performing speech detection on the audio signal based on the second number, the energy change probabilities corresponding to the multiple first differences respectively, and the zero-crossing rate change probabilities corresponding to the multiple second differences respectively.

[0016] In some embodiments, the number of non-target spectra in the multi-frame spectrum is a third number. The speech detection of the audio signal based on the second number, the energy change probabilities corresponding to the multiple first differences respectively, and the zero-crossing rate change probabilities corresponding to the multiple second differences respectively includes: respectively determining a first probability mean and a second probability mean, where the first probability mean is the average of the energy change probabilities of the first target number among the multiple energy change probabilities, and the second probability mean is the average of the energy change probabilities of the second target number among the multiple energy change probabilities; respectively determining a third probability mean and a fourth probability mean, where the third probability mean is the average of the zero-crossing rate change probabilities of the first target number among the multiple zero-crossing rate change probabilities, and the fourth probability mean is the average of the zero-crossing rate change probabilities of the second target number among the multiple zero-crossing rate change probabilities; if the second number is greater than or equal to a third number threshold, and the first probability mean is less than the product of the second probability mean and a second preset ratio or the third probability mean is less than the product of the fourth probability mean and the second preset ratio, determining that the audio signal is a speech signal; if the third number is less than the second number threshold and greater than or equal to the third number threshold, and the product of the second preset ratio and the first probability mean is greater than the second probability mean or the product of the second preset ratio and the third probability mean is greater than the fourth probability mean, determining that the audio signal is a non-speech signal.

[0017] In some embodiments, the voice detection method further includes: if a third quantity of non-target spectra in the multi-frame spectra is greater than or equal to the second quantity threshold, determining that the audio signal is a non-voice signal.

[0018] In some embodiments, performing voice detection on the audio signal based on the frequency doubling information of the first sampling point includes: determining a second target frequency doubling sampling point based on the frequency doubling information of the first sampling point, where the amplitude value of the second target frequency doubling sampling point is the maximum amplitude value in the target frame spectrum; if a fourth quantity of the second target frequency doubling sampling points is greater than or equal to the first quantity threshold, determining that the audio signal is a voice signal.

[0019] In some embodiments, the process of determining the second target frequency doubling sampling point includes: determining two candidate sampling points, where the two candidate sampling points are respectively the sampling points before and after an initial frequency doubling sampling point, and the initial frequency doubling sampling point is the sampling point of the first sampling point at a target multiple; determining the second target frequency doubling sampling point of the first sampling point at the target multiple based on the two candidate sampling points and the initial frequency doubling sampling point.

[0020] In some embodiments, sampling points among the multiple sampling points whose amplitude values are the maximum amplitude values in the target frame spectrum and whose amplitude values are greater than an amplitude threshold are assigned a first value. Determining the second target frequency doubling sampling point of the first sampling point at the target multiple based on the two candidate sampling points and the initial frequency doubling sampling point includes: if the two candidate sampling points and the initial frequency doubling sampling point include sampling points assigned the first value, determining the sampling points assigned the first value as the second target frequency doubling sampling points of the first sampling point at the target multiple.

[0021] In some embodiments, the process of determining the first sampling point includes: determining an initial sampling point, where the amplitude value of the initial sampling point is the maximum amplitude value in the target frame spectrum and the amplitude value is greater than the amplitude threshold; respectively obtaining two auxiliary sampling points corresponding to the initial sampling point, where the two auxiliary sampling points are respectively the nearest sampling point before the sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum and the nearest sampling point after the sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum; if the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold, determining the initial sampling point as the first sampling point.

[0022] In some embodiments, the target frame spectrum includes multiple spectrum maxima. The process of determining the amplitude threshold includes: determining a target maximum value from the multiple spectrum maxima, where the value of the target maximum value is the largest; determining the product of the target maximum value and a first preset ratio to obtain the amplitude threshold.

[0023] On the other hand, a voice detection device is provided, and the device includes:

[0024] A first acquisition module, configured to acquire a target frame spectrum of an audio signal, where the target frame spectrum includes amplitude values of multiple sampling points of a target frame;

[0025] A first determination module, configured to determine a first sampling point among the multiple sampling points based on the target frame spectrum, where the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum;

[0026] A second acquisition module, configured to acquire frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes amplitude values of at least one frequency doubling sampling point of the first sampling point;

[0027] A detection module, configured to perform voice detection on the audio signal based on the frequency doubling information of the first sampling point.

[0028] In some embodiments, the detection module is configured to acquire frequency doubling information of at least one second sampling point, where the frequency doubling information of the second sampling point includes amplitude values of at least one frequency doubling sampling point of the second sampling point, the amplitude value of the second sampling point is the maximum amplitude value in other frame spectra, and the other frame spectra are spectra adjacent to the target frame spectrum in the audio signal; determine a first target frequency doubling sampling point based on the frequency doubling information of the first sampling point and the frequency doubling information of the at least one second sampling point, where the amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in a multi-frame spectrum composed of the other frame spectrum and the target frame spectrum; determine a target spectrum in the multi-frame spectrum based on the first target frequency doubling sampling point, where the target spectrum includes a target sampling point among the first sampling point and the at least one second sampling point, and the first quantity of the first target frequency doubling sampling points corresponding to the target sampling point is greater than or equal to a first quantity threshold; if a second quantity is greater than or equal to a second quantity threshold, determine that the audio signal is a voice signal, where the second quantity is the quantity of target spectra included in the multi-frame spectrum.

[0029] In some embodiments, the detection module is configured to, if the second quantity is less than the second quantity threshold, acquire an energy parameter and a zero-crossing rate parameter of the multi-frame spectrum; perform voice detection on the audio signal based on the second quantity, the energy parameter, and the zero-crossing rate parameter.

[0030] In some embodiments, the energy parameter includes the energy corresponding to each of the multi-frame spectra, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to each of the multi-frame spectra; the detection module is configured to determine, based on the energy corresponding to each of the multi-frame spectra and the zero-crossing rate corresponding to each of the multi-frame spectra, a first difference and a second difference between each adjacent two spectra among the multi-frame spectra except the first frame and the last frame, where the first difference is the difference in energy between the two spectra, and the second difference is the difference in zero-crossing rate between the two spectra; and perform voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multi-frame spectra.

[0031] In some embodiments, the detection module is configured to perform linear fitting on the multiple first differences to obtain a first fitting parameter; perform linear fitting on the multiple second differences to obtain a second fitting parameter; and perform voice detection on the audio signal based on the second quantity, the first fitting parameter, and the second fitting parameter.

[0032] In some embodiments, the number of non-target spectra in the multi-frame spectra is a third quantity, and the detection module is configured to determine that the audio signal is a voice signal if the second quantity is greater than or equal to a third quantity threshold and the first fitting parameter is greater than a first fitting threshold or the second fitting parameter is less than a second fitting threshold;

[0033] If the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the first fitting parameter is less than the first fitting threshold or the second fitting parameter is greater than the second fitting threshold, it is determined that the audio signal is a non-voice signal.

[0034] In some embodiments, the detection module is configured to determine the energy change probability corresponding to each of the multiple first differences and determine the zero-crossing rate change probability corresponding to each of the multiple second differences; and perform voice detection on the audio signal based on the second quantity, the energy change probability corresponding to each of the multiple first differences, and the zero-crossing rate change probability corresponding to each of the multiple second differences.

[0035] In some embodiments, the number of non-target spectra in the multi-frame spectrum is a third number. The detection module is configured to respectively determine a first probability mean value and a second probability mean value. The first probability mean value is the average of the energy change probabilities of the number of energy changes before the target number among the multiple energy change probabilities. The second probability mean value is the average of the energy change probabilities of the number of energy changes after the target number among the multiple energy change probabilities. The third probability mean value and a fourth probability mean value are respectively determined. The third probability mean value is the average of the zero-crossing rate change probabilities of the number of zero-crossing rate changes before the target number among the multiple zero-crossing rate change probabilities. The fourth probability mean value is the average of the zero-crossing rate change probabilities of the number of zero-crossing rate changes after the target number among the multiple zero-crossing rate change probabilities. If the second number is greater than or equal to a third number threshold, and the first probability mean value is less than the product of the second probability mean value and a second preset ratio or the third probability mean value is less than the product of the fourth probability mean value and the second preset ratio, then it is determined that the audio signal is a voice signal. If the third number is less than the second number threshold and greater than or equal to the third number threshold, and the product of the second preset ratio and the first probability mean value is greater than the second probability mean value or the product of the second preset ratio and the third probability mean value is greater than the fourth probability mean value, then it is determined that the audio signal is a non-voice signal.

[0036] In some embodiments, the voice detection device further includes a second determination module, configured to determine that the audio signal is a non-voice signal if the third number of non-target spectra in the multi-frame spectrum is greater than or equal to the second number threshold.

[0037] In some embodiments, the detection module is configured to determine a second target frequency doubling sampling point based on the frequency doubling information of the first sampling point. The amplitude value of the second target frequency doubling sampling point is the maximum amplitude value in the target frame spectrum. If the fourth number of the second target frequency doubling sampling points is greater than or equal to a first number threshold, then it is determined that the audio signal is a voice signal.

[0038] In some embodiments, the voice detection device further includes a third determination module, configured to determine two candidate sampling points, which are respectively the sampling points before and after the initial frequency doubling sampling point. The initial frequency doubling sampling point is the sampling point of the first sampling point at the target multiple. Based on the two candidate sampling points and the initial frequency doubling sampling point, the second target frequency doubling sampling point of the first sampling point at the target multiple is determined.

[0039] In some embodiments, for the sampling points among the multiple sampling points whose amplitude values are the maximum amplitude values in the spectrum of the target frame and the amplitude values are greater than the amplitude threshold, the sampling points are assigned a first value. The third determination module is configured to, if among the two candidate sampling points and the initial multiple-frequency sampling point, there is a sampling point assigned the first value, determine the sampling point assigned the first value as the second target multiple-frequency sampling point of the first sampling point at the target multiple.

[0040] In some embodiments, the voice detection device further includes: a fourth determination module, configured to determine an initial sampling point, where the amplitude value of the initial sampling point is the maximum amplitude value in the spectrum of the target frame and the amplitude value is greater than the amplitude threshold;

[0041] a third acquisition module, configured to respectively acquire two auxiliary sampling points corresponding to the initial sampling point, where the two auxiliary sampling points are respectively the nearest sampling point before the sampling point and with the amplitude value being the maximum amplitude value in the spectrum of the target frame and the nearest sampling point after the sampling point and with the amplitude value being the maximum amplitude value in the spectrum of the target frame; a fifth determination module, configured to, if the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold, determine the initial sampling point as the first sampling point.

[0042] In some embodiments, the voice detection device further includes: a sixth determination module, configured to determine a target maximum value from the multiple spectrum maximum values, where the value of the target maximum value is the largest; a seventh determination module, configured to determine the product of the target maximum value and a first preset ratio to obtain the amplitude threshold.

[0043] On the other hand, an electronic device is provided, where the electronic device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories. The at least one program code is loaded and executed by the one or more processors to implement the voice detection method in any of the above implementation manners.

[0044] On the other hand, a computer-readable storage medium is provided, where at least one program code is stored in the computer-readable storage medium. The at least one program code is loaded and executed by a processor to implement the voice detection method in any of the above implementation manners.

[0045] On the other hand, a computer program product is provided, where the computer program product includes computer program code. The computer program code is stored in a computer-readable storage medium, and a processor of an electronic device reads the computer program code from the computer-readable storage medium. The processor executes the computer program code so that the electronic device executes the voice detection method in any of the above implementation manners.

[0046] An embodiment of the present application provides a voice detection method. Since the amplitude value of the voice signal at the double-frequency sampling point has a similar amplitude maximum characteristic to the amplitude value of its corresponding sampling point, detecting whether an audio signal is a voice signal based on the amplitude value of the double-frequency sampling point of the sampling point in the spectrum can effectively detect whether the audio signal is a voice signal, avoiding the situation that the short-time energy and zero-crossing rate cannot accurately detect audio signals with low signal-to-noise ratio, thereby improving the applicability of voice detection for audio signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 is a schematic diagram of the implementation environment of a voice detection method provided by an embodiment of the present application;

[0049] Figure 2 is a flowchart of a voice detection method provided by an embodiment of the present application;

[0050] Figure 3 is a flowchart of a voice detection method provided by an embodiment of the present application;

[0051] Figure 4 is a flowchart of a voice detection method provided by an embodiment of the present application;

[0052] Figure 5 is a flowchart of a voice detection method provided by an embodiment of the present application;

[0053] Figure 6 is a flowchart of a voice detection method provided by an embodiment of the present application;

[0054] Figure 7 is a schematic diagram of an audio signal and a spectrum provided by an embodiment of the present application;

[0055] Figure 8 is a block diagram of a voice detection device provided by an embodiment of the present application;

[0056] Figure 9 is a block diagram of a terminal provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.

[0058] In the description, claims and drawings of this application, terms such as "first", "second", "third" and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0059] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the audio signals involved in this application are all obtained under full authorization.

[0060] The voice detection method provided by the embodiments of this application can be executed by an electronic device. In some embodiments, the electronic device is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a human-computer interaction device, a vehicle-mounted terminal, etc., but is not limited thereto. Among them, the electronic device can perform voice detection on the audio signal. Those skilled in the art can know that the number of the above-mentioned electronic devices can be more or less. For example, the above-mentioned electronic device can be one, or the above-mentioned electronic devices can be dozens or hundreds, or more. The embodiments of this application do not limit the number and type of the electronic devices.

[0061] In some embodiments, if the voice detection method is applied to a call scenario, then the electronic device 10 is a communication terminal, such as a mobile phone, a landline phone, a walkie-talkie, etc.; correspondingly, the implementation environment of the voice detection method includes the electronic device 10 and the call counterpart 20. During the call between the electronic device 10 and the call counterpart 20, voice detection is performed according to the method provided by the embodiments of this application, so as to reduce the voice coding rate of the electronic device 10 and the call counterpart 20, save communication bandwidth, and reduce device power consumption when there is no voice signal transmission.

[0062] In some embodiments, if the voice detection method is applied to a human-computer interaction scenario, then the electronic device is a human-computer interaction device, such as a smart home appliance, a smart robot, etc.; correspondingly, after the human-computer interaction device obtains the audio signal of the object, it detects the voice signal in the audio signal, and then provides services for the object based on the voice signal.

[0063] In some embodiments, the voice detection method is applied to an audio recording scenario, and the electronic device is an audio recording device. Correspondingly, after the audio recording device obtains the recorded audio signal, it performs voice detection on the audio signal to obtain voice signals and non-voice signals, and then only performs voice recognition on the voice signals to reduce the amount of voice recognition and improve the voice recognition efficiency.

[0064] An embodiment of the present application provides a voice detection method, and the execution subject is an electronic device. Refer to Figure 2 , and the method includes:

[0065] Step 201: Obtain the target frame spectrum of the audio signal, where the target frame spectrum includes the amplitude values of multiple sampling points of the target frame.

[0066] Step 202: Based on the target frame spectrum, determine the first sampling point among the multiple sampling points, and the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum.

[0067] Step 203: Obtain the frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes the amplitude values of at least one frequency doubling sampling point of the first sampling point.

[0068] Step 204: Based on the frequency doubling information of the first sampling point, perform voice detection on the audio signal.

[0069] In some embodiments, performing voice detection on the audio signal based on the frequency doubling information of the first sampling point includes: obtaining the frequency doubling information of at least one second sampling point, where the frequency doubling information of the second sampling point includes the amplitude values of at least one frequency doubling sampling point of the second sampling point, and the amplitude value of the second sampling point is the maximum amplitude value in other frame spectra, and the other frame spectra are the spectra adjacent to the target frame spectrum in the audio signal; based on the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point, determine the first target frequency doubling sampling point, and the amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in the multi-frame spectrum composed of the other frame spectrum and the target frame spectrum; based on the first target frequency doubling sampling point, determine the target spectrum in the multi-frame spectrum, where the target spectrum includes the target sampling points among the first sampling point and at least one second sampling point, and the first quantity of the first target frequency doubling sampling points corresponding to the target sampling points is greater than or equal to the first quantity threshold; if the second quantity is greater than or equal to the second quantity threshold, then determine that the audio signal is a voice signal, and the second quantity is the number of target spectra included in the multi-frame spectrum.

[0070] In some embodiments, the voice detection method further includes: if the second quantity is less than the second quantity threshold, then obtain the energy parameter and the zero-crossing rate parameter of the multi-frame spectrum; based on the second quantity, the energy parameter, and the zero-crossing rate parameter, perform voice detection on the audio signal.

[0071] In some embodiments, the energy parameter includes the energy corresponding to each of multiple frames of spectra, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to each of multiple frames of spectra; performing voice detection on the audio signal based on the second quantity, the energy parameter, and the zero-crossing rate parameter includes: determining a first difference and a second difference between each two adjacent frames of spectra except the first frame and the last frame among the multiple frames of spectra based on the energy corresponding to each of the multiple frames of spectra and the zero-crossing rate corresponding to each of the multiple frames of spectra, where the first difference is the difference in energy between two frames of spectra, and the second difference is the difference in zero-crossing rate between two frames of spectra; performing voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra.

[0072] In some embodiments, performing voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra includes: performing linear fitting on the multiple first differences to obtain a first fitting parameter; performing linear fitting on the multiple second differences to obtain a second fitting parameter; performing voice detection on the audio signal based on the second quantity, the first fitting parameter, and the second fitting parameter.

[0073] In some embodiments, the number of non-target spectra among the multiple frames of spectra is a third quantity, and performing voice detection on the audio signal based on the second quantity, the first fitting parameter, and the second fitting parameter includes: if the second quantity is greater than or equal to a third quantity threshold, and the first fitting parameter is greater than a first fitting threshold or the second fitting parameter is less than a second fitting threshold, then determining that the audio signal is a voice signal; if the third quantity is less than the second quantity threshold and greater than or equal to a third quantity threshold, and the first fitting parameter is less than a first fitting threshold or the second fitting parameter is greater than a second fitting threshold, then determining that the audio signal is a non-voice signal.

[0074] In some embodiments, performing voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra includes: determining the energy change probability corresponding to each of the multiple first differences and determining the zero-crossing rate change probability corresponding to each of the multiple second differences; performing voice detection on the audio signal based on the second quantity, the energy change probability corresponding to each of the multiple first differences, and the zero-crossing rate change probability corresponding to each of the multiple second differences.

[0075] In some embodiments, the number of non-target spectra in the multi-frame spectrum is the third number. Based on the second number, the energy change probabilities corresponding to the multiple first differences, and the zero-crossing rate change probabilities corresponding to the multiple second differences, speech detection is performed on the audio signal, including: respectively determining a first probability mean and a second probability mean, where the first probability mean is the average of the energy change probabilities of the first target number among the multiple energy change probabilities, and the second probability mean is the average of the energy change probabilities of the last target number among the multiple energy change probabilities; respectively determining a third probability mean and a fourth probability mean, where the third probability mean is the average of the zero-crossing rate change probabilities of the first target number among the multiple zero-crossing rate change probabilities, and the fourth probability mean is the average of the zero-crossing rate change probabilities of the last target number among the multiple zero-crossing rate change probabilities; if the second number is greater than or equal to the third number threshold, and the first probability mean is less than the product of the second probability mean and the second preset ratio or the third probability mean is less than the product of the fourth probability mean and the second preset ratio, then it is determined that the audio signal is a speech signal; if the third number is less than the second number threshold and greater than or equal to the third number threshold, and the product of the second preset ratio and the first probability mean is greater than the second probability mean or the product of the second preset ratio and the third probability mean is greater than the fourth probability mean, then it is determined that the audio signal is a non-speech signal.

[0076] In some embodiments, the speech detection method further includes: if the third number of non-target spectra in the multi-frame spectrum is greater than or equal to the second number threshold, then it is determined that the audio signal is a non-speech signal.

[0077] In some embodiments, based on the multiple-frequency information of the first sampling point, speech detection is performed on the audio signal, including: based on the multiple-frequency information of the first sampling point, determining a second target multiple-frequency sampling point, where the amplitude value of the second target multiple-frequency sampling point is the maximum amplitude value in the target frame spectrum; if the fourth number of the second target multiple-frequency sampling points is greater than or equal to the first number threshold, then it is determined that the audio signal is a speech signal.

[0078] In some embodiments, the process of determining the second target multiple-frequency sampling point includes: determining two candidate sampling points, which are respectively the sampling points before and after the initial multiple-frequency sampling point, and the initial multiple-frequency sampling point is the sampling point of the first sampling point at the target multiple; based on the two candidate sampling points and the initial multiple-frequency sampling point, determining the second target multiple-frequency sampling point of the first sampling point at the target multiple.

[0079] In some embodiments, among multiple sampling points, the sampling points whose amplitude values are the maximum amplitude values in the target frame spectrum and whose amplitude values are greater than the amplitude threshold are assigned a first numerical value. Determining the second target octave sampling point of the first sampling point at a target multiple based on two candidate sampling points and an initial octave sampling point includes: if the two candidate sampling points and the initial octave sampling point include a sampling point assigned the first numerical value, determining the sampling point assigned the first numerical value as the second target octave sampling point of the first sampling point at the target multiple.

[0080] In some embodiments, the process of determining the first sampling point includes: determining an initial sampling point, the amplitude value of the initial sampling point being the maximum amplitude value in the target frame spectrum and the amplitude value being greater than the amplitude threshold; respectively obtaining two auxiliary sampling points corresponding to the initial sampling point, the two auxiliary sampling points being the nearest sampling point before the sampling point and having the maximum amplitude value in the target frame spectrum and the nearest sampling point after the sampling point and having the maximum amplitude value in the target frame spectrum; if the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold, determining the initial sampling point as the first sampling point.

[0081] In some embodiments, the target frame spectrum includes multiple spectral maximum values. The process of determining the amplitude threshold includes: determining a target maximum value from the multiple spectral maximum values, the value of the target maximum value being the largest; determining the product of the target maximum value and a first preset ratio to obtain the amplitude threshold.

[0082] An embodiment of the present application provides a voice detection method. Since the amplitude value of a voice signal at an octave sampling point has a similar maximum amplitude characteristic to the amplitude value of its corresponding sampling point, detecting whether an audio signal is a voice signal based on the amplitude value of the octave sampling point of the sampling point in the spectrum can effectively detect whether the audio signal is a voice signal, avoiding the situation that short-time energy and zero-crossing rate cannot accurately detect audio signals with low signal-to-noise ratio, thereby improving the applicability of voice detection for audio signals.

[0083] An embodiment of the present application provides a voice detection method. Refer to Figure 3 , the voice detection method includes:

[0084] Step 301: An electronic device obtains the target frame spectrum of an audio signal.

[0085] Wherein, the audio signal includes at least one of a voice signal and a non-voice signal. The target frame spectrum includes the amplitude values of multiple sampling points of the target frame. In one implementation, the electronic device has an audio acquisition component, such as a microphone, for obtaining the audio signal.

[0086] In some embodiments, the process of obtaining the target frame spectrum includes: the electronic device frames the audio signal to obtain multiple frames of signals of the audio signal, and obtains the target frame spectrum based on the multiple frames of signals; wherein, each frame of signal includes signal values of multiple sampling points, and the signal value is used to represent the intensity of the audio signal at the sampling point; optionally, the signal value is a decibel value. Among them, the target frame spectrum can be determined as needed; if the audio signal is an audio signal during a call, the target frame spectrum is the spectrum corresponding to the audio signal in the current call; if the audio signal is a previously generated voice signal, the target frame spectrum can be the spectrum corresponding to any frame of the audio signal to be detected.

[0087] Optionally, the process of the electronic device framing the audio signal includes: the electronic device determines the frame length and frame shift for framing, and frames the audio signal based on the frame length and frame shift to obtain multiple frames of signals. Optionally, the frame length is 256 and the frame shift is 160, which are not specifically limited herein. Among them, the frame length of 256 refers to the length of 256 sampling points, and the frame shift of 160 refers to moving 160 sampling points for each frame. For example, if the sampling frequency of the audio signal is 16k, that is, 16000 sampling point signal values are collected in 1 second, then the signal values of every 256 sampling points are used as one frame of signal, and then the signal values of 160 sampling points in these 256 sampling points are updated each time to obtain the next frame of signal, that is, if the signal values from the 1st sampling point to the 256th sampling point correspond to the first frame of signal, then the signal values from the 161st sampling point to the 256 + 160th sampling point correspond to the second frame of signal, and so on, to implement the framing of the audio signal and obtain multiple frames of signals of the audio signal. In this way, by overlapping the sampling points of the multiple frames of signals, a smooth transition of the multiple frames of signals is achieved.

[0088] Correspondingly, the process of the electronic device obtaining the target frame spectrum based on the multiple frames of signals includes that the electronic device determines the target frame signal to be detected from the multiple frames of signals, and transforms the target frame signal to obtain the target frame spectrum corresponding to the target frame signal. Optionally, the process of the electronic device transforming the target frame signal includes the following steps: the electronic device sequentially performs pre-emphasis, windowing, and Fourier transform on the target frame signal to be detected to obtain the first single-sided spectrum corresponding to the target frame signal as the target frame spectrum; wherein, the first single-sided spectrum is the spectrum drawn by the electronic device based on the single-sideband modulation technique.

[0089] In some other embodiments, the process of obtaining the target frame spectrum includes: the electronic device transforms the audio signal to obtain the audio spectrum corresponding to the audio signal, and frames the audio spectrum to obtain multiple frames of spectra; then the electronic device determines the spectrum corresponding to the target frame signal to be detected in the multiple frames of spectra to obtain the target frame spectrum.

[0090] Optionally, the process of the electronic device transforming the audio signal includes the following steps. The electronic device sequentially performs pre-emphasis, windowing, and Fourier transform on the audio signal to obtain the second single-sided spectrum corresponding to the audio signal, and the second single-sided spectrum includes multiple frames of spectra. Accordingly, the electronic device determines the spectrum corresponding to the target frame signal to be detected in the second single-sided spectrum to obtain the target frame spectrum. Among them, the process of the electronic device framing the audio spectrum is the same as the process of framing the audio signal, and will not be elaborated here.

[0091] Step 302: The electronic device determines the first sampling point among multiple sampling points based on the target frame spectrum.

[0092] Among them, the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum. In some embodiments, the electronic device determines the sampling point whose amplitude value is the maximum amplitude value in the target frame spectrum and the amplitude value of the target frame spectrum is greater than the amplitude threshold as the first sampling point; since there is a transition process from small to large in the sampling frequency when collecting the audio signal, the sampling frequencies of the first two sampling points are relatively low. Therefore, in some embodiments, the electronic device removes the first two sampling points among the multiple sampling points and determines the first sampling point from the remaining sampling points. If there are 256 sampling points, the electronic device determines the first sampling point from the 3rd sampling point to the 256th sampling point to avoid the influence of the first two sampling points on the first sampling point due to the too low sampling frequency.

[0093] In some other embodiments, the determination process of the first sampling point includes the following steps: The electronic device determines the initial sampling point, and the amplitude value of the initial sampling point is the maximum amplitude value in the target frame spectrum and the amplitude value is greater than the amplitude threshold. The electronic device respectively obtains two auxiliary sampling points corresponding to the initial sampling point, and the two auxiliary sampling points are respectively the nearest sampling point before the sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum and the nearest sampling point after the sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum; if the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold, the initial sampling point is determined as the first sampling point. If the amplitude value of at least one of the two auxiliary sampling points is less than or equal to the amplitude threshold, the electronic device determines that the initial sampling point is not the first sampling point. It should be noted that there are multiple initial sampling points, and the amplitude values of the multiple initial sampling points are all the maximum amplitude values in the target frame spectrum and the amplitude values are greater than the amplitude threshold. Accordingly, the electronic device determines one or more first sampling points from the multiple initial sampling points.

[0094] Among them, for any amplitude value Xn, n represents the position of the sampling point, and the maximum amplitude value satisfies X n-1 <X n And X n >X n+1For example, if the initial sampling point is the 5th sampling point, its amplitude value X5 is the maximum amplitude value in the target frame spectrum, and X5 is greater than the amplitude threshold, then two auxiliary sampling points corresponding to the initial sampling point are determined. The two auxiliary sampling points are respectively the nearest sampling point before the sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum, and the nearest sampling point after the sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum. If the amplitude values of these two auxiliary sampling points are both greater than the amplitude threshold, then the 5th sampling point is determined as the first sampling point.

[0095] In some embodiments, the electronic device respectively obtains two auxiliary sampling points before the initial sampling point and two auxiliary sampling points after the initial sampling point. The amplitude values of these four auxiliary sampling points are the maximum amplitude values in the target frame spectrum. If the amplitude values of these four nearest sampling points are all greater than the amplitude threshold, then the initial sampling point is determined as the first sampling point.

[0096] In the embodiments of the present application, the electronic device determines the first sampling point through the amplitude value of the nearest sampling point before the initial sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum and the amplitude value of the nearest sampling point after the initial sampling point and with the amplitude value being the maximum amplitude value in the target frame spectrum, which ensures that the amplitude value of the first sampling point is the maximum amplitude value and the situation where it is greater than the amplitude threshold is accurate, and avoids the inaccurate determination of the first sampling point caused by the accidental occurrence of the maximum amplitude value at the initial sampling point.

[0097] In some embodiments, the target frame spectrum includes multiple spectrum maximum values. The process of determining the amplitude threshold includes the following steps: The electronic device determines a target maximum value from the multiple spectrum maximum values, and the value of the target maximum value is the largest; The electronic device determines the product of the target maximum value and a first preset ratio to obtain the amplitude threshold.

[0098] Optionally, the first preset ratio is any ratio between 0.5 and 0.95, such as the first preset ratio is 0.5, the first preset ratio is 0.95, or the first preset ratio is 0.8, and no specific limitation is made here. In the embodiments of the present application, by determining the amplitude threshold based on the maximum value among the multiple spectrum maximum values, the obtained amplitude threshold is more in line with the actual situation of the amplitude values of multiple sampling points, and thus it is convenient to subsequently determine the first sampling point among the multiple sampling points based on the amplitude threshold.

[0099] Step 303: The electronic device obtains the frequency doubling information of the first sampling point.

[0100] Among them, the harmonic frequency information of the first sampling point includes the amplitude values of at least one harmonic frequency sampling point of the first sampling point. Optionally, the sampling points in the at least one harmonic frequency sampling point are adjacent in sequence. Optionally, the at least one harmonic frequency sampling point is 2, including the 2-fold sampling point and the 3-fold sampling point of the first sampling point respectively; the at least one harmonic frequency sampling point is 3, including the 2-fold sampling point, the 3-fold sampling point and the 4-fold sampling point of the first sampling point respectively; if the first sampling point is the 5th sampling point, its at least one harmonic frequency sampling point may include the 2-fold sampling point, the 3-fold sampling point and the 4-fold sampling point, etc., that is, the 10th sampling point, the 15th sampling point and the 20th sampling point, etc. In the embodiments of the present application, the number of at least one harmonic frequency sampling point can be set and changed as needed, and no specific limitation is made here.

[0101] Step 304: The electronic device determines a second target harmonic frequency sampling point based on the harmonic frequency information of the first sampling point.

[0102] Among them, the amplitude value of the second target harmonic frequency sampling point is the maximum amplitude value in the target frame spectrum. In some embodiments, the process of determining the second target harmonic frequency sampling point includes the following steps: The electronic device determines two candidate sampling points, which are the sampling points before and after the initial harmonic frequency sampling point respectively, and the initial harmonic frequency sampling point is the sampling point of the first sampling point at the target multiple. The electronic device determines the second target harmonic frequency sampling point of the first sampling point at the target multiple based on the two candidate sampling points and the initial harmonic frequency sampling point. Optionally, the target multiple can be at least one of 2-fold, 3-fold, 3-fold, 4-fold, 5-fold, etc. For example, when the target multiple is 2-fold and the first sampling point is the 5th sampling point, the initial harmonic frequency sampling point is the 10th sampling point, and the two candidate sampling points are the 9th sampling point and the 11th sampling point respectively.

[0103] In some other embodiments, the electronic device determines four candidate sampling points, which are the two sampling points before and the two sampling points after the initial harmonic frequency sampling point respectively; the electronic device determines the second target harmonic frequency sampling point of the first sampling point at the target multiple based on the four candidate sampling points and the initial harmonic frequency sampling point. In some embodiments, the sampling points among multiple sampling points whose amplitude values are the maximum amplitude values in the target frame spectrum and whose amplitude values are greater than the amplitude threshold are assigned a first value; the process of the electronic device determining the second target harmonic frequency sampling point of the first sampling point at the target multiple based on the four candidate sampling points and the initial harmonic frequency sampling point includes: If the four candidate sampling points and the initial harmonic frequency sampling point include the sampling points assigned the first value, the electronic device determines the sampling points assigned the first value as the second target harmonic frequency sampling point of the first sampling point at the target multiple.

[0104] In this embodiment, the second target frequency doubling sampling point of the first sampling point is determined by the sampling points before and after the first sampling point at the target multiple sampling point, fully considering the situation where the position of the second target frequency doubling sampling point deviates, so that even if the position of the second target frequency doubling sampling point deviates, the second target frequency doubling sampling point can be determined based on the sampling points before and after it, thereby improving the rationality of the determined second target frequency doubling sampling point.

[0105] In some embodiments, among multiple sampling points, the sampling points whose amplitude values are the maximum amplitude values in the target frame spectrum and whose amplitude values are greater than the amplitude threshold are assigned a first value. The electronic device determines the second target frequency doubling sampling point of the first sampling point at the target multiple based on two candidate sampling points and an initial frequency doubling sampling point, including the following steps: If the two candidate sampling points and the initial frequency doubling sampling point include sampling points assigned the first value, the electronic device determines the sampling point assigned the first value as the second target frequency doubling sampling point of the first sampling point at the target multiple.

[0106] Optionally, the target multiple can be at least one of 2 times, 3 times, 3 times, 4 times, 5 times, etc. For example, if the target multiple is 2 times and the first sampling point is the 5th sampling point, then the initial frequency doubling sampling point is the 10th sampling point, and the two candidate sampling points are the 9th sampling point and the 11th sampling point respectively. If the 9th sampling point, the 11th sampling point, and the 10th sampling point include sampling points assigned the first value, then the sampling point assigned the first value is determined as the second target frequency doubling sampling point of the first sampling point at the target multiple. For example, if the 9th sampling point is assigned the first value, then the 9th sampling point is the second target frequency doubling sampling point of the 5th sampling point at 2 times; if the 11th sampling point is assigned the first value, then the 11th sampling point is the second target frequency doubling sampling point of the 5th sampling point at 2 times. If the 10th sampling point is assigned the first value, then the 10th sampling point is the second target frequency doubling sampling point of the 5th sampling point at 2 times. In some embodiments, there can be multiple second target frequency doubling sampling points of the first sampling point at the target multiple, that is, at least two sampling points among the two candidate sampling points and the initial frequency doubling sampling point are both assigned the first value, and no specific limitation is made here.

[0107] Wherein, the first value is used to mark that the amplitude value of the sampling point is the maximum amplitude value in the target spectrum and its amplitude value is greater than the amplitude threshold. The specific value of the first value can be set and changed as needed, and no specific limitation is made here; for example, the first value is 1.

[0108] In the embodiments of the present application, by using the sampling points that meet the amplitude value condition before and after the initial frequency doubling sampling point as the second target frequency doubling sampling point of the first sampling point, the existence of the position deviation of the second target frequency doubling sampling point is allowed, thereby improving the rationality of the determined second target frequency doubling sampling point.

[0109] Step 305: If the fourth quantity of the second target frequency - doubled sampling points is greater than or equal to the first quantity threshold, the electronic device determines that the audio signal is a voice signal.

[0110] Among them, the first quantity threshold can be set and adjusted as needed; optionally, the first quantity threshold is 3. In some embodiments, the second target frequency - doubled sampling points of the fourth quantity are the sequentially adjacent frequency - doubled sampling points starting from the 2 - fold sampling point of the first sampling point. For example, if the fourth quantity is 3, the second target frequency - doubled sampling points of the fourth quantity are the 2 - fold sampling point, 3 - fold sampling point, and 4 - fold sampling point of the first sampling point in sequence.

[0111] In some embodiments, the target frame spectrum includes multiple first sampling points. If the fourth quantity of the second target frequency - doubled sampling points corresponding to at least one of the multiple first sampling points is greater than or equal to the first quantity threshold, the electronic device determines that the audio signal is a voice signal. If the fourth quantity of the second target frequency - doubled sampling points corresponding to each of the multiple first sampling points is less than the first quantity threshold, the electronic device determines that the audio signal is a non - voice signal.

[0112] In one implementation, the electronic device assigns a second value to the sampling points whose amplitude value is the maximum amplitude value but the amplitude value is less than or equal to the amplitude threshold; optionally, the sampling points whose amplitude value is not the maximum amplitude value among the multiple sampling points are also assigned the second value, and the second value is used to mark other sampling points that are not the first sampling points. Among them, the second value is different from the first value. For example, if the first value is 1, the second value is 0. In this implementation, by assigning the first value to the sampling points whose amplitude value is the maximum amplitude value in the target spectrum and whose amplitude value is greater than the amplitude threshold, the electronic device can quickly determine the first sampling point from multiple sampling points, thereby improving the efficiency of determining the first sampling point.

[0113] The embodiment of the present application provides a voice detection method. Since the voice signal has a similar maximum - amplitude characteristic in the amplitude value of the frequency - doubled sampling point and the amplitude value of its corresponding sampling point, detecting whether the audio signal is a voice signal based on the amplitude value of the frequency - doubled sampling point of the sampling point in the spectrum can effectively detect whether the audio signal is a voice signal, avoiding the situation that the short - time energy and zero - crossing rate cannot accurately detect the audio signal with a low signal - to - noise ratio, thereby improving the applicability of voice detection for audio signals.

[0114] The embodiment of the present application provides a voice detection method. Refer to Figure 4 , and this voice detection method includes:

[0115] Step 401: The electronic device obtains the target frame spectrum of the audio signal.

[0116] This step is the same as step 301 and will not be elaborated here.

[0117] Step 402: The electronic device determines a first sampling point among multiple sampling points based on the target frame spectrum.

[0118] This step is the same as step 302 and will not be elaborated here.

[0119] Step 403: The electronic device obtains the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0120] Among them, the frequency doubling information of the first sampling point includes the amplitude values of at least one frequency doubling sampling point of the first sampling point, and the frequency doubling information of the second sampling point includes the amplitude values of at least one frequency doubling sampling point of the second sampling point. The amplitude value of the second sampling point is the maximum amplitude value in other frame spectra, and the other frame spectra are the spectra adjacent to the target frame spectrum in the audio signal. Optionally, the other frame spectra are the spectra before the target frame spectrum in the audio signal or the spectra after the target frame spectrum in the audio signal, which can be set and changed as needed and are not specifically limited here.

[0121] Among them, there is at least one other frame spectrum. If the other frame spectra are the spectra before the target frame spectrum in the audio signal, then at least one other frame spectrum is the continuous spectrum adjacent to and before the target frame spectrum, and the number of at least one other frame spectrum can be set and changed as needed. For example, if there are 9 other frame spectra and the target frame spectrum is the 10th frame spectrum, then the other frame spectra are the 9th frame spectrum, the 8th frame spectrum... the 1st frame spectrum respectively. The multiple frame spectra are respectively represented as M10, M9, M8,..., M1.

[0122] The step for the electronic device to obtain the frequency doubling information of the first sampling point is the same as step 303 and will not be elaborated here. The step for the electronic device to obtain the frequency doubling information of the second sampling point is the same as the step for the electronic device to obtain the frequency doubling information of the first sampling point and will not be elaborated here.

[0123] Step 404: The electronic device determines a first target frequency doubling sampling point based on the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0124] Among them, the amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in the multiple frame spectra composed of the other frame spectra and the target frame spectrum. The step for the electronic device to determine the first target frequency doubling sampling point in this step is the same as the step for the electronic device to determine the second target frequency doubling sampling point in step 304 and will not be elaborated here.

[0125] Step 405: The electronic device determines the target spectrum in the multiple frame spectra based on the first target frequency doubling sampling point.

[0126] Among them, the target spectrum includes target sampling points among the first sampling point and at least one second sampling point, and the first quantity of the first target multiple-frequency sampling points corresponding to the target sampling points is greater than or equal to the first quantity threshold.

[0127] Step 406: If the second quantity is greater than or equal to the second quantity threshold, the electronic device determines that the audio signal is a voice signal, where the second quantity is the quantity of target spectra included in multiple frames of spectra.

[0128] In some embodiments, starting from the target frame spectrum, the electronic device determines that the second quantity of the target spectra that are successively adjacent to it before is greater than or equal to the second quantity threshold, and then determines that the audio signal is a voice signal, that is, the second quantity is the quantity of consecutive target spectra starting from the target frame spectrum. For example, if the second quantity is 8, the target frame spectra with the second quantity are M10, M9, M8, M7, M6, M5, M4, M3 in sequence.

[0129] Among them, the second quantity threshold can be set and changed as needed, and no specific limitation is made here; optionally, the second quantity threshold is 8. In one implementation, the second quantity threshold is associated with the quantity of multiple frames of spectra, and the second quantity threshold increases as the quantity of multiple frames of spectra increases. For example, if the quantity of multiple frames of spectra is 10, the second quantity threshold is 8; if the quantity of multiple frames of spectra is 20, the second quantity threshold is 16.

[0130] In some embodiments, if the third quantity of non-target spectra in multiple frames of spectra is greater than or equal to the second quantity threshold, it is determined that the audio signal is a non-voice signal.

[0131] Among them, the electronic device determines the non-target spectra in multiple frames of spectra based on the first target multiple-frequency sampling points; the non-target spectra do not include target sampling points. In one implementation, starting from the target frame spectrum, the electronic device determines that the third quantity of the non-target spectra that are successively adjacent to it before is greater than or equal to the second quantity threshold, and then determines that the audio signal is a voice signal, that is, the third quantity is the quantity of consecutive non-target spectra starting from the target frame spectrum. For example, if the third quantity is 8, the non-target frame spectra with the third quantity are M10, M9, M8, M7, M6, M5, M4, M3 in sequence.

[0132] In some embodiments, after the electronic device determines the target spectrum, it marks the target spectrum with a first identifier, which is convenient for the electronic device to determine the second quantity of the target spectrum. Optionally, the electronic device also marks the non-target spectrum with a second identifier, which is convenient for the electronic device to determine the third quantity of the non-target spectrum. Marking the target spectrum and the non-target spectrum in this way can improve the efficiency of detecting the audio signal. Among them, the first identifier and the second identifier are different, and both the first identifier and the second identifier can be text identifiers, digital identifiers or letter identifiers, and no specific limitation is made here. For example, the first identifier is 1 and the second identifier is 0.

[0133] In some embodiments, if the second quantity of the target spectrum in the multi-frame spectrum is greater than or equal to the second quantity threshold, the electronic device marks the multi-frame spectrum with a third identifier; if the third quantity of the non-target spectrum in the multi-frame spectrum is greater than or equal to the second quantity threshold, the electronic device marks the multi-frame spectrum with a fourth identifier, which facilitates subsequent voice detection of the audio signal by the electronic device based on the identifier of the multi-frame spectrum. Among them, the third identifier and the fourth identifier are different, and both the third identifier and the fourth identifier can be text identifiers, digital identifiers or letter identifiers, which are not specifically limited herein. For example, if the third identifier is 2, the fourth identifier is -2. Optionally, the multi-frame spectrum identifier is represented as Fm. When Fm = 2, it means that the multi-frame spectrum is marked with the third identifier, and when Fm = -2, it means that the multi-frame spectrum is marked with the fourth identifier. Correspondingly, if there are more than or equal to the second quantity threshold of spectra in the multi-frame spectrum that are continuously marked with the first identifier, it is determined that Fm = 2; if there are more than or equal to the second quantity threshold of spectra in the multi-frame spectrum that are continuously marked with the second identifier, it is determined that Fm = -2, which facilitates directly determining whether the audio signal is a voice signal based on the value of the identifier Fm.

[0134] In the embodiments of the present application, voice detection is performed on the audio signal through the frequency doubling information of the sampling points in the multi-frame spectrum. Since the amount of data of the frequency doubling information of the sampling points in the multi-frame spectrum is large, performing voice detection on the audio signal based on the frequency doubling information of the sampling points in the multi-frame spectrum can improve the accuracy of voice detection of the audio signal.

[0135] The embodiments of the present application provide a voice detection method. Refer to Figure 5 and this voice detection method includes:

[0136] Step 501: The electronic device acquires the target frame spectrum of the audio signal.

[0137] Step 502: The electronic device determines the first sampling point among multiple sampling points based on the target frame spectrum.

[0138] Step 503: The electronic device acquires the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0139] Step 504: The electronic device determines the first target frequency doubling sampling point based on the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0140] Step 505: The electronic device determines the target spectrum in the multi-frame spectrum based on the first target frequency doubling sampling point.

[0141] Steps 501-505 are the same as steps 401-405 and will not be elaborated here.

[0142] Step 506: If the second quantity is less than the second quantity threshold, the electronic device acquires the energy parameter and the zero-crossing rate parameter of multiple frames of spectra.

[0143] Among them, the energy parameter includes the energy corresponding to multiple frames of spectra respectively, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to multiple frames of spectra respectively. In some embodiments, the energy corresponding to each frame of spectrum is short-time energy respectively. The process by which the electronic device acquires the energy parameter of multiple frames of spectra includes the following steps: The electronic device acquires the audio signal of the current frame to be processed, and performs pre-emphasis, windowing, and Fourier transform on the audio signals corresponding to multiple sampling points of the current frame in sequence to obtain the single-sided spectrum corresponding to the audio signal of the current frame, squares the single-sided spectrum to obtain a power spectrum with the length of multiple sampling points, and then accumulates and averages the energy corresponding to each sampling point in the power spectrum to obtain the energy corresponding to this frame of spectrum. In another implementation manner, the electronic device directly squares the single-sided spectrum obtained through the processing in step 301 to obtain a power spectrum with the length of multiple sampling points.

[0144] In one implementation manner, the electronic device removes the energy of the first two sampling points among multiple sampling points, and averages the energy of the remaining multiple sampling points to obtain the energy corresponding to the current frame of spectrum. If there are 256 sampling points, the electronic device accumulates and averages the energy from the 3rd sampling point to the 256th sampling point to obtain the energy corresponding to the current frame of spectrum; this avoids the influence of the first two sampling points on the average energy of multiple sampling points due to too low sampling frequency, and further ensures the accuracy of the energy corresponding to the current frame of spectrum.

[0145] In some embodiments, the process by which the electronic device acquires the zero-crossing rate parameter of multiple frames of spectra includes the following steps: For each frame of spectrum, the electronic device determines in sequence whether the product of two adjacent sampling points is less than 0 based on the signal values of multiple sampling points of the audio signal corresponding to this frame of spectrum, that is, determines whether the signal value of the nth sampling point multiplied by the signal value of the (n - 1)th sampling point is less than 0. In one implementation manner, the electronic device uses a counter for counting. If the product of two adjacent sampling points is less than 0, the counter increments by 1, and accumulates in sequence. Finally, the number of counts obtained is the zero-crossing rate.

[0146] Step 507: The electronic device determines the first difference and the second difference between each adjacent two frames of spectra except the first frame and the last frame among multiple frames of spectra based on the energy corresponding to multiple frames of spectra respectively and the zero-crossing rate corresponding to multiple frames of spectra respectively.

[0147] Among them, the first difference is the energy difference between two frames of spectra, and the second difference is the zero-crossing rate difference between two frames of spectra.

[0148] Optionally, the first difference is the difference obtained by subtracting the energy of the previous frame spectrum from the energy of the subsequent frame spectrum. For example, if the energies corresponding to 10 frame spectra are respectively denoted as E10, E9, E8, E7, ……, E2, E1, representing the energies of the 10th frame spectrum, the 9th frame spectrum, ……, the 1st frame spectrum respectively, then the multiple first differences are respectively de9 = E10 - E9, de9 = E9 - E8, ……, de1 = E2 - E1. The second difference is the difference obtained by subtracting the zero-crossing rate of the previous frame spectrum from the zero-crossing rate of the subsequent frame spectrum. For example, if the zero-crossing rates corresponding to multiple frame spectra are respectively denoted as Z10, Z9, Z8, Z7, ……, Z2, Z1, representing the zero-crossing rates of the 10th frame spectrum, the 9th frame spectrum, ……, the 1st frame spectrum respectively, then the multiple second differences are respectively dz9 = Z10 - Z9, dz9 = Z9 - Z8, ……, dz1 = Z2 - Z1.

[0149] Step 508: The electronic device performs linear fitting on the multiple first differences to obtain a first fitting parameter; and performs linear fitting on the multiple second differences to obtain a second fitting parameter.

[0150] Among them, the first fitting parameter is the slope k1 obtained by performing linear fitting on the multiple first differences, and the second fitting parameter is the slope k2 obtained by performing linear fitting on the multiple second differences.

[0151] Step 509: The electronic device performs voice detection on the audio signal based on the second quantity, the first fitting parameter, and the second fitting parameter.

[0152] In one implementation, if the second quantity is greater than or equal to the third quantity threshold, and the first fitting parameter is greater than the first fitting threshold or the second fitting parameter is less than the second fitting threshold, the electronic device determines that the audio signal is a voice signal. In another implementation, the number of non-target spectra in multiple frame spectra is the third quantity. If the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the first fitting parameter is less than the first fitting threshold or the second fitting parameter is greater than the second fitting threshold, the electronic device determines that the audio signal is a non-voice signal.

[0153] Among them, the third quantity threshold is less than the second quantity threshold, and the third quantity threshold can be set and changed as needed, and no specific limitation is made here; optionally, the third quantity threshold is 3. Among them, the third quantity threshold is associated with the number of multiple frame spectra and is less than the second quantity threshold, and the third quantity threshold increases as the number of multiple frame spectra increases. For example, if the number of multiple frame spectra is 10, the second quantity threshold is 8, and the third quantity threshold is 3; if the number of multiple frame spectra is 20, the second quantity threshold is 16, and the third quantity threshold is 6.

[0154] Among them, the first fitting threshold and the second fitting threshold can be set and changed as needed. For example, the first fitting threshold is 1 and the second fitting threshold is -1. In this way, if the first fitting parameter is greater than 1, it indicates that the energy of multiple frames of spectra gradually increases. If the second fitting parameter is less than -1, it indicates that the zero-crossing rate of multiple frames of spectra gradually decreases. Similarly, if the first fitting parameter is less than -1, it indicates that the energy of multiple frames of spectra decreases in turn. If the second fitting parameter is greater than 1, it indicates that the zero-crossing rate of multiple frames of spectra gradually increases. In the embodiments of the present application, since the short-time energy of the voice signal is large and the zero-crossing rate is small, and since the second quantity is between the third quantity threshold and the second quantity threshold, it only represents that the audio signal may be a voice signal. Furthermore, if the first fitting parameter is greater than the first fitting threshold or the second fitting parameter is less than the second fitting threshold, it can be determined that the audio signal is a voice signal. If the first fitting parameter is less than the first fitting threshold or the second fitting parameter is greater than the second fitting threshold, it can be determined that the audio signal is a non-voice signal, thereby improving the accuracy of voice signal detection.

[0155] In the embodiments of the present application, since the first fitting parameter obtained by linearly fitting multiple first differences and the second fitting parameter obtained by linearly fitting multiple second differences can respectively effectively characterize the short-time energy change trend and the zero-crossing rate change trend of multiple frames of spectra, detecting the voice of the audio signal based on the first fitting parameter and the second fitting parameter can improve the reliability of the detection.

[0156] In some embodiments, the electronic device marks multiple frames of spectra with the second quantity greater than or equal to the third quantity threshold as the fifth identifier, and also marks multiple frames of spectra with the third quantity less than the second quantity threshold and greater than or equal to the third quantity threshold as the fifth identifier. The fifth identifier is used to indicate that the audio signal cannot be directly determined as a voice signal or a non-voice signal based on the frequency doubling information of multiple frames of spectra. In this way, detecting the audio signal can be directly performed by marking the frequency doubling information of multiple frames of spectra with the third identifier and the fourth identifier, and detecting the audio signal cannot be directly performed based on the fifth identifier marking the frequency doubling information of multiple frames of spectra. Furthermore, it is convenient to subsequently obtain the energy parameter and the zero-crossing rate parameter of the multiple frames of spectra marked with the fifth identifier for voice detection, thereby improving the efficiency of voice detection.

[0157] In some embodiments, the electronic device marks the first fitting identifier Fe1 for multiple frames of spectrum based on the first fitting parameter; wherein, the electronic device marks the sixth identifier for multiple frames of spectrum with the first fitting parameter greater than the first fitting threshold, and marks the seventh identifier for multiple frames of spectrum with the first fitting parameter less than the first fitting threshold. The electronic device marks the second fitting identifier Fz1 for multiple frames of spectrum based on the second fitting parameter; wherein, the electronic device marks the eighth identifier for multiple frames of spectrum with the second fitting parameter less than the second fitting threshold, and marks the ninth identifier for multiple frames of spectrum with the second fitting parameter greater than the second fitting threshold. For example, for multiple frames of spectrum with the first fitting parameter greater than the first fitting threshold and the second fitting parameter less than the second fitting threshold, the sixth identifier and the eighth identifier are marked. For multiple frames of spectrum with the first fitting parameter less than the first fitting threshold and the second fitting parameter greater than the second fitting threshold, the seventh identifier and the ninth identifier are marked. Optionally, the sixth identifier is different from the seventh identifier, the eighth identifier is different from the ninth identifier, the sixth identifier may be the same as or different from the eighth identifier, and the seventh identifier and the ninth identifier may be the same as or different from each other; for example, both the sixth identifier and the eighth identifier are 1, that is, Fe1 = 1, Fz1 = 1, and both the seventh identifier and the ninth identifier are -1, that is, Fe1 = -1, Fz1 = -1. Correspondingly, if the fifth identifier is marked for multiple frames of spectrum and the first fitting identifier Fe1 > 0 or the second fitting identifier Fz1 > 0, the electronic device determines that the audio signal is a speech signal. If the fifth identifier is marked for multiple frames of spectrum and the first fitting identifier Fe1 < 0, or the second fitting identifier Fz1 < 0, it is determined that the audio signal is a non-speech signal.

[0158] In this embodiment, by marking multiple frames of spectrum, it is convenient to determine whether the audio signal is a speech signal based on the first fitting identifier or the second fitting identifier of multiple spectrums when the fifth identifier is marked for multiple frames of spectrum. It is simple and direct, and can thus achieve fast speech detection of the audio signal.

[0159] In some embodiments, if the second quantity is greater than or equal to the third quantity threshold, and the first fitting parameter is greater than the first fitting threshold and the second fitting parameter is less than the second fitting threshold, the electronic device determines that the audio signal is a speech signal. In another implementation, if the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the first fitting parameter is less than the first fitting threshold and the second fitting parameter is greater than the second fitting threshold, the electronic device determines that the audio signal is a non-speech signal. In this embodiment, by comprehensively performing speech detection on the audio signal based on three criteria: frequency doubling information, energy parameter, and zero-crossing rate parameter, the reliability and accuracy of speech detection can be improved.

[0160] In an embodiment of the present application, since the short-time energy of the voice signal is large and the zero-crossing rate is small, while the short-time energy of the non-voice signal is small and the zero-crossing rate is large, when the frequency doubling information of the sampling points based on multiple frames of spectra cannot directly determine whether the audio signal is a voice signal, the short-time energy and zero-crossing rate of the multiple frames of spectra are combined to perform voice detection on the audio signal, thereby improving the flexibility and reliability of voice detection and making the voice detection of the audio signal more accurate.

[0161] An embodiment of the present application provides a voice detection method. Refer to Figure 6 , and this voice detection method includes:

[0162] Step 601: The electronic device acquires the target frame spectrum of the audio signal.

[0163] Step 602: The electronic device determines a first sampling point among multiple sampling points based on the target frame spectrum.

[0164] Step 603: The electronic device acquires the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0165] Step 604: The electronic device determines a first target frequency doubling sampling point based on the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point.

[0166] Step 605: The electronic device determines a target spectrum among multiple frames of spectra based on the first target frequency doubling sampling point.

[0167] Step 606: If the second quantity is less than the second quantity threshold, the electronic device acquires the energy parameter and zero-crossing rate parameter of the multiple frames of spectra.

[0168] Step 607: The electronic device determines a first difference and a second difference between each adjacent two frames of spectra except the first frame and the last frame among the multiple frames of spectra based on the energy corresponding to each of the multiple frames of spectra and the zero-crossing rate corresponding to each of the multiple frames of spectra.

[0169] Steps 601-607 are the same as steps 501-507, and will not be elaborated here.

[0170] Step 608: The electronic device determines the energy change probability corresponding to each of the multiple first differences and determines the zero-crossing rate change probability corresponding to each of the multiple second differences.

[0171] In one implementation, the electronic device determines the energy change probability corresponding to each first difference through the following softmax (a probability function) function based on the multiple first differences. The electronic device determines the zero-crossing rate change probability corresponding to each second difference through the following softmax function based on the multiple second differences.

[0172] Softmax function:

[0173] where X i represents the i-th difference. If the electronic device determines the energy change probability through this function, then X i represents the i-th first difference, and C represents the number of multiple first differences. If the electronic device determines the zero-crossing rate change probability through this function, then X i represents the second difference, and C represents the number of multiple second differences.

[0174] Optionally, the multiple energy change probabilities obtained by the electronic device processing multiple first differences de1, de2,..., de9 using the softmax function are respectively represented as Pe1, Pe2,..., Pe9. The multiple zero-crossing rate change probabilities obtained by the electronic device processing multiple second differences dz1, dz2,..., dz9 using the softmax function are respectively represented as Pz1, Pz2,..., Pz9.

[0175] Step 609: The electronic device performs voice detection on the audio signal based on the second quantity, the energy change probabilities corresponding to the multiple first differences, and the zero-crossing rate change probabilities corresponding to the multiple second differences.

[0176] In some embodiments, the electronic device performs voice detection on the audio signal based on the second quantity, the energy change probabilities corresponding to the multiple first differences, and the zero-crossing rate change probabilities corresponding to the multiple second differences, including the following steps: The electronic device respectively determines the first probability mean and the second probability mean. The first probability mean is the average of the energy change probabilities of the first target quantity among the multiple energy change probabilities, and the second probability mean is the average of the energy change probabilities of the second target quantity among the multiple energy change probabilities. The electronic device respectively determines the third probability mean and the fourth probability mean. The third probability mean is the average of the zero-crossing rate change probabilities of the first target quantity among the multiple zero-crossing rate change probabilities, and the fourth probability mean is the average of the zero-crossing rate change probabilities of the second target quantity among the multiple zero-crossing rate change probabilities. If the second quantity is greater than or equal to the third quantity threshold, and the first probability mean is less than the product of the second probability mean and the second preset ratio or the third probability mean is less than the product of the fourth probability mean and the second preset ratio, the electronic device determines that the audio signal is a voice signal. If the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the product of the second preset ratio and the first probability mean is greater than the second probability mean or the product of the second preset ratio and the third probability mean is greater than the fourth probability mean, the electronic device determines that the audio signal is a non-voice signal.

[0177] Optionally, the number of energy change probabilities of the previous target quantity is the same as that of the energy change probabilities of the subsequent target quantity. For example, if the number of multiple energy change probabilities is 9, the electronic device determines the average value of the first 5 energy change probabilities to obtain the first probability mean value, and the electronic device determines the average value of the last 5 energy change probabilities to obtain the second probability mean value. If the multiple energy change probabilities are Pe1, Pe2, ……, Pe9 respectively, then the first probability mean value Pem0 = (Pe1~Pe5) / 5, and the second probability mean value Pem1 = (Pe5~Pe9) / 5.

[0178] Similarly, the number of zero-crossing rate change probabilities of the previous target quantity is the same as that of the zero-crossing rate change probabilities of the subsequent target quantity. For example, if the number of multiple zero-crossing rate change probabilities is 9, the electronic device determines the average value of the first 5 zero-crossing rate change probabilities to obtain the third probability mean value, and the electronic device determines the average value of the last 5 zero-crossing rate change probabilities to obtain the fourth probability mean value. If the multiple zero-crossing rate change probabilities are Pz1, Pz2, ……, Pz9 respectively, then the third probability mean value Pzm0 = (Pz1~Pz5) / 5, and the fourth probability mean value Pzm1 = (Pz5~Pz9) / 5.

[0179] Among them, the second preset ratio is a value between 0 and 1, which can be set and changed as needed. For example, if the second preset ratio is 0.8, the relationship between the first probability mean value and the second probability mean value is satisfied as Pem0 < 0.8 * Pem1, or the relationship between the third probability mean value and the fourth probability mean value is satisfied as Pzm0 < 0.8 * Pzm1, and the electronic device determines that the audio signal is a voice signal. If the relationship between the first probability mean value and the second probability mean value is satisfied as 0.8 * Pem0 > Pem1, or the relationship between the third probability mean value and the fourth probability mean value is satisfied as 0.8 * Pzm0 > Pzm1, the electronic device determines that the audio signal is a non-voice signal.

[0180] It should be noted that although the energy change probability is the probability of the energy change corresponding to the first difference, only when the first probability mean is less than the second probability mean, it cannot accurately indicate that the energy of multiple frames of spectrum gradually increases. Only when the second probability mean is much greater than the first probability mean can it accurately indicate that the energy of multiple frames of spectrum gradually increases. Similarly, although the zero-crossing rate change probability is the probability of the zero-crossing rate change corresponding to the second difference, only when the third probability mean is less than the fourth probability mean, it cannot accurately indicate the decrease of the zero-crossing rate of multiple frames of spectrum. Only when the fourth probability mean is much greater than the third probability mean can it accurately indicate that the zero-crossing rate of multiple frames of spectrum gradually decreases. Thus, in the embodiment of the present application, if the first probability mean is less than the product of the second probability mean and the second preset ratio or the third probability mean is less than the product of the fourth probability mean and the second preset ratio, it is determined that the audio signal is a speech signal; if the product of the second preset ratio and the first probability mean is greater than the second probability mean or the product of the second preset ratio and the third probability mean is greater than the fourth probability mean, it is determined that the audio signal is a non-speech signal, thereby improving the accuracy of detecting the speech signal.

[0181] In the embodiment of the present application, since the energy change probability and the zero-crossing rate change probability can respectively represent the probability of the energy change and the probability of the zero-crossing rate change, detecting the speech of the audio signal based on the energy change probability corresponding to multiple first differences and the zero-crossing rate change probability corresponding to multiple second differences can improve the reliability of the detection.

[0182] In some embodiments, the electronic device determines the probability identification Fe2 of the energy change of multiple frames of spectrum markers based on the first probability mean and the second probability mean; wherein, the electronic device marks multiple frames of spectrum with a second preset ratio of the first probability mean less than the second probability mean as the tenth identification, and marks multiple frames of spectrum with the product of the second preset ratio and the first probability mean greater than the second probability mean as the eleventh identification. The electronic device determines the probability identification Fz2 of the zero-crossing rate change of multiple frames of spectrum markers based on the third probability mean and the fourth probability mean; wherein, the electronic device marks multiple frames of spectrum with the third probability mean less than the product of the fourth probability mean and the second preset ratio as the twelfth identification, and marks multiple frames of spectrum with the product of the second preset ratio and the third probability mean greater than the fourth probability mean as the thirteenth identification. Optionally, the tenth identification and the eleventh identification are different, the twelfth identification and the thirteenth identification are different, the tenth identification and the twelfth identification may be the same or different, and the eleventh identification and the thirteenth identification may be the same or different. For example, if both the tenth identification and the twelfth identification are 1, that is, Fe2 = 1 and Fz2 = 1, and both the eleventh identification and the thirteenth identification are -1, that is, Fe2 = -1 and Fz2 = -1. Correspondingly, if multiple frames of spectrum markers are the fifth identification and the probability identification Fe2 of energy change > 0 or the probability identification Fz2 of zero-crossing rate change > 0, the electronic device determines that the audio signal is a voice signal. If multiple frames of spectrum markers are the fifth identification and the probability identification Fe2 of energy change < 0, or the probability identification Fz2 of zero-crossing rate change < 0, it is determined that the audio signal is a non-voice signal.

[0183] In this embodiment, by marking multiple frames of spectrum, it is convenient to determine whether the audio signal is a voice signal based on the probability identification of the energy change and the probability identification of the zero-crossing rate of multiple spectra when the multiple frames of spectrum are marked as the fifth identification. It is simple and direct, and can thus achieve fast voice detection of the audio signal.

[0184] In some embodiments, if the second quantity is greater than or equal to the third quantity threshold, and the first probability mean is less than the product of the second probability mean and the second preset ratio, and the third probability mean is less than the product of the fourth probability mean and the second preset ratio, the electronic device determines that the audio signal is a voice signal. If the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the product of the second preset ratio and the first probability mean is greater than the second probability mean, and the product of the second preset ratio and the third probability mean is greater than the fourth probability mean, the electronic device determines that the audio signal is a non-voice signal. In this embodiment, by comprehensively using three criteria, namely the octave information, the energy parameter, and the zero-crossing rate parameter, to perform voice detection on the audio signal, the reliability and accuracy of voice detection can be improved.

[0185] See Figure 7 , Figure 7 which is a schematic diagram of an audio signal and spectrum provided for the embodiments of this application itself. Figure 7The upper part represents the audio signal, and the lower part represents the spectrogram; among them, the framed part in the spectrum is the speech signal in this part of the audio signal. It can be seen from the figure that the amplitude values of the distances between the sampling points and their multiple-frequency sampling points in this section of the spectrum are similar, that is, the spectrum of the speech signal has multi-order characteristics. Furthermore, speech detection can be performed on the audio signal based on the multiple-frequency information of the sampling points, and the detection result is reliable. It should be noted that through the method provided in the embodiments of the present application, speech detection of the audio signal is achieved through multiple criteria, ensuring the accuracy of the detection result. This not only avoids the situation where the detection result of an audio signal with a low signal-to-noise ratio is inaccurate only by using short-time energy and zero-crossing rate; but also avoids the process of training a speech recognition model with a large amount of sample audio data, reducing the computational amount. That is, the method provided in the embodiments of the present application fully considers both the computational amount of the data and the accuracy of the detection result, and can thus ensure the accuracy of the detection result under a certain computational amount of the data.

[0186] In the embodiments of the present application, since the short-time energy of the speech signal is large and the zero-crossing rate is small, while the short-time energy of the non-speech signal is small and the zero-crossing rate is large. In the case where it is not possible to directly determine whether the audio signal is a speech signal based on the multiple-frequency information of the sampling points of multiple frames of spectra, the short-time energy and zero-crossing rate of multiple frames of spectra are combined to perform speech detection on the audio signal, thereby improving the flexibility of speech detection and making the speech detection of the audio signal more accurate.

[0187] The embodiments of the present application also provide a speech detection device. Refer to Figure 8 , the speech detection device includes:

[0188] A first acquisition module 801, configured to acquire the target frame spectrum of the audio signal, where the target frame spectrum includes the amplitude values of multiple sampling points of the target frame;

[0189] A first determination module 802, configured to determine a first sampling point among the multiple sampling points based on the target frame spectrum, where the amplitude value of the first sampling point is the maximum amplitude value in the target frame spectrum;

[0190] A second acquisition module 803, configured to acquire the multiple-frequency information of the first sampling point, where the multiple-frequency information of the first sampling point includes the amplitude values of at least one multiple-frequency sampling point of the first sampling point;

[0191] A detection module 804, configured to perform speech detection on the audio signal based on the multiple-frequency information of the first sampling point.

[0192] In some embodiments, the detection module 804 is configured to obtain the frequency doubling information of at least one second sampling point. The frequency doubling information of the second sampling point includes the amplitude values of at least one frequency doubling sampling point of the second sampling point. The amplitude value of the second sampling point is the maximum amplitude value in the spectrum of other frames, and the spectrum of other frames is the spectrum adjacent to the target frame spectrum in the audio signal. Based on the frequency doubling information of the first sampling point and the frequency doubling information of at least one second sampling point, determine the first target frequency doubling sampling point. The amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in the multi-frame spectrum composed of the spectrum of other frames and the target frame spectrum. Based on the first target frequency doubling sampling point, determine the target spectrum in the multi-frame spectrum. The target spectrum includes the target sampling points among the first sampling point and at least one second sampling point, and the first quantity of the first target frequency doubling sampling points corresponding to the target sampling points is greater than or equal to the first quantity threshold. If the second quantity is greater than or equal to the second quantity threshold, determine that the audio signal is a voice signal, where the second quantity is the number of target spectra included in the multi-frame spectrum.

[0193] In some embodiments, the detection module 804 is configured to, if the second quantity is less than the second quantity threshold, obtain the energy parameter and the zero-crossing rate parameter of the multi-frame spectrum. Based on the second quantity, the energy parameter, and the zero-crossing rate parameter, perform voice detection on the audio signal.

[0194] In some embodiments, the energy parameter includes the energy corresponding to each multi-frame spectrum, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to each multi-frame spectrum. The detection module 804 is configured to determine the first difference and the second difference between each adjacent two-frame spectra except the first frame and the last frame in the multi-frame spectrum based on the energy corresponding to each multi-frame spectrum and the zero-crossing rate corresponding to each multi-frame spectrum. The first difference is the difference in energy between two-frame spectra, and the second difference is the difference in zero-crossing rate between two-frame spectra. Based on the second quantity and the multiple first differences and multiple second differences corresponding to the multi-frame spectrum, perform voice detection on the audio signal.

[0195] In some embodiments, the detection module 804 is configured to perform linear fitting on the multiple first differences to obtain the first fitting parameter. Perform linear fitting on the multiple second differences to obtain the second fitting parameter. Based on the second quantity, the first fitting parameter, and the second fitting parameter, perform voice detection on the audio signal.

[0196] In some embodiments, the number of non-target spectra in the multi-frame spectrum is the third quantity. The detection module 804 is configured to, if the second quantity is greater than or equal to the third quantity threshold, and the first fitting parameter is greater than the first fitting threshold or the second fitting parameter is less than the second fitting threshold, determine that the audio signal is a voice signal.

[0197] If the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the first fitting parameter is less than the first fitting threshold or the second fitting parameter is greater than the second fitting threshold, then determine that the audio signal is a non-speech signal.

[0198] In some embodiments, the detection module 804 is configured to determine the energy change probabilities corresponding to the multiple first differences and determine the zero-crossing rate change probabilities corresponding to the multiple second differences; and perform speech detection on the audio signal based on the second quantity, the energy change probabilities corresponding to the multiple first differences, and the zero-crossing rate change probabilities corresponding to the multiple second differences.

[0199] In some embodiments, the number of non-target spectra in the multi-frame spectrum is the third quantity. The detection module 804 is configured to respectively determine a first probability mean value and a second probability mean value. The first probability mean value is the average of the energy change probabilities of the first target quantity among the multiple energy change probabilities, and the second probability mean value is the average of the energy change probabilities of the last target quantity among the multiple energy change probabilities; respectively determine a third probability mean value and a fourth probability mean value. The third probability mean value is the average of the zero-crossing rate change probabilities of the first target quantity among the multiple zero-crossing rate change probabilities, and the fourth probability mean value is the average of the zero-crossing rate change probabilities of the last target quantity among the multiple zero-crossing rate change probabilities; if the second quantity is greater than or equal to the third quantity threshold, and the first probability mean value is less than the product of the second probability mean value and the second preset ratio or the third probability mean value is less than the product of the fourth probability mean value and the second preset ratio, then determine that the audio signal is a speech signal; if the third quantity is less than the second quantity threshold and greater than or equal to the third quantity threshold, and the product of the second preset ratio and the first probability mean value is greater than the second probability mean value or the product of the second preset ratio and the third probability mean value is greater than the fourth probability mean value, then determine that the audio signal is a non-speech signal.

[0200] In some embodiments, the speech detection device further includes a second determination module, configured to determine that the audio signal is a non-speech signal if the third quantity of non-target spectra in the multi-frame spectrum is greater than or equal to the second quantity threshold.

[0201] In some embodiments, the detection module 804 is configured to determine a second target multiple-frequency sampling point based on the multiple-frequency information of the first sampling point, and the amplitude value of the second target multiple-frequency sampling point is the maximum amplitude value in the target frame spectrum; if the fourth quantity of the second target multiple-frequency sampling point is greater than or equal to the first quantity threshold, then determine that the audio signal is a speech signal.

[0202] In some embodiments, the voice detection device further includes a third determination module, configured to determine two candidate sampling points, where the two candidate sampling points are respectively the sampling points before and after the initial frequency-doubled sampling point, and the initial frequency-doubled sampling point is the sampling point of the first sampling point at a target multiple; based on the two candidate sampling points and the initial frequency-doubled sampling point, determine the second target frequency-doubled sampling point of the first sampling point at the target multiple.

[0203] In some embodiments, among multiple sampling points, the sampling points whose amplitude values are the maximum amplitude values in the target frame spectrum and whose amplitude values are greater than the amplitude threshold are assigned a first value. The third determination module is configured to, if the two candidate sampling points and the initial frequency-doubled sampling point include sampling points assigned the first value, determine the sampling points assigned the first value as the second target frequency-doubled sampling point of the first sampling point at the target multiple.

[0204] In some embodiments, the voice detection device further includes:

[0205] A fourth determination module, configured to determine an initial sampling point, where the amplitude value of the initial sampling point is the maximum amplitude value in the target frame spectrum and the amplitude value is greater than the amplitude threshold;

[0206] A third acquisition module, configured to respectively acquire the nearest sampling point before the initial sampling point and whose amplitude value is the maximum amplitude value in the target frame spectrum, and the nearest sampling point after the initial sampling point and whose amplitude value is the maximum amplitude value in the target frame spectrum;

[0207] A fifth determination module, configured to, if the amplitude values of the nearest sampling point before the initial sampling point and whose amplitude value is the maximum amplitude value in the target frame spectrum and the nearest sampling point after the initial sampling point and whose amplitude value is the maximum amplitude value in the target frame spectrum are both greater than the amplitude threshold, determine the initial sampling point as the first sampling point.

[0208] In some embodiments, the voice detection device further includes:

[0209] A sixth determination module, configured to determine a target maximum value from multiple spectrum maximum values, where the value of the target maximum value is the largest;

[0210] A seventh determination module, configured to determine the product of the target maximum value and a first preset ratio to obtain the amplitude threshold.

[0211] In some embodiments, the electronic device is provided as a terminal. Figure 9The block diagram of the terminal 900 provided by an exemplary embodiment of the present application is shown. The terminal 900 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 900 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc. Generally, the terminal 900 includes: a processor 901 and a memory 902.

[0212] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0213] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 is used to store at least one program code, and the at least one program code is used to be executed by the processor 901 to implement the voice detection method provided by the method embodiments of the present application.

[0214] In some embodiments, the terminal 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.

[0215] The peripheral device interface 903 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0216] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0217] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905. The touch signals can be input as control signals to the processor 901 for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 905, which is set on the front panel of the terminal 900; in other embodiments, there can be at least two display screens 905, which are respectively set on different surfaces of the terminal 900 or are in a foldable design; in other embodiments, the display screen 905 can be a flexible display screen, which is set on the curved surface or the folding surface of the terminal 900. Even, the display screen 905 can also be set as an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 905 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0218] The camera module 906 is used to collect images or videos. Optionally, the camera module 906 includes a front camera and a rear camera. Generally, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 906 can also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0219] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 901 for processing, or input to the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 907 may further include a headphone jack.

[0220] The positioning component 908 is used to locate the current geographical location of the terminal 900 to achieve navigation or LBS (Location Based Service). The positioning component 908 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.

[0221] The power supply 909 is used to supply power to each component in the terminal 900. The power supply 909 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 909 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0222] In some embodiments, the terminal 900 further includes one or more sensors 910. The one or more sensors 910 include but are not limited to: an acceleration sensor 911, a gyroscope sensor 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.

[0223] The acceleration sensor 911 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 900. For example, the acceleration sensor 911 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 901 can control the display screen 905 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 911. The acceleration sensor 911 can also be used for game or collection of the user's motion data.

[0224] The gyroscope sensor 912 can detect the body orientation and rotation angle of the terminal 900. The gyroscope sensor 912 can cooperate with the acceleration sensor 911 to collect the 3D actions of the user on the terminal 900. Based on the data collected by the gyroscope sensor 912, the processor 901 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0225] The pressure sensor 913 can be disposed on the side frame of the terminal 900 and / or the lower layer of the display screen 905. When the pressure sensor 913 is disposed on the side frame of the terminal 900, it can detect the holding signal of the user on the terminal 900, and the processor 901 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 913. When the pressure sensor 913 is disposed on the lower layer of the display screen 905, the processor 901 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 905. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0226] The fingerprint sensor 914 is used to collect the fingerprints of the user. The processor 901 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 914, or the fingerprint sensor 914 can identify the user's identity according to the collected fingerprints. When the identified user identity is a trusted identity, the processor 901 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 914 can be disposed on the front, back, or side of the terminal 900. When there are physical buttons or manufacturer Logos on the terminal 900, the fingerprint sensor 914 can be integrated with the physical buttons or manufacturer Logos.

[0227] The optical sensor 915 is used to collect the ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera module 906 according to the ambient light intensity collected by the optical sensor 915.

[0228] The proximity sensor 916, also known as the distance sensor, is usually disposed on the front panel of the terminal 900. The proximity sensor 916 is used to collect the distance between the user and the front of the terminal 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually decreasing, the processor 901 controls the display screen 905 to switch from the lit state to the off state; when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually increasing, the processor 901 controls the display screen 905 to switch from the off state to the lit state. Those skilled in the art can understand that Figure 9 the structure shown in

[0229] does not limit the terminal 900, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0230] The embodiments of the present application also provide a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the voice detection method in any of the above implementation manners.

[0231] The embodiments of the present application also provide a computer program product, the computer program product includes computer program code, the computer program code is stored in a computer-readable storage medium, and a processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device executes the voice detection method in any of the above implementation manners. In some embodiments, the computer program product involved in the embodiments of the present application may be deployed to be executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network. The multiple electronic devices distributed at multiple locations and interconnected by a communication network may form a blockchain system.

[0232] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A voice detection method, characterized in that, The method includes: Obtaining a target frame spectrum of an audio signal, where the target frame spectrum includes amplitude values of multiple sampling points of a target frame; Based on the target frame spectrum, determining an initial sampling point, where the amplitude value of the initial sampling point is the maximum amplitude value in the target frame spectrum and the amplitude value is greater than an amplitude threshold; Respectively obtaining two auxiliary sampling points corresponding to the initial sampling point, where the two auxiliary sampling points are respectively the nearest sampling point before the initial sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum, and the nearest sampling point after the initial sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum; If the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold, determining the initial sampling point as the first sampling point among the multiple sampling points; Obtaining the frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes amplitude values of at least one frequency doubling sampling point of the first sampling point; Based on the frequency doubling information of the first sampling point, performing voice detection on the audio signal.

2. The method according to claim 1, characterized in that, The performing voice detection on the audio signal based on the frequency doubling information of the first sampling point includes: Obtaining the frequency doubling information of at least one second sampling point, where the frequency doubling information of the second sampling point includes amplitude values of at least one frequency doubling sampling point of the second sampling point, the amplitude value of the second sampling point is the maximum amplitude value in other frame spectra, and the other frame spectra are spectra adjacent to the target frame spectrum in the audio signal; Based on the frequency doubling information of the first sampling point and the frequency doubling information of the at least one second sampling point, determining a first target frequency doubling sampling point, where the amplitude value of the first target frequency doubling sampling point is the maximum amplitude value in a multi-frame spectrum composed of the other frame spectra and the target frame spectrum; Based on the first target frequency doubling sampling point, determining a target spectrum in the multi-frame spectrum, where the target spectrum includes target sampling points among the first sampling point and the at least one second sampling point, and the first quantity of the first target frequency doubling sampling points corresponding to the target sampling points is greater than or equal to a first quantity threshold; If a second quantity is greater than or equal to a second quantity threshold, determining that the audio signal is a voice signal, where the second quantity is the quantity of target spectra included in the multi-frame spectrum.

3. The method according to claim 2, characterized in that, The method further includes: If the second quantity is less than the second quantity threshold, obtaining the energy parameter and the zero-crossing rate parameter of the multi-frame spectrum; Based on the second quantity, the energy parameter, and the zero-crossing rate parameter, performing voice detection on the audio signal.

4. The method according to claim 3, characterized in that, The energy parameter includes the energy corresponding to each of the multi-frame spectra, and the zero-crossing rate parameter includes the zero-crossing rate corresponding to each of the multi-frame spectra; The performing voice detection on the audio signal based on the second quantity, the energy parameter, and the zero-crossing rate parameter includes: Based on the energy corresponding to each of the multiple frames of spectra and the zero-crossing rate corresponding to each of the multiple frames of spectra, determine a first difference and a second difference between each adjacent two frames of spectra among the multiple frames of spectra except the first frame and the last frame, where the first difference is the difference in energy between the two frames of spectra, and the second difference is the difference in zero-crossing rate between the two frames of spectra; Based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra, perform voice detection on the audio signal.

5. The method according to claim 4, characterized in that, The performing voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra includes: Perform linear fitting on the multiple first differences to obtain a first fitting parameter; Perform linear fitting on the multiple second differences to obtain a second fitting parameter; Based on the second quantity, the first fitting parameter, and the second fitting parameter, perform voice detection on the audio signal.

6. The method according to claim 4, characterized in that, The performing voice detection on the audio signal based on the second quantity and the multiple first differences and multiple second differences corresponding to the multiple frames of spectra includes: Determine the energy change probability corresponding to each of the multiple first differences and determine the zero-crossing rate change probability corresponding to each of the multiple second differences; Based on the second quantity, the energy change probability corresponding to each of the multiple first differences, and the zero-crossing rate change probability corresponding to each of the multiple second differences, perform voice detection on the audio signal.

7. The method according to claim 1, characterized in that, The performing voice detection on the audio signal based on the multiple-frequency information of the first sampling point includes: Based on the multiple-frequency information of the first sampling point, determine a second target multiple-frequency sampling point, where the amplitude value of the second target multiple-frequency sampling point is the maximum amplitude value in the target frame spectrum; If the fourth quantity of the second target multiple-frequency sampling points is greater than or equal to the first quantity threshold, determine that the audio signal is a voice signal.

8. The method according to claim 7, characterized in that, The process of determining the second target multiple-frequency sampling point includes: Determine two candidate sampling points, where the two candidate sampling points are the sampling points before and after the initial multiple-frequency sampling point respectively, and the initial multiple-frequency sampling point is the sampling point of the first sampling point at the target multiple; Based on the two candidate sampling points and the initial multiple-frequency sampling point, determine the second target multiple-frequency sampling point of the first sampling point at the target multiple.

9. The method according to claim 8, wherein, Among the multiple sampling points, the sampling point with an amplitude value being the maximum amplitude value in the target frame spectrum and the amplitude value being greater than the amplitude threshold is assigned a first value. The determining the second target multiple-frequency sampling point of the first sampling point at the target multiple based on the two candidate sampling points and the initial multiple-frequency sampling point includes: If the two candidate sampling points and the initial multiple-frequency sampling point include the sampling point assigned the first value, determine the sampling point assigned the first value as the second target multiple-frequency sampling point of the first sampling point at the target multiple.

10. The method according to any one of claims 1 or 9, wherein, The target frame spectrum includes multiple spectrum maximum values. The process of determining the amplitude threshold includes: Determine a target maximum value from the multiple spectrum maximum values, where the value of the target maximum value is the largest; Determine the product of the target maximum value and the first preset ratio to obtain the amplitude threshold value.

11. A voice detection device, wherein, The device includes: A first acquisition module, configured to acquire a target frame spectrum of an audio signal, where the target frame spectrum includes amplitude values of multiple sampling points of a target frame; A first determination module, configured to determine an initial sampling point based on the target frame spectrum, where the amplitude value of the initial sampling point is the maximum amplitude value in the target frame spectrum and the amplitude value is greater than the amplitude threshold value; respectively acquire two auxiliary sampling points corresponding to the initial sampling point, where the two auxiliary sampling points are respectively the nearest sampling point before the initial sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum and the nearest sampling point after the initial sampling point and with an amplitude value being the maximum amplitude value in the target frame spectrum; if the amplitude values of the two auxiliary sampling points are both greater than the amplitude threshold value, determine the initial sampling point as the first sampling point among the multiple sampling points; A second acquisition module, configured to acquire the frequency doubling information of the first sampling point, where the frequency doubling information of the first sampling point includes the amplitude values of at least one frequency doubling sampling point of the first sampling point; A detection module, configured to perform voice detection on the audio signal based on the frequency doubling information of the first sampling point.

12. An electronic device, wherein, The electronic device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the voice detection method according to any one of claims 1 to 10.

13. A computer-readable storage medium, wherein, At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to implement the voice detection method according to any one of claims 1 to 10.

14. A computer program product, wherein, The computer program product includes computer program code, the computer program code is stored in a computer-readable storage medium, a processor of an electronic device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the electronic device executes the voice detection method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • System and method for automatic speach to text conversion

    CN102227767A

  • Audio processing method and device, and storage medium

    CN109065068A

  • Recognition method of Chinese sound using computer

    CN85100180A