Voice signal detection method and computer-readable storage medium

By calculating the autocorrelation value and zero crossing rate of the sound signal, combined with the detection results of the previous frame of the sound signal, the problem of inaccurate speech detection in a low signal-to-noise ratio environment is solved, and accurate speech signal detection in a low signal-to-noise ratio environment is achieved.

CN115547364BActive Publication Date: 2025-08-19GEER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211205753.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-08-19
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

The existing speech detection methods are inaccurate in low signal-to-noise ratio environments, resulting in low applicability of speech detection technology.

Method used

By calculating the autocorrelation value and zero crossing rate of the sound signal, and combining the detection results of the sound signal in the previous frame, it is determined whether the sound signal is a voice signal.

Benefits of technology

It realizes accurate detection of voice signals and non-voice signals in a low signal-to-noise ratio environment, and improves the applicability of voice signal detection methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115547364B_ABST
    Figure CN115547364B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice signal detection method and a computer-readable storage medium. The voice signal detection method includes the following steps: sampling a frame of sound signal at a preset sampling rate, and using the sampled frame of sound signal as a sound signal to be detected; calculating the autocorrelation value to be detected and the zero-crossing rate to be detected of the sound signal to be detected, and obtaining a detection result obtained by performing voice signal detection on the previous frame of sound signal to be detected; and determining a detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of sound signal, wherein the detection result of the sound signal is a result indicating whether the sound signal is a speech signal. The present invention achieves accurate detection of non-speech signals and speech signals in sound signals with a low signal-to-noise ratio, thereby improving the applicability of the voice signal detection method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a speech signal detection method, apparatus, device, and computer-readable storage medium. Background Art

[0002] With the development of technology, intelligent voice interaction has gradually become part of people's lives. In intelligent voice interaction, voice detection is often required to distinguish between voice and non-voice signals to make intelligent voice interaction more accurate. Current voice detection methods are mainly based on short-time energy. During the detection process, the short-time energy of the sound signal is compared with a preset threshold to detect the voice signal and non-voice signal in the sound signal. Short-time energy-based detection methods can detect voice signals relatively accurately in high signal-to-noise ratio environments. However, in low signal-to-noise ratio environments, the noise in the sound signal significantly interferes with the voice signal, making this method unable to accurately detect the voice signal and non-voice signal in the sound signal, which reduces the applicability of voice detection technology.

[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of the present invention is to provide a speech signal detection method and a computer-readable storage medium, aiming to solve the technical problem that speech detection results are inaccurate in low signal-to-noise ratio environments, resulting in low applicability of speech detection technology.

[0005] To achieve the above object, the present invention provides a method for detecting a speech signal, wherein the speech signal detection comprises the following steps:

[0006] Sampling a frame of sound signal according to a preset sampling rate, and using the sampled frame of sound signal as the sound signal to be detected;

[0007] Calculating the autocorrelation value and zero-crossing rate of the sound signal to be detected, and obtaining a detection result obtained by performing speech signal detection on a previous frame of the sound signal to be detected;

[0008] The detection result of the sound signal to be detected is determined based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal, wherein the detection result of the sound signal is a result indicating whether the sound signal is a speech signal.

[0009] Optionally, before the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal, the method further includes:

[0010] Obtaining a preset range of autocorrelation values to be detected, wherein a minimum value in the preset range is a signal autocorrelation value of a preset non-speech signal, and a maximum value in the preset range is a signal autocorrelation value of a preset speech signal;

[0011] Obtaining a preset zero-crossing rate, wherein the preset zero-crossing rate is determined based on a signal zero-crossing rate of the preset non-speech signal and a signal zero-crossing rate of the preset speech signal;

[0012] The step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal includes:

[0013] The detection result of the sound signal to be detected is determined based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame of sound signal.

[0014] Optionally, when the detection result of the previous frame of sound signal is that the previous frame of sound signal is a non-speech signal, the step of determining the detection result of the sound signal to be detected based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame of sound signal includes:

[0015] When it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, after cumulatively adding one to the first number and setting the second number to zero, detecting whether the first number is greater than a first preset threshold, and detecting whether the sum of the first number and the third number is greater than a second preset threshold, wherein the first number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is greater than the maximum value in the preset range and whose detected zero-crossing rate is greater than the preset zero-crossing rate; the second number is the number of detected sound signals for which the first number is greater than the third preset threshold when the detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate; and the third number is the number of detected sound signals whose detected autocorrelation value is within the preset range;

[0016] When it is determined that the first number is less than or equal to the first preset threshold, or the sum of the first number and the third number is less than or equal to the second preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0017] When it is determined that the first number is greater than the first preset threshold and the sum of the first number and the third number is greater than the second preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the first number and the third number are set to zero.

[0018] Optionally, before the step of adding one to the first quantity and setting the second quantity to zero, the method further includes:

[0019] Detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate;

[0020] After the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes:

[0021] When the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal;

[0022] When the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0023] When it is determined that the autocorrelation value to be detected is within the preset range, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal, the third number is cumulatively increased by one and the second number is set to zero.

[0024] Optionally, after the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes:

[0025] When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, detecting whether the first number is greater than the third preset threshold;

[0026] When it is determined that the first number is less than or equal to the third preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal;

[0027] When it is determined that the first number is greater than the third preset threshold, adding one to the second number, and then detecting whether the second number is greater than a fourth preset threshold;

[0028] When it is determined that the second number is greater than the fourth preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and setting the first number, the second number, and the third number to zero;

[0029] When it is determined that the second number is less than or equal to the fourth preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal.

[0030] Optionally, when the detection result of the previous frame of sound signal is that the previous frame of sound signal is a speech signal, the step of determining the detection result of the sound signal to be detected based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame of sound signal includes:

[0031] When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, after cumulatively adding one to the fourth number and setting the fifth number to zero, detecting whether the fourth number is greater than a fifth preset threshold, wherein the fourth number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate, and the fifth number is the number of detected sound signals for which the fourth number is greater than a sixth preset threshold when the detected autocorrelation value is greater than the maximum value in the preset range;

[0032] When it is determined that the fourth number is less than or equal to the fifth preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0033] When it is determined that the fourth number is greater than the fifth preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and the fourth number is set to zero.

[0034] Optionally, before the step of adding one to the fourth quantity and setting the fifth quantity to zero, the method further includes:

[0035] Detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate;

[0036] After the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes:

[0037] When it is determined that the autocorrelation value to be detected is smaller than the minimum value in the preset range and the zero-crossing rate to be detected is greater than a preset zero-crossing rate, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal.

[0038] Optionally, after the step of detecting whether the autocorrelation value to be detected is within the preset range, the method further includes:

[0039] When it is determined that the autocorrelation value to be detected is greater than the minimum value in the preset range, detecting whether the fourth number is greater than the sixth preset threshold;

[0040] When it is determined that the fourth number is greater than the sixth preset threshold, adding one to the fifth number, and then detecting whether the fifth number is greater than a seventh preset threshold;

[0041] When it is determined that the fifth number is less than or equal to the seventh preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0042] When it is determined that the fifth number is greater than the seventh preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the fourth number and the fifth number are set to zero.

[0043] Optionally, the voice signal detection method is applied to a head-mounted device, and after the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of the sound signal, the method further includes:

[0044] When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a voice signal, turning on a transparency mode of the head mounted device;

[0045] When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-voice signal, an active noise reduction mode of the head mounted device is turned on.

[0046] In addition, to achieve the above-mentioned purpose, the present invention also proposes a computer-readable storage medium, on which a speech signal detection program is stored. When the speech signal detection program is executed by a processor, the steps of the speech signal detection method described above are implemented.

[0047] In the present invention, a frame of sound signal is sampled at a preset sampling rate, the sampled frame of sound signal is used as the sound signal to be detected, the autocorrelation value to be detected and the zero-crossing rate to be detected of the sound signal to be detected are calculated, and the detection result of the speech signal detection of the previous frame of sound signal to be detected is obtained. The detection result of the sound signal to be detected is determined based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal, wherein the detection result of the sound signal is a result that characterizes whether the sound signal is a speech signal. Compared with the currently commonly used short-time energy detection method, the present invention performs detection based on the autocorrelation and zero-crossing rate of the sound signal. By utilizing the different zero-crossing rates of the sound signal, it can effectively distinguish between unvoiced and voiced sounds to determine whether there is a speech signal in the sound signal. By utilizing the different autocorrelation functions of the noise signal and the speech signal, it can effectively distinguish between the noise signal and the speech signal. Therefore, the present invention achieves accurate detection of non-speech signals and speech signals in sound signals with low signal-to-noise ratio, improving the applicability of the speech signal detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 1 is a flow chart of a first embodiment of a method for detecting a speech signal according to the present invention;

[0049] Figure 2 、 Figure 3 and Figure 4 Schematic diagram of a flow chart of an embodiment of a speech signal detection method of the present invention;

[0050] Figure 5 It is a structural diagram of a voice signal detection device in a hardware operating environment involved in an embodiment of the present invention.

[0051] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0052] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0053] The embodiment of the present invention provides a method for detecting a speech signal. Figure 1 , Figure 1 This is a flow chart of the first embodiment of a voice signal detection method of the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than here. The voice signal detection method may be performed by a personal computer, a smart phone, a head-mounted device, or other device, which is not limited in this embodiment. For the convenience of description, the description of the execution subject is omitted below. In this embodiment, the voice signal detection method includes:

[0054] Step S10, sampling a frame of sound signal according to a preset sampling rate, and using the sampled frame of sound signal as the sound signal to be detected;

[0055] Current speech signal detection methods are primarily based on short-term energy. Generally, background noise has the lowest short-term energy, voiced sounds have the highest short-term energy, and unvoiced sounds fall somewhere in between. This characteristic allows for relatively accurate speech signal detection in high signal-to-noise ratio environments. However, in low signal-to-noise ratio environments, the noise present in the sound signal significantly interferes with the speech signal, making this method incapable of accurately detecting the speech signal, reducing the applicability of speech detection technology.

[0056] In this embodiment, a frame of sound signal is sampled according to a preset sampling rate, and the sampled sound signal is used as a sound signal to be detected for voice signal detection.

[0057] Furthermore, in a specific embodiment, after sampling a frame of sound signal, the sampled frame of sound signal can be subjected to signal processing, and the processed frame of sound signal can be used as the sound signal to be detected. For example, in one embodiment, the DC component of the sampled frame of sound signal can be removed to reduce interference in speech signal detection and improve the accuracy of speech signal detection. Specific settings can be made based on actual needs and are not limited here.

[0058] Step S20, calculating the autocorrelation value and zero-crossing rate of the sound signal to be detected, and obtaining a detection result obtained by performing speech signal detection on a previous frame of the sound signal to be detected;

[0059] In this embodiment, after obtaining the sound signal to be detected, the autocorrelation value (hereinafter referred to as the autocorrelation value to be detected for distinction) and the zero-crossing rate (hereinafter referred to as the zero-crossing rate to be detected for distinction) of the sound signal to be detected are calculated. The specific calculation method is not repeated in this embodiment.

[0060] In a specific embodiment, the value of the autocorrelation value to be detected is not limited and can be set according to actual needs. It is not limited here. For example, in one embodiment, the autocorrelation value to be detected can be the maximum value of the autocorrelation function of the sound signal to be detected. In this embodiment, the autocorrelation value to be detected can be calculated according to the maximum value calculation formula of the autocorrelation function: Calculate, where R t is the autocorrelation function of the t-th sampling point of the sound signal to be detected, L is the length of the sound signal to be detected, and τ is the delay amount; in another embodiment, the autocorrelation value to be detected can also be the average of the autocorrelation values of each sampling point of the sound signal to be detected.

[0061] In this embodiment, a detection result obtained by performing voice signal detection on the previous frame of the sound signal to be detected is obtained. The detection result of the previous frame of the sound signal may be a voice signal or a non-voice signal. In a specific embodiment, the detection result can be stored after each voice signal detection is completed. In this embodiment, the detection result of the previous frame of the sound signal can be obtained from the stored detection results. The details are not detailed here.

[0062] Step S30 , determining a detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of sound signal, wherein the detection result of the sound signal is a result indicating whether the sound signal is a speech signal.

[0063] In this embodiment, after obtaining the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of sound signal, the detection result of the sound signal to be detected (hereinafter referred to as the detection result of the sound signal to be detected for distinction) can be determined based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of sound signal.

[0064] In a specific embodiment, the detection conditions in the speech signal detection process can be determined based on the detection results of the previous frame of sound signal. For example, in one embodiment, different thresholds can be determined based on different detection results of the previous frame of sound signal, so that the autocorrelation value and zero-crossing rate to be detected can be compared with the threshold to determine the detection result of the sound signal to be detected. The specific settings can be made according to actual needs and are not limited here.

[0065] In a specific embodiment, the autocorrelation value to be detected and the zero-crossing rate to be detected can be compared with the detection conditions in the voice signal detection to determine the detection result of the voice signal to be detected. Specifically, in one embodiment, the value obtained by combining the autocorrelation value to be detected and the zero-crossing rate to be detected can be used as the detection content of the voice signal detection to determine the detection result of the voice signal to be detected. For example, the product or sum of the autocorrelation value to be detected and the zero-crossing rate to be detected can be used as the detection content; in another embodiment, the autocorrelation value to be detected and the zero-crossing rate to be detected can be used as the detection content respectively, and the detection result of the voice signal to be detected can be determined based on the respective detection results of the two. The specific configuration can be made according to actual needs and is not limited here.

[0066] It should be noted that, by sampling a frame of sound signal at a preset sampling rate, the sampled frame of sound signal is used as the sound signal to be detected, the autocorrelation value to be detected and the zero-crossing rate to be detected of the sound signal to be detected are calculated, and the detection result of the speech signal detection of the previous frame of sound signal to be detected is obtained. The detection result of the sound signal to be detected is determined based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal, wherein the detection result of the sound signal is a result that characterizes whether the sound signal is a speech signal. Compared with the currently commonly used short-time energy detection method, this embodiment performs detection based on the autocorrelation and zero-crossing rate of the sound signal. By utilizing the different zero-crossing rates of the sound signal, it is possible to effectively distinguish between unvoiced and voiced sounds to determine whether there is a speech signal in the sound signal. By utilizing the different autocorrelation functions of the noise signal and the speech signal, it is possible to effectively distinguish between the noise signal and the speech signal. Therefore, this embodiment achieves accurate detection of non-speech signals and speech signals in sound signals with a low signal-to-noise ratio, thereby improving the applicability of the speech signal detection method.

[0067] Furthermore, in one embodiment, the voice signal detection method is applied to a head-mounted device, and after step S30, the method further includes:

[0068] Step S40 : When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a voice signal, turning on the transparency mode of the head mounted device.

[0069] In this embodiment, the voice signal detection method is applied to a head-mounted device, which can be a head-mounted display, such as a virtual reality device; or it can be an earphone device, such as an ear-hook earphone device, an in-ear earphone device, etc., and there is no specific limitation here.

[0070] In this embodiment, the sound signal to be detected can be obtained by sampling the ambient sound signal of the external environment in which the head-mounted device is located. Specifically, in this embodiment, after determining the detection result of the sound signal to be detected, the head-mounted device can be controlled according to the detection result of the sound signal to be detected.

[0071] Specifically, when the detection result of the sound signal to be detected is determined to be a voice signal, it can be determined that there is a voice signal in the external environment where the head-mounted device is located. At this time, the transparency mode of the head-mounted device can be turned on so that the user can communicate normally with the outside world when using the head-mounted device.

[0072] Step S50 : When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, an active noise reduction mode of the head mounted device is turned on.

[0073] In this embodiment, when the detection result of the sound signal to be detected is determined to be a non-voice signal, it can be determined that there is no voice signal in the external environment where the head-mounted device is located. At this time, the active noise reduction mode of the head-mounted device can be turned on to reduce external interference to the user when using the head-mounted device and improve the user's experience of using the head-mounted device.

[0074] It should be noted that the voice signal detection method is applied to a head-mounted device. When it is determined that the detection result of the sound signal to be detected is a voice signal, the transparency mode of the head-mounted device can be turned on to enable the user to communicate normally with the outside world when using the head-mounted device; when it is determined that the detection result of the sound signal to be detected is a non-voice signal, the active noise reduction mode of the head-mounted device can be turned on to reduce external interference to the user when using the head-mounted device, thereby improving the user's experience of using the head-mounted device.

[0075] In this embodiment, a frame of sound signal is sampled at a preset sampling rate, the sampled frame of sound signal is used as the sound signal to be detected, the autocorrelation value to be detected and the zero-crossing rate to be detected of the sound signal to be detected are calculated, and the detection result of the speech signal detection of the previous frame of sound signal to be detected is obtained. The detection result of the sound signal to be detected is determined based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal. Compared with the currently commonly used short-time energy detection method, this embodiment performs detection based on the autocorrelation and zero-crossing rate of the sound signal. By utilizing the different zero-crossing rates of the sound signal, it is possible to effectively distinguish between unvoiced and voiced sounds to determine whether there is a speech signal in the sound signal. By utilizing the different autocorrelation functions of the noise signal and the speech signal, it is possible to effectively distinguish between the noise signal and the speech signal. Therefore, this embodiment achieves accurate detection of non-speech signals and speech signals in sound signals with a low signal-to-noise ratio, thereby improving the applicability of the speech signal detection method.

[0076] Furthermore, based on the above-mentioned first embodiment, a second embodiment of the speech signal detection method of the present invention is proposed. In this embodiment, before step S30, the method further includes:

[0077] Step S60, obtaining a preset range of autocorrelation values to be detected, wherein the minimum value in the preset range is the signal autocorrelation value of a preset non-speech signal, and the maximum value in the preset range is the signal autocorrelation value of a preset speech signal;

[0078] In this embodiment, the autocorrelation value to be detected and the zero-crossing rate to be detected may be used as detection contents respectively, and the detection result of the sound signal to be detected is determined according to the detection results of the two.

[0079] Specifically, in this embodiment, a preset range of autocorrelation values to be detected is obtained. In this embodiment, the minimum value in the preset range may be the signal autocorrelation value of a preset non-speech signal, and the maximum value in the preset range may be the signal autocorrelation value of a preset speech signal. The preset speech signal and the preset non-speech signal may be any sound signal and are not limited herein.

[0080] In a specific embodiment, the value of the signal autocorrelation value of the preset non-speech signal is not limited. For example, it can be the average of the autocorrelation values of each signal point in the preset non-speech signal, or it can be the maximum value of the autocorrelation values of each signal point in the preset non-speech signal. There is no specific limitation here. Similarly, the value of the signal autocorrelation value of the preset speech signal is also not limited.

[0081] Step S70, obtaining a preset zero-crossing rate, wherein the preset zero-crossing rate is determined based on the preset signal zero-crossing rate of the non-speech signal and the preset signal zero-crossing rate of the speech signal;

[0082] In this embodiment, a preset zero-crossing rate is obtained. Specifically, in this embodiment, the preset zero-crossing rate can be determined based on a preset signal zero-crossing rate of a non-speech signal and a preset signal zero-crossing rate of a speech signal. For example, in one embodiment, the value of the preset zero-crossing rate can be a value between the preset signal zero-crossing rate of a non-speech signal and the preset signal zero-crossing rate of a speech signal. The value can be set according to actual needs and is not limited here.

[0083] In a specific implementation, the preset range and the preset zero-crossing rate may be obtained through laboratory testing or may be set based on engineer experience, and are not limited herein.

[0084] The step S30 includes:

[0085] Step S301, determining the detection result of the sound signal to be detected based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame sound signal.

[0086] In this embodiment, after obtaining the preset range and the preset zero-crossing rate, the detection result of the sound signal to be detected can be determined based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame of sound signal.

[0087] Specifically, in one embodiment, the size between the autocorrelation value to be detected and the preset range can be compared, and the size between the zero-crossing rate to be detected and the preset zero-crossing rate can be compared. According to the comparison result of the two values, the detection result of the sound signal to be detected is determined. For example, in one embodiment, if it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, then the detection result of the sound signal to be detected is determined to be a non-speech signal; if it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, then the detection result of the sound signal to be detected is determined to be a speech signal; if the autocorrelation value to be detected is within the preset range, then the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal. The specific detection conditions can be set according to actual needs and are not limited here.

[0088] In another embodiment, after comparing the autocorrelation value to be detected with a preset range, and comparing the zero-crossing rate to be detected with a preset zero-crossing rate, other parameters may be further tested based on the comparison result of the two values to determine the detection result of the sound signal to be detected. For example, in one embodiment, if it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, it may be detected whether the product of the autocorrelation value to be detected and the zero-crossing rate to be detected is less than a preset threshold value. If so, it is considered that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal. The specific detection conditions can be set according to the actual needs of the user and are not limited here.

[0089] In this embodiment, a preset range of autocorrelation values to be detected is obtained, wherein the minimum value in the preset range is the signal autocorrelation value of the preset non-speech signal, and the maximum value in the preset range is the signal autocorrelation value of the preset speech signal; and a preset zero-crossing rate is obtained, wherein the preset zero-crossing rate is determined based on the signal zero-crossing rate of the preset non-speech signal and the signal zero-crossing rate of the preset speech signal. The detection result of the sound signal to be detected is determined based on the size of the autocorrelation value to be detected relative to the preset range, the size of the zero-crossing rate to be detected relative to the preset zero-crossing rate, and the detection result of the previous frame of sound signal. In this embodiment, dual-threshold detection is performed on the sound signal to be detected. By utilizing the different zero-crossing rates of the sound signal, it is possible to effectively distinguish between unvoiced and voiced sounds to determine whether there is a speech signal in the sound signal. At the same time, by utilizing the different autocorrelation functions of the noise signal and the speech signal, it is possible to effectively distinguish between the noise signal and the speech signal, thereby improving the accuracy of detecting non-speech signals and speech signals in sound signals with low signal-to-noise ratios and improving the applicability of the speech signal detection method.

[0090] Furthermore, based on the first and / or second embodiments described above, a third embodiment of the speech signal detection method of the present invention is proposed. In this embodiment, when the detection result of the previous frame of sound signal is that the previous frame of sound signal is a non-speech signal, step S301 includes:

[0091] Step S3011: When it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, after cumulatively adding one to the first number and setting the second number to zero, detecting whether the first number is greater than a first preset threshold, and detecting whether the sum of the first number and a third number is greater than a second preset threshold, wherein the first number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is greater than the maximum value in the preset range and whose detected zero-crossing rate is greater than the preset zero-crossing rate; the second number is the number of detected sound signals for which the first number is greater than the third preset threshold when the detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate; and the third number is the number of detected sound signals whose detected autocorrelation value is within the preset range;

[0092] In this embodiment, each frame of sound signal that has completed speech signal detection is called a detected sound signal, the autocorrelation value of any frame of detected sound signal is called a detected autocorrelation value, and the zero-crossing rate of any frame of detected sound signal is called a detected zero-crossing rate.

[0093] In this embodiment, the detected sound signals whose detected autocorrelation value is greater than the maximum value in the preset range and whose detected zero-crossing rate is greater than the preset zero-crossing rate can be considered as speech signals. The number of detected sound signals that meet this condition is referred to as the first number below.

[0094] When the detected autocorrelation value is less than the minimum value within a preset range and the detected zero-crossing rate is less than a preset zero-crossing rate, a first number of detected sound signals greater than a preset threshold (hereinafter referred to as the third preset threshold for distinction) can be considered non-speech signals between speech signals. The number of detected sound signals meeting this condition is hereinafter referred to as the second number. In a specific embodiment, the value of the third preset threshold is greater than or equal to zero.

[0095] The detected sound signals whose detected autocorrelation values are within the preset range are considered to be relatively quiet speech signals. The number of detected sound signals that meet this condition is referred to as the third number hereinafter.

[0096] Specifically, in this embodiment, when it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, it is considered that the sound signal to be detected may be a speech signal. At this time, the first number is accumulated and one is added and the second number is set to zero.

[0097] After determining that the sound signal to be detected is a speech signal, the number of times a speech signal appears before the sound signal to be detected is detected to determine whether the sound signal to be detected is a frame of speech signal among multiple frames of valid speech signals, thereby determining the detection result of the sound signal to be detected. Specifically, it is detected whether the first number is greater than a first preset threshold, and whether the sum of the first number and the third number is greater than a second preset threshold. The first preset threshold and the second preset threshold are greater than or equal to zero, and the specific values can be set according to actual needs.

[0098] Step S3012: When it is determined that the first number is less than or equal to the first preset threshold, or the sum of the first number and the third number is less than or equal to the second preset threshold, determining that the detection result of the to-be-detected sound signal is consistent with the detection result of the previous frame of sound signal;

[0099] When it is determined that the first number is less than or equal to the first preset threshold, or the sum of the first number and the third number is less than or equal to the second preset threshold, it is determined that the sound signal to be detected is not a frame of speech signal in multiple frames of valid speech signals. At this time, the sound signal to be detected may be an invalid speech signal that is short and does not contain important information, for example, a sudden impact sound in a quiet environment. At this time, the detection result of the sound signal to be detected is determined to be consistent with the detection result of the previous frame of sound signal. In this embodiment, the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal.

[0100] Step S3013: When it is determined that the first number is greater than the first preset threshold and the sum of the first number and the third number is greater than the second preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the first number and the third number are set to zero.

[0101] In this embodiment, when it is determined that the first number is greater than a first preset threshold and the sum of the first number and the third number is greater than a second preset threshold, it is determined that the sound signal to be detected is a sound signal frame among the multiple frames of valid sound signals. In this case, the detection result of the sound signal to be detected can be determined as that the sound signal to be detected is a speech signal. Each count is set to zero, that is, the first number and the third number are set to zero, to prevent the current counts from affecting the detection result of the next sound signal frame.

[0102] It should be noted that when the previous frame of sound signal is a non-speech signal, the sound signal to be detected can be determined to be a speech signal when it is determined that the autocorrelation value to be detected is greater than the maximum value in a preset range and the zero-crossing rate to be detected is greater than a preset zero-crossing rate. By detecting the number of times a speech signal appears, it can be determined whether the sound signal to be detected does not contain a speech signal in a frame of multiple frames of speech signals, thereby determining the detection result of the sound signal to be detected, which can improve the accuracy of the detection result.

[0103] Furthermore, in one embodiment, before step S3011, the following steps are further included:

[0104] Step S3014, detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate;

[0105] In this embodiment, it is detected whether the autocorrelation value to be detected is within a preset range, and it is detected whether the zero-crossing rate to be detected is greater than a preset zero-crossing rate, so that different operations are performed according to different results.

[0106] After step S3014, the method further includes:

[0107] Step S3015: When the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0108] When the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, the sound signal to be detected may be a sudden and loud noise, for example, an impact sound in a quiet environment. At this time, the detection result of the sound signal to be detected is determined to be consistent with the detection result of the previous frame sound signal. In this embodiment, the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal.

[0109] Step S3016 : when the autocorrelation value to be detected is smaller than the minimum value in the preset range and the zero-crossing rate to be detected is larger than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal.

[0110] When the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the sound signal to be detected may be a lighter non-speech signal in the environment, such as noise. At this time, the detection result of the sound signal to be detected is determined to be consistent with the detection result of the previous frame sound signal. In this embodiment, the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal.

[0111] Furthermore, in one embodiment, after detecting whether the autocorrelation value to be checked is within the preset range in step S3014, the method further includes:

[0112] Step S3017: When it is determined that the autocorrelation value to be detected is within the preset range, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal, the third number is cumulatively increased by one and the second number is set to zero.

[0113] When it is determined that the autocorrelation value to be detected is within a preset range, the sound signal to be detected is considered to be a relatively quiet speech signal. In this case, the third number is cumulatively incremented by one and the second number is set to zero to prevent technical interference with the detection result of non-speech signals between speech frames in subsequent speech signal detection. In this embodiment, the detection result of the sound signal to be detected is determined to be consistent with the detection result of the previous frame of sound signal, that is, the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal.

[0114] It should be noted that by counting possible voice signals (i.e., counting the third number), it is possible to determine whether the sound signal to be detected is a frame of voice signal in multiple frames of voice signals based on the first number and the third number, thereby improving the accuracy of the detection result of the sound signal to be detected.

[0115] Furthermore, in one embodiment, after step S3014, the following steps are further included:

[0116] Step S3018: When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, detecting whether the first number is greater than the third preset threshold;

[0117] In this embodiment, when it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, it is considered that the sound signal to be detected may be a non-speech signal or a non-speech signal between speech signals. Therefore, it is detected whether the first number is greater than the third preset threshold to determine whether the sound signal to be detected is a non-speech signal or a non-speech signal between speech signals.

[0118] Step S3019: When it is determined that the first number is less than or equal to the third preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal;

[0119] When it is determined that the first number is less than or equal to the third preset threshold, it can be determined that there are no multiple frames of speech signals in the detected sound signal before the sound signal to be detected. At this time, the detection result of the sound signal to be detected is determined to be a non-speech signal.

[0120] Step S3020: When it is determined that the first number is greater than the third preset threshold, the second number is incremented by one, and then it is detected whether the second number is greater than a fourth preset threshold.

[0121] When it is determined that the first number is greater than the third preset threshold, it can be determined that multiple frames of speech signals appear before the sound signal to be detected. In this case, the sound signal to be detected may be a non-speech signal between speech signals, and the second number is accumulated by one. To further determine whether the sound signal to be detected is a non-speech signal between speech signals, it is tested whether the second number is greater than a fourth preset threshold.

[0122] Step S3021: When it is determined that the second number is greater than the fourth preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and setting the first number, the second number, and the third number to zero;

[0123] When it is determined that the second number is greater than the fourth preset threshold, it can be determined that non-speech signals appear between multiple frames of continuous speech signals. At this time, the detection result of the sound signal to be detected is determined to be a non-speech signal, and the first number, the second number and the third number are set to 0.

[0124] Step S3022: When it is determined that the second number is less than or equal to the fourth preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal.

[0125] When it is determined that the second number is less than or equal to the fourth preset threshold, it is determined that no non-speech signal appears between multiple frames of continuous speech signals. At this time, the sound signal to be detected is determined to be a non-speech signal between speech signals, and the detection result of the sound signal to be detected is determined to be consistent with the detection result of the previous frame of sound signal. In this embodiment, the sound to be detected is a non-speech signal.

[0126] It should be noted that, by detecting whether the first number is greater than the third preset threshold when it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, and when it is determined that the first number is greater than the third preset threshold, adding one to the second number, and then detecting whether the second number is greater than the fourth preset threshold, it can be determined whether the voice signal to be detected is a non-speech signal between voice signals or a non-speech signal, so that different operations are performed for different situations to improve the accuracy of the result of seeing the voice signal in the album.

[0127] In this embodiment, when the previous frame of the sound signal is a non-speech signal, the sound signal to be detected can be determined to be a speech signal when it is determined that the autocorrelation value to be detected is greater than the maximum value within a preset range and the zero-crossing rate to be detected is greater than a preset zero-crossing rate. By detecting the number of times a speech signal appears, it is determined whether the sound signal to be detected does not contain a speech signal in a frame of the multiple frames of speech signals, thereby determining the detection result of the sound signal to be detected, which can improve the accuracy of the detection result.

[0128] Furthermore, based on the first and / or second embodiments described above, a fourth embodiment of the speech signal detection method of the present invention is proposed. In this embodiment, when the detection result of the previous frame of sound signal is that the previous frame of sound signal is a speech signal, step S301 includes:

[0129] Step S3023: When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, after cumulatively adding one to the fourth number and setting the fifth number to zero, detecting whether the fourth number is greater than a fifth preset threshold, wherein the fourth number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate, and the fifth number is the number of detected sound signals for which the fourth number is greater than a sixth preset threshold when the detected autocorrelation value is greater than the maximum value in the preset range;

[0130] In this embodiment, the detected sound signal whose detected autocorrelation value is less than the minimum value in the preset range and whose detected zero crossing rate is less than the preset zero crossing rate is considered to be a non-speech signal. The number of detected sound signals that meet this condition is hereinafter referred to as the fourth number.

[0131] When the detected autocorrelation value is greater than the maximum value within the preset range, the fourth number of detected sound signals greater than the sixth preset threshold is considered to be silence signals between non-speech signals. The number of detected sound signals meeting this condition is hereinafter referred to as the fifth number. The value of the fifth preset threshold is greater than or equal to zero, and the specific value is not limited herein.

[0132] Specifically, in this embodiment, when it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, it can be determined that the sound signal to be detected is a non-speech signal. At this time, the fourth number is accumulated and added by one, and the fifth number is set to zero to avoid the counting of speech signals between non-speech signals affecting the detection results.

[0133] After determining that the sound signal to be detected is a non-speech signal, the number of non-speech signals that appear before the sound signal to be detected is detected to determine whether the sound signal to be detected is a non-speech signal in the multiple frames of non-speech signals. Specifically, the fourth number is detected to determine whether it is greater than a fifth preset threshold. In a specific embodiment, the value of the fifth preset threshold is greater than or equal to zero, and can be set as needed and is not limited here.

[0134] Step S3024: When it is determined that the fourth number is less than or equal to the fifth preset threshold, determining that the detection result of the to-be-detected sound signal is consistent with the detection result of the previous frame of sound signal;

[0135] When it is determined that the fourth number is less than or equal to the fifth preset threshold, it is determined that the sound signal to be detected is not a frame of non-speech signal among multiple frames of non-speech signals. At this time, it is considered that the sound signal to be detected may be a non-speech signal between speech signals. Therefore, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal. In this embodiment, the detection result of the sound signal to be detected is determined to be that the sound signal to be detected is a speech signal.

[0136] Step S3025: When it is determined that the fourth number is greater than the fifth preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and the fourth number is set to zero.

[0137] When it is determined that the fourth number is greater than the fifth preset threshold, the sound signal to be detected is determined to be a non-speech signal among the multiple frames of non-speech signals. In this case, the detection result of the sound signal to be detected can be determined as the non-speech signal. The fourth number is set to zero to prevent the counting of the non-speech signal from affecting the detection result of the next sound signal frame.

[0138] It should be noted that when the previous frame of the sound signal is a speech signal, the sound signal to be detected can be determined to be a non-speech signal by determining that the autocorrelation value to be detected is less than the minimum value in a preset range and the zero-crossing rate to be detected is less than a preset zero-crossing rate. In this case, the number of occurrences of the non-speech signal is detected to determine whether the sound signal to be detected is a non-speech signal frame among multiple frames of non-speech signals, thereby determining the detection result of the sound signal to be detected, which can improve the accuracy of the detection result.

[0139] Furthermore, in one embodiment, before step S3023, the following steps are further included:

[0140] Step S3026, detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate;

[0141] In this embodiment, it is detected whether the autocorrelation value to be detected is within a preset range, and it is detected whether the zero-crossing rate to be detected is greater than a preset zero-crossing rate, so that different operations are performed according to different results.

[0142] After step S3026, the method further includes:

[0143] Step S3027 : When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal.

[0144] When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, at this time, the sound signal to be detected may be a relatively light noise, and the detection result of the sound signal to be detected is considered to be consistent with the detection result of the previous frame sound signal. In this embodiment, the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal.

[0145] Furthermore, in one embodiment, after step S313, the following steps are further included:

[0146] Step S3028: When it is determined that the autocorrelation value to be detected is greater than the minimum value in the preset range, detecting whether the fourth number is greater than the sixth preset threshold;

[0147] In this embodiment, when it is determined that the autocorrelation value to be detected is greater than the minimum value within a preset range, the sound signal to be detected may be a speech signal or a short speech signal between non-speech signals. In this case, whether the fourth number is greater than a sixth preset threshold is detected to determine whether the sound signal to be detected is a speech signal or a short speech signal between non-speech signals. In a specific embodiment, the value of the sixth preset threshold is greater than or equal to zero and can be set according to actual needs and is not limited here.

[0148] Step S3029: When it is determined that the fourth number is greater than the sixth preset threshold, the fifth number is incremented by one, and then it is detected whether the fifth number is greater than a seventh preset threshold.

[0149] When it is determined that the first number is greater than the sixth preset threshold, a short voice signal between the non-voice signals of the sound signal to be detected is determined. At this time, after the fifth number is accumulated and added by one, it is detected whether the fifth number is greater than the seventh preset threshold to detect whether a voice signal appears between multiple consecutive frames of non-voice signals in the target sound signal to determine the detection result of the sound signal to be detected.

[0150] Step S3030: When it is determined that the fifth number is less than or equal to the seventh preset threshold, determining that the detection result of the to-be-detected sound signal is consistent with the detection result of the previous frame of sound signal;

[0151] When it is determined that the fifth number is less than or equal to the seventh preset threshold, it is determined that there is no speech signal between multiple consecutive frames of non-speech signals in the target sound signal, and the sound signal to be detected may be a short speech signal that does not contain valid information. At this time, it is considered that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal. In this embodiment, the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal.

[0152] Step S3031: When it is determined that the fifth number is greater than the seventh preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the fourth number and the fifth number are set to zero.

[0153] When it is determined that the fifth number is greater than the seventh preset threshold, it is determined that a speech signal appears between multiple consecutive frames of non-speech signals in the target sound signal. At this time, the detection result of the sound signal to be detected is determined to be that the sound signal to be detected is a speech signal, and the fourth number and the fifth number are set to zero to avoid the current count affecting the detection result of the next frame of sound signal.

[0154] It should be noted that when the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, the fourth number is detected to be greater than the sixth preset threshold value to determine whether the sound signal to be detected is a speech signal or a short-term speech signal between non-speech signals. When it is determined that the fourth number is less than or equal to the sixth preset threshold value, the sound signal to be detected is determined to be a short-term speech signal between non-speech signals. At this time, after adding one to the fifth number, it is detected whether the fifth number is greater than the seventh preset threshold value to detect whether the sound signal to be detected is a speech signal between non-speech signals for multiple consecutive frames, so as to determine the detection result of the sound signal to be detected of the target sound signal, thereby improving the accuracy of the detection result.

[0155] In this embodiment, when the previous frame of the sound signal is a speech signal, the sound signal to be detected can be determined to be a non-speech signal by determining that the autocorrelation value to be detected is less than a minimum value within a preset range and the zero-crossing rate to be detected is less than a preset zero-crossing rate. In this case, the number of occurrences of the non-speech signal is detected to determine whether the sound signal to be detected is a non-speech signal frame among multiple frames of non-speech signals, thereby determining the detection result of the sound signal to be detected, which can improve the accuracy of the detection result.

[0156] Furthermore, in one embodiment, the voice signal detection method is applied to a head-mounted device. In this embodiment, the active noise reduction mode and the transparency mode of the head-mounted device are controlled according to the detection result of the sound signal to be detected. Figure 2 、 Figure 3 and Figure 4 The process diagram shown is:

[0157] Initialization parameters: Get the minimum and maximum values in the preset range and the preset zero-crossing rate (i.e. Figure 2As shown in the figure, the voice frame count (count) is set to 0, the silent frame mute (number of frames) is set to 0, and the detection thresholds T1, T2 and Z0 are determined). In this embodiment, the process of determining the detection threshold can be: picking up a preset non-voice signal through a head-mounted microphone, analyzing and calculating the signal autocorrelation value and signal zero-crossing rate of the non-voice signal, and picking up a preset voice signal through a head-mounted microphone, analyzing and calculating the signal autocorrelation value and signal zero-crossing rate of the voice signal. The autocorrelation value of the preset non-voice signal is used as the minimum value of the preset range, and the autocorrelation value of the preset voice signal is used as the maximum value of the preset range. According to the zero-crossing rates of the preset non-voice signal and the preset voice signal, a value between the zero-crossing rate of the preset non-voice signal and the zero-crossing rate of the preset voice signal is determined as the preset zero-crossing rate. Furthermore, in a specific embodiment, different head-mounted devices have different gains and microphone sensitivities, which will result in different preset ranges and preset zero-crossing rates for different head-mounted devices. The preset ranges and preset zero-crossing rates corresponding to different head-mounted devices can be obtained according to the above method for different head-mounted devices.

[0158] A frame of sound signal is sampled according to a preset sampling rate. In this embodiment, the preset sampling rate may be 400 samples / s (i.e. Figure 2 Take a frame of MIC (Microphone) data as shown in ).

[0159] The sampled sound signal is processed and the processed sound signal is used as the sound signal to be detected (i.e. Figure 2 As shown in , the DC component is removed and a Hamming window is added to a frame of data.

[0160] Calculate the zero-crossing rate to be detected (i.e. Figure 2 Calculate the number of zero crossings in the current frame as shown in ).

[0161] Calculate the autocorrelation value to be detected. In this embodiment, the maximum value of the autocorrelation function of the sound signal to be detected is used as the autocorrelation value to be detected (ie Figure 2 Find the autocorrelation of a frame of data as shown in , and take the maximum autocorrelation).

[0162] Get the detection result of the previous frame of sound signal (ie Figure 2 The current state is determined as shown in the voice_flag, so that the head-mounted device can be tested differently based on the current state. Furthermore, in one embodiment, the current state can also be determined based on the power-on mode of the head-mounted device. That is, when the head-mounted device is in active noise reduction mode, the current state is determined to be 0; when the head-mounted device is in transparency mode, the current state is determined to be 1. Furthermore, in one embodiment, when the detection result of the previous frame of sound signal is inconsistent with the mode of the head-mounted device, the current state can be determined based on the mode of the head-mounted device.

[0163] In this embodiment, when the detection result of the previous frame of sound signal is a non-speech signal (i.e. Figure 3 (When the current state is 0, refer to Figure 3 As shown in the flow chart, at this time: when it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the first number is cumulatively added by one and the second number is set to zero, and then it is detected whether the first number is greater than the first preset threshold value, and whether the sum of the first number and the third number is greater than the second preset threshold value (i.e. Figure 3 As shown in , when the autocorrelation to be detected is greater than T2, the zero-crossing rate to be detected is greater than Z0, the voice frame count is +1, the silent frame between voices is mute=0, and the sum of the voice frame count and the possible voice frames is detected to be greater than min_voice, and the voice frame is detected to be greater than 5);

[0164] When it is determined that the first number is less than or equal to the first preset threshold, or the sum of the first number and the third number is less than or equal to the second preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal (i.e. Figure 3 As shown in , when the sum of the voice frame count and the possible voice frames is less than or equal to min_voice, or the detected voice frame is less than or equal to 5, then return to take the next frame of sound signal);

[0165] When it is determined that the first number is greater than the first preset threshold and the sum of the first number and the third number is greater than the second preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the first number and the third number are set to zero (i.e. Figure 3 As shown in , when the sum of the voice frames and the possible voice frames is greater than min_voice, or the detected voice frames are greater than 5, the state voice_flag is 1, the voice frame count = 0 and the possible voice frame = 0).

[0166] When it is determined that the autocorrelation value to be detected is within the preset range, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal, the third number is cumulatively increased by one and the second number is set to zero (i.e. Figure 3 As shown in , when the autocorrelation to be detected is between T1 and T2, there may be a speech frame + one, and a silent frame between voices (mute = 0);

[0167] When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, it is detected whether the first number is greater than a third preset threshold (i.e. Figure 3 Detect whether the voice frame count is greater than 0);

[0168] When it is determined that the first number is less than or equal to the third preset threshold, the detection result of the sound signal to be detected is determined to be a non-speech signal (ie Figure 3 As shown in , when the voice frame count is less than or equal to 0, take the next frame of sound signal);

[0169] When it is determined that the first number is greater than the third preset threshold, the second number is accumulated and added by one, and then the second number is checked to see if it is greater than the fourth preset threshold (i.e. Figure 3 As shown in , mute+1 is added to the silent frame between voices, and the mute frame between voices is detected to see if it is greater than mute_T);

[0170] When it is determined that the second number is greater than the fourth preset threshold, the detection result of the sound signal to be detected is determined to be a non-speech signal, and the first number, the second number and the third number are set to zero (i.e. Figure 3 , the voice frame count = 0, the silent frame mute = 0, and the possible voice frame = 0);

[0171] When it is determined that the second number is less than or equal to the fourth preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal (ie Figure 3 As shown in , when the inter-voice silence frame mute is less than or equal to mute_T, the next frame of sound signal is taken).

[0172] Determine that the detection result of the previous frame of sound signal is a speech signal (i.e. Figure 4 , when it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, the fourth number is cumulatively added by one and the fifth number is set to zero, and then it is detected whether the fourth number is greater than the fifth preset threshold (i.e. Figure 4 As shown in , when the autocorrelation to be detected is less than T1 and the zero-crossing rate to be detected is less than Z0, the silent frame is mute+1 and the speech count between the silent frames is 0);

[0173] When it is determined that the fourth number is less than or equal to the fifth preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal;

[0174] When it is determined that the fourth number is greater than the fifth preset threshold, the detection result of the sound signal to be detected is determined to be a non-speech signal, and the fourth number is set to zero (ie Figure 4 Set the silent frames to 0 as shown in .

[0175] When it is determined that the autocorrelation value to be detected is greater than the minimum value in the preset range, it is detected whether the fourth number is greater than the sixth preset threshold (i.e. Figure 4 Detect whether the silent frame is greater than 0);

[0176] When it is determined that the fourth number is greater than the sixth preset threshold, the fifth number is accumulated and added by one, and then it is detected whether the fifth number is greater than the seventh preset threshold (i.e. Figure 4 As shown in , the silent room voice count is increased by one, and the silent room voice count is detected to see if it is greater than count_T);

[0177] When it is determined that the fifth number is less than or equal to the seventh preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal (ie Figure 4 Take the next frame of sound signal as shown in );

[0178] When it is determined that the fifth number is greater than the seventh preset threshold, the detection result of the sound signal to be detected is determined to be a speech signal, and the fourth number and the fifth number are set to zero (ie Figure 4 The voice_flag shown in is 1).

[0179] The embodiment of the present invention also provides a voice signal detection device, referring to Figure 5 The speech signal detection device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0180] Those skilled in the art will understand that Figure 5 The structure shown in the figure does not constitute a limitation on the speech signal detection device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0181] like Figure 5 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a data storage module, a network communication module, a user interface module, and a voice signal detection program.

[0182] exist Figure 5 In the voice signal detection device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the voice signal detection device of the present invention can be set in the voice signal detection device, and the voice signal detection device calls the voice signal detection program stored in the memory 1005 through the processor 1001 and executes the steps of the voice signal detection method provided in the embodiment of the present invention.

[0183] The various embodiments of the speech signal detection device of the present invention may refer to the various embodiments of the speech signal detection method of the present invention, and will not be described in detail here.

[0184] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a speech signal detection program is stored. When the speech signal detection program is executed by a processor, the steps of the speech signal detection method described above are implemented.

[0185] The various embodiments of the computer-readable storage medium of the present invention may refer to the various embodiments of the speech signal detection method of the present invention, and will not be described in detail here.

[0186] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0187] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0188] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0189] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A speech signal detection method, characterized in that: The speech signal detection method comprises the following steps: Sampling a frame of sound signal according to a preset sampling rate, and using the sampled frame of sound signal as the sound signal to be detected; Calculating the autocorrelation value and zero-crossing rate of the sound signal to be detected, and obtaining a detection result obtained by performing speech signal detection on a previous frame of the sound signal to be detected; Determining a detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and a detection result of a previous frame of the sound signal, wherein the detection result of the sound signal is a result indicating whether the sound signal is a speech signal; When the detection result of the previous frame of sound signal is that the previous frame of sound signal is a speech signal, the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of sound signal includes: When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, after cumulatively adding one to the fourth number and setting the fifth number to zero, detecting whether the fourth number is greater than a fifth preset threshold, wherein the fourth number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate, and the fifth number is the number of detected sound signals for which the fourth number is greater than a sixth preset threshold when the detected autocorrelation value is greater than the maximum value in the preset range; When it is determined that the fourth number is less than or equal to the fifth preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal; When it is determined that the fourth number is greater than the fifth preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and the fourth number is set to zero.

2. The speech signal detection method according to claim 1, wherein: Before the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal, the method further includes: Obtaining a preset range of autocorrelation values to be detected, wherein a minimum value in the preset range is a signal autocorrelation value of a preset non-speech signal, and a maximum value in the preset range is a signal autocorrelation value of a preset speech signal; A preset zero-crossing rate is obtained, wherein the preset zero-crossing rate is determined based on the signal zero-crossing rate of the preset non-speech signal and the signal zero-crossing rate of the preset speech signal.

3. The speech signal detection method according to claim 2, wherein: When the detection result of the previous frame of sound signal is that the previous frame of sound signal is a non-speech signal, the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected and the detection result of the previous frame of sound signal further includes: When it is determined that the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, after cumulatively adding one to the first number and setting the second number to zero, detecting whether the first number is greater than a first preset threshold, and detecting whether the sum of the first number and the third number is greater than a second preset threshold, wherein the first number is the number of detected sound signals for which speech signal detection has been completed and whose detected autocorrelation value is greater than the maximum value in the preset range and whose detected zero-crossing rate is greater than the preset zero-crossing rate; the second number is the number of detected sound signals for which the first number is greater than the third preset threshold when the detected autocorrelation value is less than the minimum value in the preset range and whose detected zero-crossing rate is less than the preset zero-crossing rate; and the third number is the number of detected sound signals whose detected autocorrelation value is within the preset range; When it is determined that the first number is less than or equal to the first preset threshold, or the sum of the first number and the third number is less than or equal to the second preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal; When it is determined that the first number is greater than the first preset threshold and the sum of the first number and the third number is greater than the second preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the first number and the third number are set to zero.

4. The speech signal detection method according to claim 3, wherein: Before the step of adding one to the first quantity and setting the second quantity to zero, the method further includes: Detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate; After the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes: When the autocorrelation value to be detected is greater than the maximum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal; When the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is greater than the preset zero-crossing rate, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal; When it is determined that the autocorrelation value to be detected is within the preset range, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal, the third number is cumulatively increased by one and the second number is set to zero.

5. The speech signal detection method according to claim 4, wherein: After the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes: When it is determined that the autocorrelation value to be detected is less than the minimum value in the preset range and the zero-crossing rate to be detected is less than the preset zero-crossing rate, detecting whether the first number is greater than the third preset threshold; When it is determined that the first number is less than or equal to the third preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal; When it is determined that the first number is greater than the third preset threshold, adding one to the second number, and then detecting whether the second number is greater than a fourth preset threshold; When it is determined that the second number is greater than the fourth preset threshold, determining that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-speech signal, and setting the first number, the second number, and the third number to zero; When it is determined that the second number is less than or equal to the fourth preset threshold, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal.

6. The speech signal detection method according to claim 1, wherein: Before the step of adding one to the fourth quantity and setting the fifth quantity to zero, the method further includes: Detecting whether the autocorrelation value to be detected is within the preset range, and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate; After the steps of detecting whether the autocorrelation value to be detected is within the preset range and detecting whether the zero-crossing rate to be detected is greater than the preset zero-crossing rate, the method further includes: When it is determined that the autocorrelation value to be detected is smaller than the minimum value in the preset range and the zero-crossing rate to be detected is greater than a preset zero-crossing rate, it is determined that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame sound signal.

7. The speech signal detection method according to claim 6, wherein: After the step of detecting whether the autocorrelation value to be detected is within the preset range, the method further includes: When it is determined that the autocorrelation value to be detected is greater than the minimum value in the preset range, detecting whether the fourth number is greater than the sixth preset threshold; When it is determined that the fourth number is greater than the sixth preset threshold, adding one to the fifth number, and then detecting whether the fifth number is greater than a seventh preset threshold; When it is determined that the fifth number is less than or equal to the seventh preset threshold, determining that the detection result of the sound signal to be detected is consistent with the detection result of the previous frame of sound signal; When it is determined that the fifth number is greater than the seventh preset threshold, it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a speech signal, and the fourth number and the fifth number are set to zero.

8. The speech signal detection method according to any one of claims 1 to 7, wherein: The speech signal detection method is applied to a head-mounted device, and after the step of determining the detection result of the sound signal to be detected based on the autocorrelation value to be detected, the zero-crossing rate to be detected, and the detection result of the previous frame of the sound signal, further includes: When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a voice signal, turning on a transparency mode of the head mounted device; When it is determined that the detection result of the sound signal to be detected is that the sound signal to be detected is a non-voice signal, an active noise reduction mode of the head mounted device is turned on.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a speech signal detection program, which, when executed by a processor, implements the steps of the speech signal detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice activity detection method and device and voice recognition method and device

    CN108346425A

  • Voice recognition method and device and device for voice recognition

    CN113707130A