Voice detection method, device and computer-readable storage medium
By collecting and analyzing the time domain characteristic parameters of noise data and speech data in the speech detection environment and adjusting the speech detection threshold, the problem of poor adaptability of fixed thresholds is solved, and higher speech detection accuracy is achieved.
Patent Information
- Application Number
- CN202211493538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-11-25
AI Technical Summary
In the prior art, when the user environment changes, the speech detection method has poor adaptability to the fixed threshold, resulting in low speech detection accuracy.
The environmental noise data and voice data are collected in the speech detection environment, the time domain characteristic parameters are extracted, the speech detection threshold is adjusted according to these parameters, and the threshold interval with stronger adaptability is established, and the speech detection is performed through the first speech detection threshold and the second speech detection threshold.
It improves the accuracy of voice detection, can better distinguish between user voice and environmental noise, and improves the detection performance of the product in different environments.
Smart Images

Figure CN115881167B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a speech detection method, device, and computer-readable storage medium. Background Art
[0002] Speech detection generally refers to the detection of speech endpoints. Its purpose is to find the starting and ending points of a speech signal and to detect whether it exists. Speech detection methods mainly include threshold-based methods and pattern recognition-based methods.
[0003] For threshold-based speech detection methods, configuring an appropriate detection threshold directly affects the accuracy of the detection results. However, as the user's environment changes, the background noise level also changes. This can cause the fixed threshold calibrated at the factory to be less adaptable to different environments, resulting in low speech detection accuracy. Summary of the Invention
[0004] The main purpose of the present invention is to provide a speech detection method, device and computer-readable storage medium, aiming to provide a solution for adjusting the speech detection threshold in a speech detection environment to improve the accuracy of speech detection.
[0005] To achieve the above object, the present invention provides a voice detection method, which comprises the following steps:
[0006] Collecting environmental noise data and speech data in a speech detection environment;
[0007] Extracting a first time-domain feature parameter from the ambient noise data, and extracting a second time-domain feature parameter from the speech data;
[0008] determining a first speech detection threshold according to the first time domain feature parameter;
[0009] Each threshold interval is established according to the first speech detection threshold and the preset first speech endpoint value, and the first speech endpoint value is adjusted according to the distribution of the second time domain feature parameter in each threshold interval to obtain the second speech detection threshold, so as to perform speech detection based on the first speech detection threshold and the second speech detection threshold.
[0010] Optionally, the step of collecting ambient noise data and speech data in a speech detection environment includes:
[0011] When a first recording instruction is detected, collecting the environmental noise data according to the first recording instruction;
[0012] When a second recording instruction is detected, the voice data is collected according to the second recording instruction.
[0013] Optionally, the steps of extracting the first time domain feature parameter from the ambient noise data and extracting the second time domain feature parameter from the speech data include:
[0014] Dividing the environmental noise data into first audio frames, and determining first time-domain feature parameters of each of the first audio frames;
[0015] The speech data is divided into second audio frames, and second time domain feature parameters of each second audio frame are determined.
[0016] Optionally, the step of determining a first speech detection threshold according to the first time domain feature parameter includes:
[0017] Obtaining a removal parameter, and removing audio frames whose corresponding first time-domain feature parameters exceed the removal parameter from each of the first audio frames to obtain remaining audio frames;
[0018] Calculate an average value of the first time domain feature parameters corresponding to the remaining audio frames as the first speech detection threshold.
[0019] Optionally, the step of establishing respective threshold intervals according to the first speech detection threshold and a preset first speech endpoint value includes:
[0020] Taking the first speech detection threshold as a lower endpoint, taking the first speech endpoint value as an upper endpoint, and taking an interval smaller than the lower endpoint as a first threshold interval;
[0021] taking an interval greater than or equal to the lower endpoint and less than or equal to the upper endpoint as a second threshold interval;
[0022] An interval greater than the upper endpoint is used as a third threshold interval.
[0023] Optionally, the step of adjusting the first speech endpoint value according to the distribution of the second time domain feature parameter in each threshold interval to obtain the second speech detection threshold includes:
[0024] Recording the number of first frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the first threshold range, recording the number of second frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the second threshold range, and recording the number of third frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the third threshold range;
[0025] determining a maximum number of frames among the first number of frames, the second number of frames, and the third number of frames;
[0026] If the third frame number is the maximum frame number and the third frame number meets the preset frame number threshold condition, the first speech endpoint value is used as the second speech detection threshold.
[0027] Optionally, after the step of determining the maximum number of frames among the first number of frames, the second number of frames, and the third number of frames, the method further includes:
[0028] If the first number of frames is the maximum number of frames, performing the step of collecting voice data in the voice detection environment;
[0029] If the second number of frames is the maximum number of frames, the first speech endpoint value is reduced to obtain a second speech endpoint value, and a new threshold interval is established based on the second speech endpoint value and the first speech detection threshold.
[0030] Optionally, before the step of extracting the second time domain feature parameter from the voice data, the method further includes:
[0031] Collect motion data based on inertial sensing module;
[0032] If the motion data meets the preset environmental threshold condition, prompt information about the change of the voice detection environment is output.
[0033] To achieve the above object, the present invention further provides a speech detection device, comprising:
[0034] An acquisition module, used for acquiring environmental noise data and speech data in a speech detection environment;
[0035] an extraction module, configured to extract a first time-domain feature parameter from the ambient noise data and a second time-domain feature parameter from the speech data;
[0036] a determination module, configured to determine a first speech detection threshold according to the first time domain feature parameter;
[0037] An adjustment module is used to establish each threshold interval according to the first speech detection threshold and the preset first speech endpoint value, adjust the first speech endpoint value according to the distribution of the second time domain feature parameter in each threshold interval, and obtain a second speech detection threshold, so as to perform speech detection based on the first speech detection threshold and the second speech detection threshold.
[0038] To achieve the above-mentioned purpose, the present invention also provides a speech detection device, which includes: a memory, a processor, and a speech detection program stored in the memory and executable on the processor. When the speech detection program is executed by the processor, the steps of the speech detection method described above are implemented.
[0039] In addition, to achieve the above-mentioned purpose, the present invention further proposes a computer-readable storage medium, on which a speech detection program is stored. When the speech detection program is executed by a processor, the steps of the speech detection method described above are implemented.
[0040] In the present invention, environmental noise data and voice data are collected in a voice detection environment; a first time domain feature parameter is extracted from the environmental noise data, and a second time domain feature parameter is extracted from the voice data; a first voice detection threshold is determined based on the first time domain feature parameter; each threshold interval is established based on the first voice detection threshold and a preset first voice endpoint value, and the first voice endpoint value is adjusted based on the distribution of the second time domain feature parameter in each threshold interval to obtain a second voice detection threshold, so as to perform voice detection based on the first voice detection threshold and the second voice detection threshold. The collected environmental noise data and voice data contain the environmental background noise in the voice detection environment, and the process of extracting the time domain feature parameters corresponds to the time domain detection method used in voice detection. During the voice detection process, the first voice detection threshold and the second voice detection threshold can replace the original fixed threshold in the product, and more accurately distinguish the user's voice from the noise in the environment, thereby improving the voice detection accuracy of the product in the voice detection environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic diagram of the hardware operating environment involved in an embodiment of the present invention;
[0042] Figure 2 Schematic diagram of the flow of the first embodiment of the speech detection method of the present invention;
[0043] Figure 3 1 is a flow chart of a second embodiment of a speech detection method according to the present invention;
[0044] Figure 4 1 is a flow chart of a third embodiment of a speech detection method according to the present invention;
[0045] Figure 5 1 is a flow chart of a fourth embodiment of a speech detection method according to the present invention;
[0046] Figure 6 1 is a flow chart of a fifth embodiment of a speech detection method according to the present invention;
[0047] Figure 7 Schematic diagram of a speech detection device according to an embodiment of the present invention.
[0048] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] like Figure 1 As shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.
[0051] It should be noted that the voice detection device in the embodiment of the present invention can be a headset, a smart phone, a personal computer, a server and other devices, and is not specifically limited here.
[0052] like Figure 1 As shown, the speech detection device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or it may be a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally also be a storage device independent of the aforementioned processor 1001.
[0053] Those skilled in the art will understand that Figure 1 The device structure shown in the figure does not constitute a limitation on the voice detection device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0054] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module and a voice detection program. The operating system is a program that manages and controls the hardware and software resources of the device and supports the operation of the voice detection program and other software or programs. Figure 1 In the device shown, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server; and the processor 1001 can be used to call the voice detection program stored in the memory 1005 and execute the steps of the voice detection method in the following embodiments.
[0055] Based on the above structure, various embodiments of the speech detection method are proposed.
[0056] Reference Figure 2 , Figure 2 FIG. 1 is a flow chart of a first embodiment of a speech detection method according to the present invention.
[0057] The embodiment of the present invention provides an embodiment of a voice detection method. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than here. In this embodiment, the execution subject of the voice detection method can be a headset, a personal computer, a smart phone or other device, which is not limited in this embodiment. For ease of description, the following embodiments are described with the headset as the execution subject. In this embodiment, the voice detection method includes:
[0058] Step S10, collecting environmental noise data and speech data in a speech detection environment;
[0059] The voice detection environment refers to the environment in which the headset performs voice detection, which is generally the same as the environment in which the headset is used. Voice detection refers to the process of detecting the starting and ending points of speech. Ambient noise data refers to the noise-free data collected by the headset in the voice detection environment. Voice data refers to the data containing both speech and noise collected by the headset in the voice detection environment.
[0060] This embodiment does not limit the method of collecting environmental noise data and voice data. For example, the headset can receive the user's voice command through voice recognition and collect environmental noise data and voice data according to the voice command.
[0061] Step S20, extracting a first time domain feature parameter from the ambient noise data, and extracting a second time domain feature parameter from the speech data;
[0062] The first time domain feature parameter refers to the parameter obtained by performing time domain feature extraction on the ambient noise data. The second time domain feature parameter refers to the parameter obtained by performing time domain feature extraction on the speech data. Time domain feature extraction refers to the operation of extracting time domain features related to speech signal processing. This embodiment does not limit the method of extracting the first time domain feature parameter and the method of extracting the second time domain feature parameter, and the method of extracting the first time domain feature parameter and the method of extracting the second time domain feature parameter can be the same. In a feasible implementation, the ambient noise data and the speech data can be input into a time domain feature extraction program, and the first time domain feature parameter and the second time domain feature parameter can be output by the time domain feature extraction program.
[0063] Step S30, determining a first speech detection threshold according to the first time domain feature parameter;
[0064] The first voice detection threshold refers to the detection threshold used in the voice detection process. Before the earphone product leaves the factory, the tester can perform debugging in a laboratory environment and set a set of fixed voice detection thresholds in the earphone, including a lower limit and an upper limit. During the voice detection process, when the signal value of the detected voice signal exceeds the upper limit, it indicates that the voice has begun. When the signal value is lower than the lower limit and remains for a certain period of time, it indicates that the voice has ended. Therefore, the setting of the voice detection threshold will directly affect the accuracy of voice detection. The first voice detection threshold can be the lower limit used in voice detection. This embodiment does not limit the method for determining the first voice detection threshold. In a feasible implementation method, the median of the first detection parameter can be directly used as the first voice detection threshold. In this way, the first voice detection threshold obtained is the lower limit value that adapts to the background noise in the environment. When the noise in the environment is large, it avoids the situation where the noise is mistaken for a voice signal.
[0065] Step S40, establish each threshold interval according to the first speech detection threshold and the preset first speech endpoint value, adjust the first speech endpoint value according to the distribution of the second time domain feature parameter in each threshold interval, and obtain the second speech detection threshold, so as to perform speech detection based on the first speech detection threshold and the second speech detection threshold.
[0066] The preset first speech endpoint value is a threshold value set based on the upper limit of speech detection. The first speech endpoint value can be obtained by using a fixed threshold value stored in the headset as the first speech endpoint value. A threshold interval refers to an interval established with the first speech detection threshold and the first speech endpoint value as the dividing point. The distribution refers to the distribution of the number of second time-domain feature parameters falling within each threshold interval. The second speech detection threshold refers to the detection threshold used during speech detection and can be the upper limit used for speech detection. Adjustment of the first speech endpoint value can be performed by increasing, decreasing, or fixing. This embodiment does not limit the method for adjusting based on the distribution. In one feasible implementation, if there are second time-domain feature parameters that exceed the first speech endpoint value and the number of such second time-domain feature parameters exceeds a preset number threshold, the first speech endpoint value can be determined to be the second speech detection threshold. If there are second time-domain feature parameters that exceed the first speech endpoint value and the number of such second time-domain feature parameters does not exceed the preset number threshold, the first speech endpoint value can be reduced and the number distribution relationship re-determined until the adjusted first speech endpoint value meets the threshold condition for serving as the second speech detection threshold.
[0067] In this embodiment, environmental noise data and voice data are collected in a voice detection environment; a first time domain feature parameter is extracted from the environmental noise data, and a second time domain feature parameter is extracted from the voice data; a first voice detection threshold is determined based on the first time domain feature parameter; each threshold interval is established based on the first voice detection threshold and a preset first voice endpoint value, and the first voice endpoint value is adjusted based on the distribution of the second time domain feature parameter in each threshold interval to obtain a second voice detection threshold, so as to perform voice detection based on the first voice detection threshold and the second voice detection threshold. The collected environmental noise data and voice data contain the ambient noise background in the voice detection environment, and the process of extracting the time domain feature parameters corresponds to the time domain detection method used in voice detection. During the voice detection process, the first voice detection threshold and the second voice detection threshold can replace the original fixed threshold in the product, more accurately distinguishing the user's voice from the noise in the environment, thereby improving the voice detection accuracy of the product in the voice detection environment.
[0068] Further, based on the above first embodiment, a second embodiment of the speech detection method of the present invention is proposed. In this embodiment, referring to Figure 3 , the method comprising:
[0069] Step S101, when a first recording instruction is detected, collecting the environmental noise data according to the first recording instruction;
[0070] The first recording instruction refers to an instruction for recording ambient noise data received by the headset. When a user uses the headset, the headset can be used as an independent device to play audio, or it can communicate with a smart terminal through a wired connection or Bluetooth connection, and the smart terminal controls the headset to perform various operations. The smart terminal that communicates with the headset can be a smartphone, personal computer, tablet computer, or other device. The following is an example of the communication connection between a smartphone and headset.
[0071] The smartphone may include an application for calibrating the voice detection threshold of the headset. When the user wants to calibrate, he or she can follow the prompts of the application to perform corresponding operations in the operation interface of the smartphone. For example, by pressing and holding a button in the operation interface, the smartphone can notify the headset to record the ambient noise data. The smartphone can send a first recording instruction to the headset via SPP (Serial Port Profile). When the headset receives the first recording instruction, it collects the audio signal through the microphone and converts the audio signal into ambient noise data. It is understandable that the smartphone can display a prompt message in the operation interface to remind the user not to make any sound during the collection process, so that the collected ambient noise data is voice-free data.
[0072] The first recording instruction may include control information regarding the recording duration. If the recording duration exceeds a first preset duration, the headset may be controlled to stop recording. The first preset duration may be set to 20 seconds to prevent the recorded audio signal from being too long. The duration of the ambient noise data must also be limited. If the duration of the ambient noise data is detected to be too short, the collected ambient noise data may be discarded, prompting the user to re-record.
[0073] Step S102: When a second recording instruction is detected, the voice data is collected according to the second recording instruction.
[0074] The second recording instruction refers to the instruction for recording voice data received by the headset. Referring to the usage scenario of the above-mentioned smartphone connected to the headset, after recording the environmental noise data, you can continue to record the audio data of the voice in the environment. The smartphone can send the second recording instruction to the headset through SPP. When the headset receives the second recording instruction, it collects the audio signal through the microphone and converts the audio signal into voice data. It is understandable that the smartphone can display text information in the operation interface, prompt the user to read the text content aloud, and send out a continuous voice signal, so that the collected voice data is continuous voice data.
[0075] The second recording instruction may include control information regarding the recording duration. If the recording duration exceeds a second preset duration, the headset may be controlled to stop recording. The second preset duration may be set to 10 seconds to prevent the recorded audio signal from being too long. The duration of the voice data must also be limited. If this is detected, the captured voice data may be discarded, prompting the user to re-record.
[0076] It is understandable that the collection of environmental noise data and speech data should be performed in the same speech detection environment. In a feasible implementation manner, before the step of extracting the second time domain feature parameter from the speech data, the following steps may also be included:
[0077] Step S103, collecting motion data based on the inertial sensor module;
[0078] An inertial sensing module refers to a sensing module capable of measuring inertial parameters such as speed, acceleration, and angular velocity, for example, a three-axis accelerometer or a six-axis accelerometer. Motion data refers to data generated by the movement of the earphones or the phone when the user uses the earphones. The motion data can be speed, acceleration, and / or angular velocity. The inertial sensing module can collect motion data according to a preset sampling period. The inertial sensing module that collects motion data can be the inertial sensing module in the mobile phone or the inertial sensing module in the earphones.
[0079] Step S104: If the motion data meets the preset environmental threshold condition, output prompt information about the change of the voice detection environment.
[0080] Environmental threshold conditions are conditions used to determine whether the user's environment has changed. In one embodiment, motion data can be used to calculate the distance traveled within a preset time period. If the distance traveled within the preset time period exceeds a preset distance threshold, a prompt message indicating a change in the detected environment is displayed, reminding the user to maintain the same environment during the detection process. This prompt message can be output via a mobile phone or headphones, and can be displayed as text or as a voice reminder.
[0081] In this embodiment, the user can control the earphones to collect data when the user feels that the voice detection accuracy has decreased, and use the data in the real environment to calibrate the voice detection threshold so that the calibrated detection threshold is adapted to the environment. The user is reminded through prompt information to maintain consistency of the environment during the voice detection process, thereby improving the user experience.
[0082] Further, based on the above-mentioned first and / or second embodiments, a third embodiment of the speech detection method of the present invention is proposed. In this embodiment, referring to Figure 4 , the method comprising:
[0083] Step S201: Divide the environmental noise data into first audio frames, and determine first time-domain feature parameters of each first audio frame;
[0084] The first audio frame refers to the audio frame obtained by dividing the ambient noise data according to the preset frame length and frame shift. The first time domain characteristic parameter can be a short-time autocorrelation, a short-time energy or a zero-crossing rate. Taking the short-time autocorrelation as an example, the short-time autocorrelation of each audio frame can be calculated using the autocorrelation function, and the maximum autocorrelation is used as the first time domain characteristic parameter. The autocorrelation function is a function used to measure the similarity of the signal's own time waveform. In the process of calculating the maximum autocorrelation value, the short-time autocorrelation of each first audio frame can be calculated using the formula of the autocorrelation function R(τ)=E[x(t)x(t+τ)] to form an autocorrelation sequence, and then the maximum value in the autocorrelation sequence is found, which is the maximum autocorrelation value.
[0085] Step S202: Divide the speech data into second audio frames, and determine second time-domain feature parameters of each second audio frame.
[0086] The second audio frame refers to the audio frame obtained by dividing the voice data according to a preset frame length and frame shift. The frame length and frame shift used to divide the second audio frame can be the same as those of the first audio frame. The second time domain feature parameter can be a short-time autocorrelation, short-time energy, or zero-crossing rate. The type of the second time domain feature parameter is the same as that of the first time domain feature parameter to unify the measurement standard of the voice detection threshold. The determination of the second time domain feature parameter can refer to the calculation process of the first time domain feature parameter described above, and will not be repeated here.
[0087] In this embodiment, the ambient noise data and speech data are divided into audio frames, and the audio frames are processed to obtain time domain feature parameters. The time domain feature parameters can be short-time autocorrelation, short-time energy or zero-crossing rate, which can correspond to the time domain detection parameters of the detection threshold in speech detection, ensuring that the obtained speech detection threshold is adapted to the speech detection algorithm in the product.
[0088] Furthermore, based on the above-mentioned first, second and / or third embodiments, a fourth embodiment of the speech detection method of the present invention is proposed. In this embodiment, referring to Figure 5 , the method comprising:
[0089] Step S301: obtaining a removal parameter, removing audio frames whose corresponding first time-domain feature parameters exceed the removal parameter from the first audio frames, to obtain remaining audio frames;
[0090] The removal parameter refers to the time domain feature parameter used as a reference standard. Each frame of data has its own time domain feature parameter. Taking the time domain feature parameter of the short-time autocorrelation type as an example, the maximum autocorrelation value exceeding the removal parameter can be considered to have the possibility of accidental occurrence. Each first audio frame is traversed, and the first audio frame corresponding to the maximum autocorrelation value exceeding the removal parameter is removed. The data of the remaining audio frames obtained is used as the basis for calibrating the first speech detection threshold. The removal parameter can be set according to the actual situation, or the average value of the maximum autocorrelation values of all first audio frames can be calculated first, and 10 times the average value is used as the removal parameter.
[0091] Step S302: Calculate an average value of the first time-domain feature parameters corresponding to the remaining audio frames as the first speech detection threshold.
[0092] The first time-domain feature parameters corresponding to the remaining audio frames can be considered as parameters that serve as a reference for threshold calibration. The sum of the maximum autocorrelation values corresponding to each remaining audio frame is calculated and then divided by the number of remaining audio frames to obtain the average value, which is the first speech detection threshold in the speech detection environment. During speech detection, the first speech detection threshold serves as a lower limit for determining the presence and absence of speech signals.
[0093] In this embodiment, a time domain analysis method is used to process the environmental noise data in the speech detection environment. The environmental noise data is data in the real environment. The first speech detection threshold obtained after processing takes the noise in the environment into account the influence of the real environment, which can improve the accuracy of the speech end endpoint detection.
[0094] Furthermore, based on the first, second, third and / or fourth embodiments above, a fifth embodiment of the speech detection method of the present invention is proposed. In this embodiment, referring to Figure 6 , the method comprising:
[0095] Step S401, taking the first speech detection threshold as a lower endpoint, taking the first speech endpoint value as an upper endpoint, and taking an interval smaller than the lower endpoint as a first threshold interval;
[0096] The first speech endpoint value can be set according to the first speech detection threshold. For example, the first speech endpoint value and the first speech detection threshold are in a multiple relationship, and the first speech endpoint value is set to 3 times the first speech detection threshold. In the following, T1 is used to represent the first speech detection threshold, and T2 is used to represent the first speech endpoint value. The relationship between T1 and T2 can be expressed as T2=nT1, where n represents the multiple, and n can be set to 3. The first threshold interval can be expressed as (-∞, T1). The distribution of the second time domain feature parameter in each threshold interval can reflect the rationality of the setting of the first speech endpoint value.
[0097] Step S402: taking an interval greater than or equal to the lower endpoint and less than or equal to the upper endpoint as a second threshold interval;
[0098] The second threshold interval can be expressed as [T1, T2].
[0099] Step S403: taking the interval greater than the upper endpoint as the third threshold interval;
[0100] The third threshold interval can be expressed as (T2, +∞).
[0101] Step S404: Record the number of first frames in each of the second audio frames whose corresponding second time domain feature parameter is within the first threshold range, record the number of second frames in each of the second audio frames whose corresponding second time domain feature parameter is within the second threshold range, and record the number of third frames in each of the second audio frames whose corresponding second time domain feature parameter is within the third threshold range.
[0102] cnt may be used to represent the total number of frames of the second audio frame, cnt0 may represent the first frame number, cnt1 may represent the second frame number, and cnt2 may represent the third frame number.
[0103] Step S405, determining the maximum number of frames among the first number of frames, the second number of frames, and the third number of frames;
[0104] The maximum number of frames can be obtained by comparing the size relationship among cnt0, cnt1 and cnt2. The maximum number of frames can reflect the distribution of voice data in the above threshold range.
[0105] Step S406: If the maximum number of frames is the third number of frames and the third number of frames meets a preset frame number threshold condition, then the first speech endpoint value is used as the second speech detection threshold;
[0106] When the maximum number of frames is the third number of frames, the second time domain feature parameter corresponding to more second audio frames is greater than the first voice endpoint value, which is more consistent with the situation where the first voice endpoint value is the upper limit value of voice detection. If the third number of frames also meets the frame number threshold condition, then this first voice endpoint value is reasonably set and can be used as the second voice detection threshold. The frame number threshold condition refers to the condition that the third number of frames must meet to be the maximum number of frames. The frame number threshold condition can be that the third number of frames is greater than two-thirds of the total number of frames. In a feasible implementation, if the maximum number of frames is the third number of frames and the third number of frames does not meet the frame number threshold condition, the first voice endpoint value can be reduced and each threshold interval can be re-established until the maximum number of frames is the third number of frames and meets the frame number threshold condition.
[0107] Step S407: If the first number of frames is the maximum number of frames, executing the step of collecting voice data in the voice detection environment;
[0108] When the first frame number is the maximum frame number, the second time domain feature parameters corresponding to more second audio frames are less than the first voice detection threshold, and the first voice detection threshold is the lower limit value, indicating that the collected voice data is invalid. The step of obtaining voice data collected in the voice detection environment can be executed to re-collect valid voice data.
[0109] Step S408: If the second number of frames is the maximum number of frames, the first speech endpoint value is reduced to obtain a second speech endpoint value, and a new threshold interval is established based on the second speech endpoint value and the first speech detection threshold.
[0110] When the second frame number is the maximum frame number, the second time domain feature parameters corresponding to more second audio frames are between the first speech detection threshold and the first speech endpoint value. When the voice data is valid, the first speech endpoint value will be too high as the second speech detection threshold. The first speech endpoint value can be reduced, and the reduced second speech endpoint value can be used to divide each threshold interval. Then the frame number distribution of the second audio frame will also change. The second frame number may no longer be the maximum frame number. The maximum frame number can be re-determined until the second speech detection threshold is set appropriately.
[0111] In this embodiment, the first speech endpoint value is first used as the speech detection threshold to be calibrated, and the first speech endpoint value is adjusted according to the distribution of the second time domain feature parameter in each threshold interval, so that the obtained second speech detection threshold adapts to the influence of noise in the user's environment, avoiding the second speech detection threshold being too high to recognize the starting point of the speech.
[0112] In addition, the embodiment of the present invention also provides a voice detection device, referring to Figure 7 , the speech detection device comprises:
[0113] The acquisition module 10 is used to collect environmental noise data and speech data in a speech detection environment;
[0114] An extraction module 20, configured to extract a first time domain feature parameter from the ambient noise data and a second time domain feature parameter from the speech data;
[0115] a determination module 30, configured to determine a first speech detection threshold according to the first time domain feature parameter;
[0116] The adjustment module 40 is used to establish each threshold interval according to the first speech detection threshold and the preset first speech endpoint value, adjust the first speech endpoint value according to the distribution of the second time domain feature parameter in each threshold interval, and obtain the second speech detection threshold, so as to perform speech detection based on the first speech detection threshold and the second speech detection threshold.
[0117] Furthermore, the acquisition module 10 is further configured to:
[0118] When a first recording instruction is detected, collecting the environmental noise data according to the first recording instruction;
[0119] When a second recording instruction is detected, the voice data is collected according to the second recording instruction.
[0120] Furthermore, the extraction module 20 is further configured to:
[0121] Dividing the environmental noise data into first audio frames, and determining first time-domain feature parameters of each of the first audio frames;
[0122] The speech data is divided into second audio frames, and second time domain feature parameters of each second audio frame are determined.
[0123] Furthermore, the determining module 30 is further configured to:
[0124] Obtaining a removal parameter, and removing audio frames whose corresponding first time-domain feature parameters exceed the removal parameter from each of the first audio frames to obtain remaining audio frames;
[0125] Calculate an average value of the first time domain feature parameters corresponding to the remaining audio frames as the first speech detection threshold.
[0126] Furthermore, the adjustment module 40 is further configured to:
[0127] Taking the first speech detection threshold as a lower endpoint, taking the first speech endpoint value as an upper endpoint, and taking an interval smaller than the lower endpoint as a first threshold interval;
[0128] taking an interval greater than or equal to the lower endpoint and less than or equal to the upper endpoint as a second threshold interval;
[0129] An interval greater than the upper endpoint is used as a third threshold interval.
[0130] Furthermore, the adjustment module 40 is further configured to:
[0131] Recording the number of first frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the first threshold range, recording the number of second frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the second threshold range, and recording the number of third frames in which the second time domain feature parameter corresponding to each of the second audio frames is within the third threshold range;
[0132] determining a maximum number of frames among the first number of frames, the second number of frames, and the third number of frames;
[0133] If the third frame number is the maximum frame number and the third frame number meets the preset frame number threshold condition, the first speech endpoint value is used as the second speech detection threshold.
[0134] Furthermore, the adjustment module 40 is further configured to:
[0135] If the first number of frames is the maximum number of frames, performing the step of collecting voice data in the voice detection environment;
[0136] If the second number of frames is the maximum number of frames, the first speech endpoint value is reduced to obtain a second speech endpoint value, and a new threshold interval is established based on the second speech endpoint value and the first speech detection threshold.
[0137] Furthermore, the speech detection device further includes a detection module, which is used to:
[0138] Collect motion data based on inertial sensing module;
[0139] If the motion data meets the preset environmental threshold condition, prompt information about the change of the voice detection environment is output.
[0140] The various embodiments of the speech detection device of the present invention may refer to the various embodiments of the speech detection method of the present invention, and will not be described in detail here.
[0141] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a speech detection program is stored. When the speech detection program is executed by a processor, the steps of the speech detection method described above are implemented.
[0142] The various embodiments of the speech detection device and computer-readable storage medium of the present invention may refer to the various embodiments of the speech detection method of the present invention, and will not be described in detail here.
[0143] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0144] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0146] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A speech detection method, characterized in that: The voice detection method comprises the following steps: Collecting environmental noise data and speech data in a speech detection environment; Extracting a first time-domain feature parameter from the ambient noise data, and extracting a second time-domain feature parameter from the speech data; determining a first speech detection threshold according to the first time domain feature parameter; Establishing respective threshold intervals with the first speech detection threshold as a lower endpoint and a preset first speech endpoint value as an upper endpoint, wherein the threshold intervals include at least a third threshold interval greater than the upper endpoint; For each of the threshold intervals, recording the number of frames in which the corresponding second time domain feature parameter in each second audio frame of the voice data is within the threshold interval; If the third frame number among the frame numbers is not the maximum frame number or the third frame number does not meet the preset frame number threshold condition, re-collect voice data or reduce the first voice endpoint value; If the third frame number among each of the frame numbers is the maximum frame number and the third frame number meets the preset frame number threshold condition, the first speech endpoint value is used as the second speech detection threshold to perform speech detection based on the first speech detection threshold and the second speech detection threshold, wherein the third frame number is the frame number in which the second time domain feature parameter in each of the second audio frames is within the third threshold range.
2. The speech detection method according to claim 1, wherein: The step of collecting environmental noise data and voice data in a voice detection environment comprises: When a first recording instruction is detected, collecting the environmental noise data according to the first recording instruction; When a second recording instruction is detected, the voice data is collected according to the second recording instruction.
3. The speech detection method according to claim 1, wherein: The steps of extracting the first time domain feature parameter from the ambient noise data and extracting the second time domain feature parameter from the voice data include: Dividing the environmental noise data into first audio frames, and determining first time-domain feature parameters of each of the first audio frames; The speech data is divided into second audio frames, and second time domain feature parameters of each second audio frame are determined.
4. The speech detection method according to claim 3, wherein: The step of determining a first speech detection threshold according to the first time domain feature parameter comprises: Obtaining a removal parameter, and removing audio frames whose corresponding first time-domain feature parameters exceed the removal parameter from each of the first audio frames to obtain remaining audio frames; Calculate an average value of the first time domain feature parameters corresponding to the remaining audio frames as the first speech detection threshold.
5. The speech detection method according to claim 3, wherein: The step of establishing each threshold interval with the first speech detection threshold as the lower endpoint and the preset first speech endpoint value as the upper endpoint comprises: taking an interval smaller than the lower endpoint as a first threshold interval; taking an interval greater than or equal to the lower endpoint and less than or equal to the upper endpoint as a second threshold interval; An interval greater than the upper endpoint is used as a third threshold interval.
6. The speech detection method according to claim 5, wherein: The step of recording, for each threshold interval, the number of frames in which the corresponding second time domain feature parameter in each second audio frame of the speech data is within the threshold interval includes: Record the first number of frames in which the corresponding second time domain feature parameter in each of the second audio frames is in the first threshold interval, record the second number of frames in which the corresponding second time domain feature parameter in each of the second audio frames is in the second threshold interval, and record the third number of frames in which the corresponding second time domain feature parameter in each of the second audio frames is in the third threshold interval.
7. The speech detection method according to claim 6, wherein: The step of recollecting voice data or reducing the first voice endpoint value if the third frame number among the frame numbers is not the maximum frame number or the third frame number does not meet the preset frame number threshold condition includes: If the first number of frames is the maximum number of frames, performing the step of collecting voice data in the voice detection environment; If the second number of frames is the maximum number of frames, the first speech endpoint value is reduced to obtain a second speech endpoint value, and a new threshold interval is established based on the second speech endpoint value and the first speech detection threshold.
8. The speech detection method according to any one of claims 1 to 7, wherein: Before the step of extracting the second time domain feature parameter from the voice data, the method further includes: Collect motion data based on inertial sensing module; If the motion data meets the preset environmental threshold condition, prompt information about the change of the voice detection environment is output.
9. A voice detection device, characterized in that: The speech detection device includes: a memory, a processor, and a speech detection program stored in the memory and executable on the processor. When the speech detection program is executed by the processor, the steps of the speech detection method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a speech detection program, which, when executed by a processor, implements the steps of the speech detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for detecting speech endpoints and system
CN102522081A
Adaptive threshold setting voice endpoint detection method and device thereof, and readable storage medium
CN108847218A