Method, device, medium, program product and system for responding to speech termination point
By using the speech recognition model and a weighted finite state converter in the speech activation detection, the duration of the speech activation detection is dynamically adjusted, and the problem of misjudgment of speech termination points caused by speech pause is solved, and the integrity of the audio file and the accuracy of voice interaction control are improved.
Patent Information
- Application Number
- CN202411261638.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-09
AI Technical Summary
In the prior art, voice activation detection misjudgment of voice termination points due to speech pauses, resulting in incomplete audio files.
By inputting the collected audio files into the speech recognition model, the target decoding results are obtained, and the accuracy of the speech termination point is judged based on the weighted finite state converter, and the duration of the speech activation detection is dynamically adjusted to cancel the misjudged speech termination point response.
Improve the completeness of audio files obtained by voice activation detection, improve the accuracy of voice interaction control, and reduce the audio file blanking time caused by misjudgment.
Smart Images

Figure CN119132347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, medium, program product and system for responding to a voice termination point. Background Art
[0002] Voice Activity Detection (VAD) is a technology used to identify voice activity in audio signals. Its main purpose is to distinguish between speech and non-speech parts in an audio stream, such as silence, noise or background sound.
[0003] In the related technical solutions, the voice interaction system usually includes a voice activation detection module. The voice activation detection module removes non-voice data so that the subsequent automatic speech recognition module only needs to process voice data, thereby improving processing efficiency and reducing the probability of misrecognition.
[0004] Voice activation detection can detect whether an audio file contains speech and determine the start and end points of the speech. In some interactive scenarios, such as when a person pauses for a long time due to thinking while speaking, voice activation detection will determine that there is a speech end point due to the pause in speaking, making the audio file obtained by voice activation detection incomplete. Summary of the invention
[0005] The present invention aims to at least solve the problem in the prior art or related art that voice activation detection may determine the existence of a voice termination point due to a speaking pause, resulting in an incomplete audio file obtained by the voice activation detection.
[0006] To this end, a first aspect of the present invention is to provide a method for responding to a speech termination point.
[0007] A second aspect of the present invention is to provide an apparatus for responding to a speech end point.
[0008] A third aspect of the present invention is to provide another apparatus for responding to a speech end point.
[0009] A fourth aspect of the present invention provides a readable storage medium.
[0010] A fifth aspect of the present invention provides a computer program product.
[0011] A sixth aspect of the present invention is to provide a voice interaction system.
[0012] In view of this, according to a first aspect of the present invention, the present invention provides a method for responding to a speech termination point, comprising: inputting a collected audio file into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, the target decoding result comprising a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, the first text being a text obtained by the speech recognition model performing speech recognition on the audio file; based on the state corresponding to the first text in one or more target decoding results in the first weighted finite state converter being a termination state, determining a first duration according to the probability value of the first text in at least one target decoding result; within a first duration after a first moment, canceling the response to the detected speech termination point, the first moment being the moment of collection of the audio file.
[0013] The present invention proposes a method for responding to a speech end point. By running the method for responding to a speech end point, a first duration can be determined according to a collected audio file, and within the first duration after the collection time of the audio file, if a speech end point is detected, the response to the detected speech end point is canceled. In this technical solution, within the first duration, by canceling the response to the detected speech end point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related technical solution that the audio file detected by the voice activation detection is incomplete due to the presence of a speaking pause.
[0014] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0015] In addition, the method for responding to a speech termination point proposed in the present application also has the following additional technical features.
[0016] In some technical solutions, optionally, the speech termination point includes: a speech termination point given by a speech recognition model based on a neural network; and / or a speech termination point given by an automatic speech recognition model.
[0017] In this technical solution, the speech ending points given by different models can all be stopped from responding, thereby improving the accuracy of the speech ending points detected by the stop response.
[0018] In some technical solutions, optionally, the collected audio file is input into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, specifically including: inputting the collected audio file into the speech recognition model to extract at least one candidate text output by the first conversion layer in the speech recognition model, the first conversion layer being used to convert the audio file into the candidate text; decoding each candidate text separately to obtain at least one target decoding result.
[0019] In this technical solution, the speech recognition model includes a first conversion layer, wherein the first conversion layer is used to convert an audio file into a text to be selected. After the collected audio file is input into the speech recognition model, by extracting the content output by the first conversion layer, the text recognized from the audio file can be obtained, that is, at least one text to be selected. In this process, the collected audio file can be converted into text.
[0020] By decoding each candidate text, a target decoding result can be obtained from each candidate text.
[0021] In some technical solutions, optionally, prefix beam search decoding is used to decode each candidate text separately to obtain at least one target decoding result.
[0022] In some technical schemes, optionally, based on the fact that the state corresponding to the first text in one or more target decoding results in the first weighted finite state converter is a termination state, the first duration is determined according to the probability value of the first text in at least one target decoding result, specifically including: obtaining the first weighted finite state converter, the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the samples in the sample set are texts obtained when the speech termination point is incorrectly responded to; comparing the first state attribute with the target state attribute representing the termination state in the first weighted finite state converter, the first state attribute is related to the state corresponding to the first text in the first weighted finite state converter in at least one target decoding result; based on the first state attribute being the same as the target state attribute, determining that the state corresponding to the first text in the first weighted finite state converter is a termination state; determining the first duration according to the probability value of the first text in at least one target decoding result.
[0023] In this technical solution, the samples in the sample set are texts obtained when the speech end point is incorrectly responded to. It can be understood that the sample set is a set of samples that are incorrectly identified as having a speech end point during the speech activation detection process. In other words, each sample in the sample set should not be determined to have a speech end point, but is incorrectly determined to have a speech end point during the speech activation detection process.
[0024] Based on this, the first weighted finite state converter constructed using the sample set can represent the possible situation where the speech termination point is incorrectly determined to occur. Based on this, after the target decoding result is identified from the audio file, the state corresponding to the first text in the target decoding result in the first weighted finite state converter is compared with the first weighted finite state converter to determine whether the first text is a possible situation where the speech termination point is incorrectly determined to occur.
[0025] Specifically, as can be seen from the above, each node in the first weighted finite state converter has a state attribute. Therefore, after determining the first text, the state corresponding to the first text in the first weighted finite state converter can be determined, that is, the state attribute of the corresponding node can be determined. If the first state attribute is the same as the target state attribute representing the termination state in the first weighted finite state converter, it is considered that the first text is a complete path in the first weighted finite state converter, that is, the first text is a case where a speech termination point is erroneously determined to occur. At this time, the first duration is determined according to the probability value of the first text so as to stop responding to the detected speech termination point within the first duration.
[0026] In this process, by stopping the response to the detected speech end point, the speech activation detection can be continued, thereby delaying the time of responding to the speech end point, and finally, the audio file obtained by the speech activation detection contains a longer and more complete speech.
[0027] In addition, in the above technical solution, the first duration is determined according to the probability value of the first text. Compared with using a fixed duration, it can dynamically adjust the speech length of the audio file obtained by voice activation detection and reduce the audio file obtained by voice activation detection containing longer blanks.
[0028] In this process, the accuracy of the audio file obtained by voice activation detection can be improved, thereby improving the accuracy and recognition efficiency of the speech recognition model.
[0029] In some technical solutions, optionally, the sample set includes a second text and / or a third text; wherein the second text is a partial text extracted from a historical audio file, and the historical audio file also includes a fourth text, the fourth text is a text located after the second text, and the time interval between the second text and the fourth text is greater than or equal to a second duration; wherein the same interactive session includes a first sub-audio file and a second sub-audio file with the same prefix, and the third text is a text extracted from the first sub-audio file.
[0030] In this technical solution, the sample set includes three possible situations, that is, the sample set includes only the second text, the sample set includes only the third text, or the sample set includes the second text and the third text.
[0031] In some technical solutions, optionally, the historical audio file, that is, the audio file collected historically, is obtained by acquiring the audio file collected historically so as to analyze the audio file collected historically. Specifically, when performing speech recognition on the historical audio file, the recognized text and the timestamp corresponding to the recognized text are obtained. If it is detected that the time interval between the text before the timestamp (that is, the second text in this application) and the text after the timestamp (that is, the fourth text in this application) is greater than or equal to the second duration, it is considered that there is a speaking pause. For the speaking pause, it should not be determined that there is a speech termination point at the pause. Based on this, the text before the timestamp (that is, the second text in this application) is used as a sample in the sample set. Among them, the second duration can be set according to actual use needs.
[0032] In some technical solutions, optionally, the second sub audio file is located after the first sub audio file, the first sub audio file is a sub audio file not recognized by the natural language model, and the second sub audio file is a sub audio file recognized by the natural language model.
[0033] In this technical solution, the first sub-audio file and the second sub-audio file can be understood as voice commands repeatedly sent by the user to the device in an interactive session. When the first sub-audio file is a sub-audio file that is not recognized by the natural language model, and the second sub-audio file is a sub-audio file that is recognized by the natural language model, it is considered that the first sub-audio file is a possible case where a speech termination point is erroneously determined. Therefore, the text extracted from the first sub-audio file, that is, the third text, is used as a sample in the sample set.
[0034] In this technical solution, the sample set can cover common possible situations, so that the constructed first weighted finite state converter can meet the discrimination needs of most scenarios, thereby improving the accuracy of voice activation detection.
[0035] In some technical solutions, optionally, a deterministic operation algorithm and a minimization operation algorithm are used to optimize the weighted finite state converter.
[0036] In this technical solution, the use of a determinization algorithm and a minimization algorithm can use the least number of states to express a weighted finite state converter. In this process, the least number of state entries makes the weighted finite state converter more compact.
[0037] In some technical schemes, optionally, the first duration is determined based on the probability value of the first text in at least one target decoding result, specifically including: determining the maximum probability value among the probability values of at least one first text; obtaining the maximum delay duration and the minimum delay duration; and determining the first duration based on the maximum probability value, the maximum delay duration and the minimum delay duration.
[0038] In the above technical solution, the maximum probability value among the probability values of at least one first text is determined so as to select the first text with the highest output possibility, so as to determine the first duration based on the probability value of the first text with the highest possibility.
[0039] In this process, the first duration can be dynamically determined according to the probability value of the first text to improve the accuracy of the speech termination point detected by the stop response.
[0040] According to a second aspect of the present invention, the present invention provides a device for responding to a speech termination point, comprising: a processing unit, for inputting a collected audio file into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, the target decoding result comprising a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, the first text being a text obtained by the speech recognition model performing speech recognition on the audio file; a determination unit, for determining a first duration based on the state corresponding to the first text in the first weighted finite state converter in one or more target decoding results being a termination state, and according to the probability value of the first text in at least one target decoding result; a response unit, for canceling the response to the detected speech termination point within a first duration after a first moment, the first moment being the moment of collection of the audio file.
[0041] The present invention proposes a device for responding to a speech end point, which can determine a first duration according to a collected audio file, and cancel the response to the detected speech end point if a speech end point is detected within the first duration after the audio file is collected. In this technical solution, by canceling the response to the detected speech end point within the first duration, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related technical solution that the audio file detected by the voice activation detection is incomplete due to the presence of a speaking pause.
[0042] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0043] In addition, the device for responding to a speech termination point proposed in the present application also has the following additional technical features.
[0044] In some technical solutions, optionally, the speech termination point includes: a speech termination point given by a speech recognition model based on a neural network; and / or a speech termination point given by an automatic speech recognition model.
[0045] In this technical solution, the speech ending points given by different models can all be stopped from responding, thereby improving the accuracy of the speech ending points detected by the stop response.
[0046] In some technical solutions, optionally, the processing unit is specifically used to: input the collected audio file into the speech recognition model to extract at least one candidate text output by the first conversion layer in the speech recognition model, and the first conversion layer is used to convert the audio file into the candidate text; decode each candidate text separately to obtain at least one target decoding result.
[0047] In this technical solution, the speech recognition model includes a first conversion layer, wherein the first conversion layer is used to convert an audio file into a text to be selected. After the collected audio file is input into the speech recognition model, by extracting the content output by the first conversion layer, the text recognized from the audio file can be obtained, that is, at least one text to be selected. In this process, the collected audio file can be converted into text.
[0048] By decoding each candidate text, a target decoding result can be obtained from each candidate text.
[0049] In some technical solutions, optionally, prefix beam search decoding is used to decode each candidate text separately to obtain at least one target decoding result.
[0050] In some technical schemes, optionally, the determination unit is specifically used to: obtain a first weighted finite state converter, the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the sample in the sample set is a text obtained when an erroneous response speech termination point is given; compare a first state attribute with a target state attribute representing a termination state in the first weighted finite state converter, the first state attribute being related to a state corresponding to the first text in the first weighted finite state converter in at least one target decoding result; determine that the state corresponding to the first text in the first weighted finite state converter is a termination state based on the fact that the first state attribute is the same as the target state attribute; and determine a first duration according to a probability value of the first text in at least one target decoding result.
[0051] In this technical solution, the samples in the sample set are texts obtained when the speech end point is incorrectly responded to. It can be understood that the sample set is a set of samples that are incorrectly identified as having a speech end point during the speech activation detection process. In other words, each sample in the sample set should not be determined to have a speech end point, but is incorrectly determined to have a speech end point during the speech activation detection process.
[0052] Based on this, the first weighted finite state converter constructed using the sample set can represent the possible situation where the speech termination point is incorrectly determined to occur. Based on this, after the target decoding result is identified from the audio file, the state corresponding to the first text in the target decoding result in the first weighted finite state converter is compared with the first weighted finite state converter to determine whether the first text is a possible situation where the speech termination point is incorrectly determined to occur.
[0053] Specifically, as can be seen from the above, each node in the first weighted finite state converter has a state attribute. Therefore, after determining the first text, the state corresponding to the first text in the first weighted finite state converter can be determined, that is, the state attribute of the corresponding node can be determined. If the first state attribute is the same as the target state attribute representing the termination state in the first weighted finite state converter, it is considered that the first text is a complete path in the first weighted finite state converter, that is, the first text is a case where a speech termination point is erroneously determined to occur. At this time, the first duration is determined according to the probability value of the first text so as to stop responding to the detected speech termination point within the first duration.
[0054] In this process, by stopping the response to the detected speech end point, the speech activation detection can be continued, thereby delaying the time of responding to the speech end point, and finally, the audio file obtained by the speech activation detection contains a longer and more complete speech.
[0055] In addition, in the above technical solution, the first duration is determined according to the probability value of the first text. Compared with using a fixed duration, it can dynamically adjust the speech length of the audio file obtained by voice activation detection and reduce the audio file obtained by voice activation detection containing longer blanks.
[0056] In this process, the accuracy of the audio file obtained by voice activation detection can be improved, thereby improving the accuracy and recognition efficiency of the speech recognition model.
[0057] In some technical solutions, optionally, the sample set includes a second text and / or a third text; wherein the second text is a partial text extracted from a historical audio file, and the historical audio file also includes a fourth text, the fourth text is a text located after the second text, and the time interval between the second text and the fourth text is greater than or equal to a second duration; wherein the same interactive session includes a first sub-audio file and a second sub-audio file with the same prefix, and the third text is a text extracted from the first sub-audio file.
[0058] In this technical solution, the sample set includes three possible situations, that is, the sample set includes only the second text, the sample set includes only the third text, or the sample set includes the second text and the third text.
[0059] In some technical solutions, optionally, the historical audio file, that is, the audio file collected historically, is obtained by acquiring the audio file collected historically so as to analyze the audio file collected historically. Specifically, when performing speech recognition on the historical audio file, the recognized text and the timestamp corresponding to the recognized text are obtained. If it is detected that the time interval between the text before the timestamp (that is, the second text in this application) and the text after the timestamp (that is, the fourth text in this application) is greater than or equal to the second duration, it is considered that there is a speaking pause. For the speaking pause, it should not be determined that there is a speech termination point at the pause. Based on this, the text before the timestamp (that is, the second text in this application) is used as a sample in the sample set. Among them, the second duration can be set according to actual use needs.
[0060] In some technical solutions, optionally, the second sub audio file is located after the first sub audio file, the first sub audio file is a sub audio file not recognized by the natural language model, and the second sub audio file is a sub audio file recognized by the natural language model.
[0061] In this technical solution, the first sub-audio file and the second sub-audio file can be understood as voice commands repeatedly sent by the user to the device in an interactive session. When the first sub-audio file is a sub-audio file that is not recognized by the natural language model, and the second sub-audio file is a sub-audio file that is recognized by the natural language model, it is considered that the first sub-audio file is a possible case where a speech termination point is erroneously determined. Therefore, the text extracted from the first sub-audio file, that is, the third text, is used as a sample in the sample set.
[0062] In this technical solution, the sample set can cover common possible situations, so that the constructed first weighted finite state converter can meet the discrimination needs of most scenarios, thereby improving the accuracy of voice activation detection.
[0063] In some technical solutions, optionally, a deterministic operation algorithm and a minimization operation algorithm are used to optimize the weighted finite state converter.
[0064] In this technical solution, the use of a determinization algorithm and a minimization algorithm can use the least number of states to express a weighted finite state converter. In this process, the least number of state entries makes the weighted finite state converter more compact.
[0065] In some technical solutions, optionally, the determination unit is specifically used to: determine the maximum probability value among the probability values of at least one first text; obtain the maximum delay duration and the minimum delay duration; and determine the first duration based on the maximum probability value, the maximum delay duration and the minimum delay duration.
[0066] In the above technical solution, the maximum probability value among the probability values of at least one first text is determined so as to select the first text with the highest output possibility, so as to determine the first duration based on the probability value of the first text with the highest possibility.
[0067] In this process, the first duration can be dynamically determined according to the probability value of the first text to improve the accuracy of the speech termination point detected by the stop response.
[0068] According to the third aspect of the present invention, the present invention provides a device for responding to a speech termination point, comprising a processor and a memory, wherein the memory stores programs or instructions that can be executed on the processor, and when the programs or instructions are executed by the processor, the steps of the method for responding to a speech termination point as described above are implemented.
[0069] According to a fourth aspect of the present invention, the present invention provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method for responding to a voice termination point as described above are implemented.
[0070] According to a fifth aspect of the present invention, the present invention provides a computer program product, which is stored in a storage medium and implements the steps of the method for responding to a speech termination point as described above when the computer program product is executed by at least one processor.
[0071] According to a sixth aspect of the present invention, the present invention provides a voice interaction system, comprising: a device for responding to a voice termination point as described in any one of the above items; and / or a readable storage medium as described above; and / or a computer program product as described above.
[0072] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0074] Figure 1 A schematic flow chart of a method for responding to a speech termination point in an embodiment of the present invention is shown;
[0075] Figure 2 A schematic diagram showing an example of a first weighted finite state converter in an embodiment of the present invention;
[0076] Figure 3 A schematic diagram showing an audio file in an embodiment of the present invention;
[0077] Figure 4 A schematic block diagram of a device for responding to a speech termination point in an embodiment of the present invention is shown;
[0078] Figure 5 A schematic block diagram of another device for responding to a speech termination point in an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0079] In order to more clearly understand the above aspects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0080] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0081] In one embodiment of the present application, Figure 1 As shown, a method for responding to a speech termination point is provided, comprising:
[0082] Step 102, inputting the collected audio file into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, wherein the target decoding result includes a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, wherein the first text is a text obtained by the speech recognition model through speech recognition of the audio file;
[0083] Step 104, based on the state corresponding to the first text in the one or more target decoding results in the first weighted finite state switch being a terminal state, determine the first duration according to the probability value of the first text in at least one target decoding result;
[0084] Step 106, canceling the response to the detected speech end point within a first time period after a first moment, where the first moment is the acquisition moment of the audio file.
[0085] The present invention proposes a method for responding to a speech termination point. By running the method for responding to a speech termination point, a first duration can be determined according to a collected audio file, and within the first duration after the collection time of the audio file, if a speech termination point is detected, the response to the detected speech termination point is canceled. In this embodiment, within the first duration, by canceling the response to the detected speech termination point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related embodiments that the audio file detected by the voice activation detection is incomplete due to the presence of a speaking pause.
[0086] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0087] The audio file may be understood as a file obtained by collecting the user's speaking voice.
[0088] In one embodiment, optionally, the speech recognition model can be an automatic speech recognition (ASR) model, wherein the goal of the automatic speech recognition model is to convert the vocabulary content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0089] In one of the embodiments, optionally, the target decoding result is a result obtained by a speech recognition model decoding an audio file.
[0090] Exemplarily, when a user says “I want to watch TV”, the collected audio file includes the sound of “I want to watch TV”. At this time, the result obtained by the speech recognition model after decoding the audio file is “I want to watch TV”.
[0091] In one embodiment, optionally, the constructed first weighted finite state converter includes a starting node and an ending node, and there are multiple nodes between the starting node and the ending node. The multiple nodes are connected to the starting node and the ending node. The starting node and the multiple nodes, the multiple nodes and the ending node, and the multiple nodes can be connected by text, and each node has a corresponding state attribute.
[0092] Based on this, the state corresponding to the first text in the first weighted finite state converter can be understood as whether the first text is a path including a start node and an end node in the first weighted finite state converter.
[0093] In the above embodiment, judging whether the state corresponding to the first text in the first weighted finite state converter is a termination state can indirectly judge whether the first text is a text that needs to determine a first duration and stops responding to the detected speech termination point within the first duration after the first moment. In this process, the accuracy of stopping responding to the detected speech termination point can be improved.
[0094] In the above embodiment, the speech recognition model has the problem of recognition accuracy. Therefore, when performing speech recognition on the collected audio file, there is a situation where the output first text is inaccurate. Based on this, the probability value of the first text is used to represent the possibility of recognizing the first text from the collected audio file.
[0095] In some embodiments, the speech ending point corresponds to the speech starting point, wherein the speech ending point refers to the moment when the speech disappears when the audio file is subjected to speech activation detection. Similarly, the speech starting point can be understood as the moment when the speech appears when the audio file is subjected to speech activation detection.
[0096] In some embodiments, the acquisition time may be understood as the time when the audio file is acquired.
[0097] In some embodiments, optionally, the termination state is a state when the first text is connected to the termination node.
[0098] In some embodiments, optionally, the speech termination point includes: a speech termination point given by a speech recognition model based on a neural network; and / or a speech termination point given by an automatic speech recognition model.
[0099] Among them, the speech endpoint given by the neural network-based speech recognition model is also known as the NNVAD Endpoint, and the speech endpoint given by the automatic speech recognition model is also known as the ASR Endpoint.
[0100] In this embodiment, the speech ending points given by different models may all be stopped from responding, thereby improving the accuracy of the speech ending points detected by the stop response.
[0101] In some embodiments, optionally, the collected audio file is input into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, specifically including: inputting the collected audio file into the speech recognition model to extract at least one candidate text output by the first conversion layer in the speech recognition model, the first conversion layer being used to convert the audio file into the candidate text; decoding each candidate text separately to obtain at least one target decoding result.
[0102] In this embodiment, the speech recognition model includes a first conversion layer, wherein the first conversion layer is used to convert an audio file into a text to be selected. After the collected audio file is input into the speech recognition model, by extracting the content output by the first conversion layer, the text recognized from the audio file can be obtained, that is, at least one text to be selected. In this process, the collected audio file can be converted into text.
[0103] By decoding each candidate text, a target decoding result can be obtained from each candidate text.
[0104] In some embodiments, optionally, prefix beam search decoding is used to decode each candidate text separately to obtain at least one target decoding result.
[0105] Among them, prefix beam search decoding is a method based on beam search. It searches all possible output sequences and uses a scoring function to evaluate the possibility of each sequence. During the search process, the algorithm retains the scores of the prefixes and selects the most likely output sequence based on these scores. Among them, the score of the prefix is also the probability value of the first text in this application.
[0106] Among them, the main advantage of the prefix beam search algorithm is that it can process input and output sequences of different lengths and can dynamically adjust the size of the search space during the search process, which makes it more efficient than other search algorithms when processing long sequences.
[0107] In this process, the use of prefix beam search decoding can output the most likely candidate text among at least one candidate text as a target decoding result, so as to improve the accuracy and recognition efficiency of the speech recognition model.
[0108] In one of the embodiments, the text to be selected can be understood as the text that may be output from the first conversion layer when the speech recognition model recognizes the collected audio file.
[0109] In one embodiment, the first conversion layer is a Connectionist Temporal Classification layer, ie, a CTC layer.
[0110] In some embodiments, optionally, based on the fact that the state corresponding to the first text in one or more target decoding results in the first weighted finite state converter is a termination state, the first duration is determined according to the probability value of the first text in at least one target decoding result, specifically including: obtaining the first weighted finite state converter, the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the samples in the sample set are texts obtained when the speech termination point is incorrectly responded to; comparing the first state attribute with the target state attribute representing the termination state in the first weighted finite state converter, the first state attribute is related to the state corresponding to the first text in the first weighted finite state converter in at least one target decoding result; based on the first state attribute being the same as the target state attribute, determining that the state corresponding to the first text in the first weighted finite state converter is a termination state; determining the first duration according to the probability value of the first text in at least one target decoding result.
[0111] In this embodiment, the samples in the sample set are texts obtained when the speech end point is incorrectly responded to. It can be understood that the sample set is a set of samples that are incorrectly identified as having a speech end point during the speech activation detection process. In other words, each sample in the sample set should not be determined to have a speech end point, but is incorrectly determined to have a speech end point during the speech activation detection process.
[0112] Based on this, the first weighted finite state converter constructed using the sample set can represent the possible situation where the speech termination point is incorrectly determined to occur. Based on this, after the target decoding result is identified from the audio file, the state corresponding to the first text in the target decoding result in the first weighted finite state converter is compared with the first weighted finite state converter to determine whether the first text is a possible situation where the speech termination point is incorrectly determined to occur.
[0113] Specifically, as can be seen from the above, each node in the first weighted finite state converter has a state attribute. Therefore, after determining the first text, the state corresponding to the first text in the first weighted finite state converter can be determined, that is, the state attribute of the corresponding node can be determined. If the first state attribute is the same as the target state attribute representing the termination state in the first weighted finite state converter, it is considered that the first text is a complete path in the first weighted finite state converter, that is, the first text is a case where a speech termination point is erroneously determined to occur. At this time, the first duration is determined according to the probability value of the first text so as to stop responding to the detected speech termination point within the first duration.
[0114] In this process, by stopping the response to the detected speech end point, the speech activation detection can be continued, thereby delaying the time of responding to the speech end point, and finally, the audio file obtained by the speech activation detection contains a longer and more complete speech.
[0115] In addition, in the above embodiment, the first duration is determined based on the probability value of the first text. Compared with using a fixed duration, it is possible to dynamically adjust the speech length of the audio file obtained by voice activation detection and reduce the longer blank time contained in the audio file obtained by voice activation detection.
[0116] In this process, the accuracy of the audio file obtained by voice activation detection can be improved, thereby improving the accuracy and recognition efficiency of the speech recognition model.
[0117] Exemplarily, the state corresponding to the first text in the first weighted finite state converter is a first state attribute, wherein the first state attribute is a first encoding, and the target state attribute is a second encoding. If the first encoding is the same as the second encoding, the state corresponding to the first text in the first weighted finite state converter is a terminal state.
[0118] Specifically, the second code is 13. If the first state attribute is the first code, and the first code is also 13, then the state corresponding to the first text in the first weighted finite state converter is the terminal state.
[0119] In one embodiment, Figure 2 A schematic diagram showing an example of a first weighted finite state converter in an embodiment of the present application is shown as follows: Figure 2 As shown, the single-layer node 0 is the starting node and the double-layer node 13 is the ending node. Each path from the starting node to the ending node represents a sample. For example, "play a song" corresponds to 0→2→3→9→13; "please play" corresponds to 0→4→10→13.
[0120] Among them, the advantage of the first weighted finite state converter is that the sample set can be compressed and the path can be shared by multiple texts, which is convenient for decoding.
[0121] In some embodiments, optionally, the sample set includes a second text and / or a third text; wherein the second text is a partial text extracted from a historical audio file, and the historical audio file also includes a fourth text, the fourth text is a text located after the second text, and the time interval between the second text and the fourth text is greater than or equal to a second duration; wherein the same interactive session includes a first sub-audio file and a second sub-audio file with the same prefix, and the third text is a text extracted from the first sub-audio file.
[0122] In this embodiment, the sample set includes three possible situations, that is, the sample set includes only the second text, the sample set includes only the third text, or the sample set includes both the second text and the third text.
[0123] In some embodiments, optionally, the historical audio file, that is, the audio file collected historically, is obtained by acquiring the audio file collected historically so as to analyze the audio file collected historically. Specifically, when performing speech recognition on the historical audio file, the recognized text and the timestamp corresponding to the recognized text are obtained. If it is detected that the time interval between the text with the front timestamp (that is, the second text in this application) and the text with the back timestamp (that is, the fourth text in this application) is greater than or equal to the second duration, it is considered that there is a speaking pause. For the speaking pause, it should not be determined that there is a speech termination point at the pause. Based on this, the text with the front timestamp (that is, the second text in this application) is used as a sample in the sample set. Among them, the second duration can be set according to actual use needs.
[0124] For example, the user controls the air conditioner to set the timer for three hours. At this time, a voice command for setting the timer for three hours will be issued, that is, "Set_time_____three hours", where each "_" represents a time interval of 100ms. The time interval between set and time is 100ms, and the time interval between time and three is 500ms. When the value of the second time length is 300ms, "scheduled" is used as the second text and as the sample in the sample set.
[0125] Among them, the interactive session can be understood as a voice conversation between a user and a device. For example, the interactive session is a session generated during a period of time before the user sends a voice command to the device and after the device feeds back information to the user.
[0126] In some embodiments, optionally, the second sub audio file is located after the first sub audio file, the first sub audio file is a sub audio file not recognized by the natural language model, and the second sub audio file is a sub audio file recognized by the natural language model.
[0127] In this embodiment, the first sub-audio file and the second sub-audio file can be understood as voice commands repeatedly sent by the user to the device in an interactive session. When the first sub-audio file is a sub-audio file not recognized by the natural language model, and the second sub-audio file is a sub-audio file recognized by the natural language model, it is considered that the first sub-audio file may be incorrectly determined to have a voice termination point. Therefore, the text extracted from the first sub-audio file, that is, the third text, is used as a sample in the sample set.
[0128] For example, the first sub-audio file is "Scheduled time _____ three hours", and the second sub-audio file is "Scheduled time _____ three hours", where "Scheduled time _____ three hours" is not recognized by the natural language model, while "Scheduled time _____ three hours" is recognized by the natural language model, then "Scheduled time _____ three hours" is used as the sample in the sample set.
[0129] In this embodiment, the sample set can cover common possible situations, so that the constructed first weighted finite state converter can meet the discrimination needs of most scenarios, thereby improving the accuracy of voice activation detection.
[0130] In some embodiments, optionally, the sample set can also be expanded using a large language model and manually expanded based on the second text and the third text.
[0131] In this embodiment, the sample set can include more samples, and the first weighted finite state converter constructed based on the sample set can meet the determination needs in more scenarios, thereby improving the accuracy of voice activation detection.
[0132] In some embodiments, optionally, a deterministic operation algorithm and / or a minimization operation algorithm is used to optimize the weighted finite state converter.
[0133] In this embodiment, the use of a determinization algorithm and a minimization algorithm can use the least number of states to express the weighted finite state switch. In this process, the least number of state entries makes the weighted finite state switch more compact.
[0134] In some embodiments, optionally, determining the first duration based on the probability value of the first text in at least one target decoding result specifically includes: determining the maximum probability value among the probability values of at least one first text; obtaining the maximum delay duration and the minimum delay duration; determining the first duration based on the maximum probability value, the maximum delay duration and the minimum delay duration.
[0135] In the above embodiment, the maximum probability value among the probability values of at least one first text is determined so as to select the first text with the highest possibility for output, and the first duration is determined based on the probability value of the first text with the highest possibility.
[0136] In this process, the first duration can be dynamically determined according to the probability value of the first text to improve the accuracy of the speech termination point detected by the stop response.
[0137] In some embodiments, optionally, the first duration is calculated using the following formula:
[0138] ydt = exp(max(log_p)) × (max_dt - min_dt) + min_dt
[0139] Among them, ydt represents the first duration, log_p represents the probability value of at least one first text, max(log_p) represents the maximum probability value among the probability values of at least one first text, max_dt represents the maximum delay duration, min_dt represents the minimum delay duration, and exp() is the exponential function with the natural constant e as the base.
[0140] Exemplarily, the probability value of the first text is represented by a logarithmic probability score. In the case where the audio file is "Turn up the temperature", the target decoding result is shown in Table 1:
[0141] Table 1
[0142] Serial number Decoding results fraction log_p Is it a terminated state? 0 Turn up the temperature -0.000575 TRUE 1 Turn up the heat -8.280709 FALSE 2 Raise the temperature -10.467508 TRUE 3 High temperature -10.491906 FALSE 4 Turn up the temperature -10.617484 FALSE
[0143] Then ydt = exp(max(-0.000575, -10.467508)) × (max_dt - min_dt) + min_dt.
[0144] Among them, the maximum delay duration and the minimum delay duration are preset parameters, and their specific values are not elaborated here.
[0145] In one of the embodiments, in the method for the response voice termination point, if the state corresponding to any one of the first texts in one or more target decoding results in the first weighted finite state transducer is a termination state, a semantic non-stop signal is generated, where the semantic non-stop signal is used to generate the first duration.
[0146] Such as Figure 3 shown, when the first text in the target decoding result is "Adjust the air conditioner's wind speed to the maximum", the starting point of the voice is recognized at "Adjust", and since the semantic non-stop signal is generated, T1 or T2 becomes ineffective.
[0147] Among them, T1, that is, NNVAD Endpoint, is the voice termination point given by the neural network-based speech recognition model, and T2, that is, ASR Endpoint, is the voice termination point given by the automatic speech recognition model. The voice feature is Fbank, where Fbank is a commonly used voice feature in ASR. Among them, NNVAD refers to neural network-based voice activity detection (VAD).
[0148] In one of the embodiments, such as Figure 4As shown, the present invention provides a device 400 for responding to a speech termination point, including: a processing unit 402, used to input a collected audio file into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, the target decoding result including a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, the first text being a text obtained by the speech recognition model performing speech recognition on the audio file; a determination unit 404, used to determine a first duration based on the state corresponding to the first text in the first weighted finite state converter in one or more target decoding results being a termination state, according to the probability value of the first text in at least one target decoding result; a response unit 406, used to cancel the response to the detected speech termination point within a first duration after a first moment, the first moment being the moment of collecting the audio file.
[0149] The present invention proposes a device 400 for responding to a speech end point, which can determine a first duration according to a collected audio file, and cancel the response to the detected speech end point if a speech end point is detected within the first duration after the collection time of the audio file. In this embodiment, within the first duration, by canceling the response to the detected speech end point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related embodiments that the audio file detected by the voice activation detection is incomplete due to the presence of a speaking pause.
[0150] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0151] The audio file may be understood as a file obtained by collecting the user's speaking voice.
[0152] In one embodiment, optionally, the speech recognition model can be an automatic speech recognition (ASR) model, wherein the goal of the automatic speech recognition model is to convert the vocabulary content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0153] In one of the embodiments, optionally, the target decoding result is a result obtained by a speech recognition model decoding an audio file.
[0154] Exemplarily, when a user says “I want to watch TV”, the collected audio file includes the sound of “I want to watch TV”. At this time, the result obtained by the speech recognition model after decoding the audio file is “I want to watch TV”.
[0155] In one embodiment, optionally, the constructed first weighted finite state converter includes a starting node and an ending node, and there are multiple nodes between the starting node and the ending node. The multiple nodes are connected to the starting node and the ending node. The starting node and the multiple nodes, the multiple nodes and the ending node, and the multiple nodes can be connected by text, and each node has a corresponding state attribute.
[0156] Based on this, the state corresponding to the first text in the first weighted finite state converter can be understood as whether the first text is a path including a start node and an end node in the first weighted finite state converter.
[0157] In the above embodiment, judging whether the state corresponding to the first text in the first weighted finite state converter is a termination state can indirectly judge whether the first text is a text that needs to determine a first duration and stops responding to the detected speech termination point within the first duration after the first moment. In this process, the accuracy of stopping responding to the detected speech termination point can be improved.
[0158] In the above embodiment, the speech recognition model has the problem of recognition accuracy. Therefore, when performing speech recognition on the collected audio file, there is a situation where the output first text is inaccurate. Based on this, the probability value of the first text is used to represent the possibility of recognizing the first text from the collected audio file.
[0159] In some embodiments, the speech ending point corresponds to the speech starting point, wherein the speech ending point refers to the moment when the speech disappears when the audio file is subjected to speech activation detection. Similarly, the speech starting point can be understood as the moment when the speech appears when the audio file is subjected to speech activation detection.
[0160] In some embodiments, the acquisition time may be understood as the time when the audio file is acquired.
[0161] In some embodiments, optionally, the termination state is a state when the first text is connected to the termination node.
[0162] In some embodiments, optionally, the speech termination point includes: a speech termination point given by a speech recognition model based on a neural network; and / or a speech termination point given by an automatic speech recognition model.
[0163] Among them, the speech endpoint given by the neural network-based speech recognition model is also known as the NNVAD Endpoint, and the speech endpoint given by the automatic speech recognition model is also known as the ASR Endpoint.
[0164] In this embodiment, the speech ending points given by different models may all be stopped from responding, thereby improving the accuracy of the speech ending points detected by the stop response.
[0165] In some embodiments, optionally, the processing unit 402 is specifically used to: input the collected audio file into the speech recognition model to extract at least one candidate text output by the first conversion layer in the speech recognition model, and the first conversion layer is used to convert the audio file into the candidate text; decode each candidate text separately to obtain at least one target decoding result.
[0166] In this embodiment, the speech recognition model includes a first conversion layer, wherein the first conversion layer is used to convert an audio file into a text to be selected. After the collected audio file is input into the speech recognition model, by extracting the content output by the first conversion layer, the text recognized from the audio file, that is, at least one text to be selected, can be obtained. In this process, the collected audio file can be converted into text. By decoding each text to be selected, a target decoding result can be selected from each text to be selected.
[0167] In some embodiments, optionally, prefix beam search decoding is used to decode each candidate text separately to obtain at least one target decoding result.
[0168] Among them, prefix beam search decoding is a method based on beam search. It searches all possible output sequences and uses a scoring function to evaluate the possibility of each sequence. During the search process, the algorithm retains the scores of the prefixes and selects the most likely output sequence based on these scores. Among them, the score of the prefix is also the probability value of the first text in this application.
[0169] Among them, the main advantage of the prefix beam search algorithm is that it can process input and output sequences of different lengths and can dynamically adjust the size of the search space during the search process, which makes it more efficient than other search algorithms when processing long sequences.
[0170] In this process, the use of prefix beam search decoding can output the most likely candidate text among at least one candidate text as a target decoding result, so as to improve the accuracy and recognition efficiency of the speech recognition model.
[0171] In one of the embodiments, the text to be selected can be understood as the text that may be output from the first conversion layer when the speech recognition model recognizes the collected audio file.
[0172] In one embodiment, the first conversion layer is a Connectionist Temporal Classification layer, ie, a CTC layer.
[0173] In some embodiments, optionally, the determination unit 404 is specifically used to: obtain a first weighted finite state converter, the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the sample in the sample set is a text obtained when an error response speech termination point is given; compare the first state attribute with the target state attribute representing the termination state in the first weighted finite state converter, the first state attribute is related to the state corresponding to the first text in the first weighted finite state converter in at least one target decoding result; based on the first state attribute being the same as the target state attribute, determine that the state corresponding to the first text in the first weighted finite state converter is a termination state; determine the first duration according to the probability value of the first text in at least one target decoding result.
[0174] In this embodiment, the samples in the sample set are texts obtained when the speech end point is incorrectly responded to. It can be understood that the sample set is a set of samples that are incorrectly identified as having a speech end point during the speech activation detection process. In other words, each sample in the sample set should not be determined to have a speech end point, but is incorrectly determined to have a speech end point during the speech activation detection process.
[0175] Based on this, the first weighted finite state converter constructed using the sample set can represent the possible situation where the speech termination point is incorrectly determined to occur. Based on this, after the target decoding result is identified from the audio file, the state corresponding to the first text in the target decoding result in the first weighted finite state converter is compared with the first weighted finite state converter to determine whether the first text is a possible situation where the speech termination point is incorrectly determined to occur.
[0176] Specifically, as can be seen from the above, each node in the first weighted finite state converter has a state attribute. Therefore, after determining the first text, the state corresponding to the first text in the first weighted finite state converter can be determined, that is, the state attribute of the corresponding node can be determined. If the first state attribute is the same as the target state attribute representing the termination state in the first weighted finite state converter, it is considered that the first text is a complete path in the first weighted finite state converter, that is, the first text is a case where a speech termination point is erroneously determined to occur. At this time, the first duration is determined according to the probability value of the first text so as to stop responding to the detected speech termination point within the first duration.
[0177] In this process, by stopping the response to the detected speech end point, the speech activation detection can be continued, thereby delaying the time of responding to the speech end point, and finally, the audio file obtained by the speech activation detection contains a longer and more complete speech.
[0178] In addition, in the above embodiment, the first duration is determined based on the probability value of the first text. Compared with using a fixed duration, it is possible to dynamically adjust the speech length of the audio file obtained by voice activation detection and reduce the longer blank time contained in the audio file obtained by voice activation detection.
[0179] In this process, the accuracy of the audio file obtained by voice activation detection can be improved, thereby improving the accuracy and recognition efficiency of the speech recognition model.
[0180] Exemplarily, the state corresponding to the first text in the first weighted finite state converter is a first state attribute, wherein the first state attribute is a first encoding, and the target state attribute is a second encoding. If the first encoding is the same as the second encoding, the state corresponding to the first text in the first weighted finite state converter is a terminal state.
[0181] Specifically, the second code is 13. If the first state attribute is the first code, and the first code is also 13, then the state corresponding to the first text in the first weighted finite state converter is the terminal state.
[0182] In some embodiments, optionally, the sample set includes a second text and / or a third text; wherein the second text is a partial text extracted from a historical audio file, and the historical audio file also includes a fourth text, and the fourth text is a text located after the second text, and the time interval between the second text and the fourth text is greater than or equal to the second duration; wherein the same interactive session includes a first sub-audio file and a second sub-audio file with the same prefix, and the third text is a text extracted from the first sub-audio file. In this embodiment, the sample set includes three possible situations, that is, the sample set includes only the second text, the sample set includes only the third text, or the sample set includes the second text and the third text.
[0183] In some embodiments, optionally, the historical audio file, that is, the audio file collected historically, is obtained by acquiring the audio file collected historically so as to analyze the audio file collected historically. Specifically, when performing speech recognition on the historical audio file, the recognized text and the timestamp corresponding to the recognized text are obtained. If it is detected that the time interval between the text with the front timestamp (that is, the second text in this application) and the text with the back timestamp (that is, the fourth text in this application) is greater than or equal to the second duration, it is considered that there is a speaking pause. For the speaking pause, it should not be determined that there is a speech termination point at the pause. Based on this, the text with the front timestamp (that is, the second text in this application) is used as a sample in the sample set. Among them, the second duration can be set according to actual use needs.
[0184] For example, the user controls the air conditioner to set the timer for three hours. At this time, a voice command for setting the timer for three hours will be issued, that is, "Set_time_____three hours", where each "_" represents a time interval of 100ms. The time interval between set and time is 100ms, and the time interval between time and three is 500ms. When the value of the second time length is 300ms, "scheduled" is used as the second text and as the sample in the sample set.
[0185] Among them, the interactive session can be understood as a voice conversation between a user and a device. For example, the interactive session is a session generated during a period of time before the user sends a voice command to the device and after the device feeds back information to the user.
[0186] In some embodiments, optionally, the second sub audio file is located after the first sub audio file, the first sub audio file is a sub audio file not recognized by the natural language model, and the second sub audio file is a sub audio file recognized by the natural language model.
[0187] In this embodiment, the first sub-audio file and the second sub-audio file can be understood as voice commands repeatedly sent by the user to the device in an interactive session. When the first sub-audio file is a sub-audio file not recognized by the natural language model, and the second sub-audio file is a sub-audio file recognized by the natural language model, it is considered that the first sub-audio file may be incorrectly determined to have a voice termination point. Therefore, the text extracted from the first sub-audio file, that is, the third text, is used as a sample in the sample set.
[0188] For example, the first sub-audio file is "Scheduled time _____ three hours", and the second sub-audio file is "Scheduled time _____ three hours", where "Scheduled time _____ three hours" is not recognized by the natural language model, while "Scheduled time _____ three hours" is recognized by the natural language model, then "Scheduled time _____ three hours" is used as the sample in the sample set.
[0189] In this embodiment, the sample set can cover common possible situations, so that the constructed first weighted finite state converter can meet the discrimination needs of most scenarios, thereby improving the accuracy of voice activation detection.
[0190] In some embodiments, optionally, the sample set can also be expanded using a large language model and manually expanded based on the second text and the third text.
[0191] In this embodiment, the sample set can include more samples, and the first weighted finite state converter constructed based on the sample set can meet the determination needs in more scenarios, thereby improving the accuracy of voice activation detection.
[0192] In some embodiments, optionally, a deterministic operation algorithm and / or a minimization operation algorithm is used to optimize the weighted finite state converter.
[0193] In this embodiment, the use of a determinization algorithm and a minimization algorithm can use the least number of states to express the weighted finite state switch. In this process, the least number of state entries makes the weighted finite state switch more compact.
[0194] In some embodiments, optionally, the determination unit 404 is specifically used to: determine the maximum probability value among the probability values of at least one first text; obtain the maximum delay duration and the minimum delay duration; and determine the first duration based on the maximum probability value, the maximum delay duration and the minimum delay duration.
[0195] In the above embodiment, the maximum probability value among the probability values of at least one first text is determined so as to select the first text with the highest possibility for output, and the first duration is determined based on the probability value of the first text with the highest possibility.
[0196] In this process, the first duration can be dynamically determined according to the probability value of the first text to improve the accuracy of the speech termination point detected by the stop response.
[0197] In some embodiments, optionally, the first duration is calculated using the following formula:
[0198] ydt=exp(max(log_p))×(max_dt-min_dt)+min_dt
[0199] Among them, ydt represents the first duration, log_p represents the probability value of at least one first text, max(log_p) represents the maximum probability value among the probability values of at least one first text, max_dt represents the maximum delay duration, min_dt represents the minimum delay duration, and exp() is an exponential function with the natural constant e as the base.
[0200] In one embodiment, if Figure 5 As shown, the present invention provides another device 500 for responding to a speech termination point, including a processor 502 and a memory 504, wherein the memory 504 stores programs or instructions that can be executed on the processor 502, and when the programs or instructions are executed by the processor, the steps of any of the above methods are implemented.
[0201] In this embodiment, the first duration can be determined according to the collected audio file, and within the first duration after the collection time of the audio file, if a voice end point is detected, the response to the detected voice end point is canceled. In this technical solution, within the first duration, by canceling the response to the detected voice end point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related technical solution that the audio file detected by the voice activation detection is incomplete due to the presence of speech pauses.
[0202] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0203] Among them, the memory can be used to store software programs and various data. The memory mainly includes a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area can store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory can include a volatile memory or a non-volatile memory, or the memory can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0204] In some embodiments, optionally, the present invention provides a readable storage medium storing a program or instruction, which, when executed by a processor, implements the steps of any of the above-mentioned methods for responding to a voice termination point.
[0205] In this embodiment, the first duration can be determined according to the collected audio file, and within the first duration after the collection time of the audio file, if a voice end point is detected, the response to the detected voice end point is canceled. In this technical solution, within the first duration, by canceling the response to the detected voice end point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related technical solution that the audio file detected by the voice activation detection is incomplete due to the presence of speech pauses.
[0206] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0207] In some embodiments, optionally, a computer program product is provided, which is stored in a storage medium, and when the computer program product is executed by at least one processor, the steps of the method for responding to a speech termination point as described above are implemented.
[0208] In some embodiments, optionally, the present invention provides a voice interaction system, comprising: a device for responding to a voice termination point as described in any one of the above items; and / or a readable storage medium as described above; and / or a computer program product as described above.
[0209] In this embodiment, the voice interaction system can determine the first duration according to the collected audio file, and within the first duration after the collection time of the audio file, if a voice end point is detected, the response to the detected voice end point is canceled. In this technical solution, within the first duration, by canceling the response to the detected voice end point, the voice activation detection can detect a more complete audio file, thereby improving the problem in the related technical solution that the audio file detected by the voice activation detection is incomplete due to the presence of speaking pauses.
[0210] When a more complete audio file is detected, the user's intention can be understood more accurately, thereby improving the accuracy of voice interaction control.
[0211] In some embodiments, optionally, the voice interaction system can be deployed in a user device, wherein the user device can be an air conditioner. In this case, the user can use the voice interaction system to voice control the air conditioner to adjust operating parameters, such as adjusting the cooling temperature, heating temperature, adjusting the wind speed, turning on and off user-selected functions, and switching modes between cooling, heating, dehumidification and air supply.
[0212] In some embodiments, optionally, the user device may be a humidifier, and the user may use the voice interaction system to voice control the humidifier to adjust operating parameters, such as adjusting the humidity setting value or turning off the humidifier.
[0213] In some embodiments, optionally, the user device may be an air purifier, and the user may use the voice interaction system to voice control the air purifier to adjust operating parameters, such as adjusting the purification mode of the air purifier.
[0214] In some embodiments, optionally, the user device may be a sweeping robot, and the user may use the voice interaction system to voice-control the sweeping robot to perform sweeping functions, cleansing or other selected functions.
[0215] The term "first" or "second" in the specification and claims of the present application may include one or more of the features explicitly or implicitly. In the textual description of the present invention, unless otherwise specified, "plurality" means two or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the objects connected before and after are in an "or" relationship.
[0216] In the text description of the present invention, it is understood that, except for explicit provisions and limitations, the terms "installation", "connection" and "connection" should be understood in a broad sense. For example, it can be fixed connection, detachable connection, or integral connection; it can be mechanical structure connection or electrical connection; it can be direct connection between the two, or indirect connection between the two through an intermediate medium, or it can be internal communication between two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0217] In the claims, specification and drawings of the present invention, the description of the terms "one embodiment", "some embodiments", "specific embodiments" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In the claims, specification and drawings of the present invention, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0218] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for responding to a speech termination point, characterized in that: include: Inputting the collected audio file into a speech recognition model to obtain at least one target decoding result output by the speech recognition model, wherein the target decoding result includes a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, wherein the first text is a text obtained by the speech recognition model performing speech recognition on the audio file; Based on the state corresponding to the first text in one or more of the target decoding results in the first weighted finite state converter being a terminal state, determining a first duration according to a probability value of the first text in the at least one target decoding result; canceling the response to the detected speech end point within the first time period after the first moment, the first moment being the acquisition moment of the audio file; The state corresponding to the first text in the first weighted finite state converter based on the one or more target decoding results is a terminal state, and the first duration is determined according to the probability value of the first text in the at least one target decoding result, specifically comprising: Acquire the first weighted finite state converter, where the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the samples in the sample set are texts obtained when a speech termination point is incorrectly responded to; Comparing a first state attribute with a target state attribute representing a terminal state in the first weighted finite state converter, the first state attribute being related to a state corresponding to the first text in the at least one target decoding result in the first weighted finite state converter; Based on the first state attribute being the same as the target state attribute, determining that the state corresponding to the first text in the first weighted finite state converter is a terminal state; A first duration is determined according to a probability value of the first text in the at least one target decoding result.
2. The method for responding to a speech termination point according to claim 1, characterized in that: The voice termination point includes: The speech termination point given by a speech recognition model based on a neural network; and / or the speech termination point given by an automatic speech recognition model.
3. The method for responding to a speech termination point according to claim 1, characterized in that: The step of inputting the collected audio file into the speech recognition model to obtain at least one target decoding result output by the speech recognition model specifically includes: Inputting the collected audio file into a speech recognition model to extract at least one text to be selected output by a first conversion layer in the speech recognition model, wherein the first conversion layer is used to convert the audio file into the text to be selected; Each candidate text is decoded separately to obtain at least one target decoding result.
4. The method for responding to a speech termination point according to claim 3, characterized in that: Prefix beam search decoding is used to decode each candidate text respectively to obtain at least one target decoding result.
5. The method for responding to a speech termination point according to claim 1, characterized in that: The sample set includes a second text and / or a third text; The second text is a partial text extracted from a historical audio file, and the historical audio file further includes a fourth text, the fourth text is a text located after the second text, and the time interval between the second text and the fourth text is greater than or equal to a second duration; The same interactive session includes a first sub-audio file and a second sub-audio file having the same prefix, and the third text is text extracted from the first sub-audio file.
6. The method for responding to a speech termination point according to claim 5, characterized in that: The second sub audio file is located after the first sub audio file, the first sub audio file is a sub audio file that is not recognized by the natural language model, and the second sub audio file is a sub audio file that is recognized by the natural language model.
7. The method for responding to a speech termination point according to claim 1, characterized in that: The weighted finite state converter is optimized by using a deterministic operation algorithm and a minimization operation algorithm.
8. The method for responding to a speech termination point according to claim 1, characterized in that: The determining the first duration according to the probability value of the first text in the at least one target decoding result specifically includes: Determining a maximum probability value among the probability values of at least one of the first texts; Get the maximum delay time and the minimum delay time; The first duration is determined based on the maximum probability value, the maximum delay duration, and the minimum delay duration.
9. A device for responding to a speech termination point, characterized in that: include: a processing unit, configured to input the collected audio file into a speech recognition model, and obtain at least one target decoding result output by the speech recognition model, wherein the target decoding result includes a first text, a probability value of the first text, and a state corresponding to the first text in a first weighted finite state converter, wherein the first text is a text obtained by the speech recognition model performing speech recognition on the audio file; A determining unit, configured to determine a first duration according to a probability value of the first text in at least one target decoding result, based on the state corresponding to the first text in the first weighted finite state converter being a terminal state. A response unit, configured to cancel a response to a detected speech end point within the first time period after a first moment, wherein the first moment is a time at which the audio file is collected; The determination unit is also used to: obtain a first weighted finite state converter, the first weighted finite state converter is a weighted finite state converter constructed using a sample set, and the samples in the sample set are texts obtained when an error response speech termination point is given; compare a first state attribute with a target state attribute representing a termination state in the first weighted finite state converter, the first state attribute being related to a state corresponding to the first text in the first weighted finite state converter in at least one target decoding result; determine that the state corresponding to the first text in the first weighted finite state converter is a termination state based on the first state attribute being the same as the target state attribute; and determine a first duration according to a probability value of the first text in at least one target decoding result.
10. A device for responding to a speech termination point, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method for responding to a speech termination point according to any one of claims 1 to 8 are implemented.
11. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the method for responding to a speech termination point according to any one of claims 1 to 8 are implemented.
12. A computer program product, the computer program product being stored in a storage medium, characterized in that: When the computer program product is executed by at least one processor, the steps of the method for responding to a speech termination point according to any one of claims 1 to 8 are implemented.
13. A voice interaction system, characterized in that: include: The device for responding to speech termination points as claimed in claim 9 or 10; and / or The readable storage medium according to claim 11; and / or The computer program product of claim 12.
Citation Information
Patent Citations
Speech recognition method and device, storage medium and electronic equipment
CN112634876A
Speech recognition method, related device, electronic equipment and storage medium
CN115798480A