A voice interaction method, system, device and medium for adjusting the broadcast voice with confidence
By introducing a confidence adjustment mechanism in the voice interaction system, the poor user experience caused by voice recognition errors is solved. Through confidence judgment and broadcast library selection, the system's response accuracy and humanized interaction are improved.
Patent Information
- Application Number
- CN202111658705.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-12-30
AI Technical Summary
When the artificial intelligence voice recognition error occurs, the reply text and instructions do not match the user's expectations, resulting in a worse user experience. When the recognition fails, the wake-up words need to be repeated, affecting the user experience.
The first voice recognition is performed by receiving voice commands, and the confidence is output, and the confidence is judged based on the confidence level or the broadcast sound is selected from the broadcast sound library. If the confidence level is low, the second recognition or the user is consulted until the confidence level meets the threshold.
Improve user experience, adjust the broadcast sound through confidence, reduce the problem of inconsistency caused by speech recognition errors, and enhance the system's response accuracy and humanized interaction.
Smart Images

Figure CN114333822B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent voice interaction, and in particular to a voice interaction method, system, device and medium for adjusting the broadcast voice with confidence. Background Art
[0002] Intelligent voice interaction is a new generation of interaction mode based on voice input, and feedback results can be obtained by speaking. Common voice interaction devices include intelligent speakers, voice dialogue intelligent household appliances, in-vehicle intelligent rearview mirrors, intelligent customer service, dialogue robots, etc. The voice recognition and dialogue system of artificial intelligence technology is the basis for the realization of the functions of intelligent speaker products. The general working process is as follows: sound collection, noise reduction, voice wake-up, speech-to-text conversion, semantic understanding, reply text and instructions, text-to-speech conversion, play sound, use speech recognition for speech-to-text conversion, select the most likely text or semantics, and then perform semantic understanding, and then perform reply text and instructions and play the sound of reply text and instructions. Since speech recognition is a process with a certain error rate, sound collection will naturally be affected by external noise, and the signal-to-noise ratio of the received message is a finite value. And the speech recognition technology of artificial intelligence also has a certain error rate. At present, the correct rate of offline artificial intelligence speech recognition is generally about 80% - 95%, and the correct rate using cloud technology may be higher, but it is not 100%. When the speech recognition of artificial intelligence in the existing voice interaction device fails to recognize correctly, the reply text and instructions and the conversion to broadcast voice for broadcasting do not meet the expectations of the user, causing discomfort to the user or a deterioration in the user experience; sometimes, when the speech recognition of artificial intelligence fails to recognize successfully for a period of time, the device will, based on the timeout setting, exit the speech recognition state, and the user may need to say the wake-up word again to recognize again, which will also cause a deterioration in the user experience. Summary of the Invention
[0003] Aiming at the defects in the prior art, the purpose of the present invention is to provide a voice interaction method, system, device and medium for adjusting the broadcast voice with confidence.
[0004] A voice interaction method for adjusting the broadcast voice with confidence provided by the present invention includes the following steps:
[0005] Step S1, receiving an external voice command;
[0006] Step S2, performing first speech recognition on the received voice command according to features and outputting a recognition result and corresponding confidence;
[0007] Step S3, making a judgment based on the recognition result and the corresponding confidence, whether to output a corresponding action, and whether to select a broadcast voice from the broadcast voice library for broadcasting;
[0008] Step S4: Determine whether to continue with the second voice recognition, and output the recognition result and the corresponding confidence level. If it is decided to continue with the second voice recognition, execute Steps S2 to S4; otherwise, return to Step S1, wait to be awakened, and proceed to the next operation until the corresponding result is output.
[0009] Among them, the confidence level is used to judge the correct probability of the semantics corresponding to the voice received by the system.
[0010] Furthermore, in Step S2,
[0011] Judge the confidence level between the received voice command and the recognized semantics. When the target confidence level is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence level is lower than the preset threshold, the system selects a voice announcement of the inquiry type from the voice announcement library and conducts a voice interaction with the user.
[0012] Furthermore, in Step S2,
[0013] The types of the target confidence level include: when the target confidence level is higher than the preset threshold, one type of result is output as non-human voice, and the other type of result is output as a single recognized semantics;
[0014] when the target confidence level is lower than the preset threshold, one type of result is output as a single recognized semantics, and the other type of result is output as multiple recognition results.
[0015] Furthermore, in Step S3,
[0016] Step S31: If the system identifies a single recognition result and the confidence level of the single recognition result is greater than or equal to the preset confidence threshold, the voice announcement directly executes the predetermined instruction action or outputs a short prompt tone with an affirmative tone and performs the instruction action; and / or,
[0017] Step S32: If the system identifies a single recognition result and the confidence level of the single recognition result is less than the preset confidence threshold, the voice announcement outputs a prompt tone with a questioning tone and waits for the user's next instruction; and / or,
[0018] Step S33: If the system identifies multiple recognition results and the confidence levels of the multiple recognition results are less than the preset confidence threshold, the voice announcement outputs a prompt tone with a questioning tone for the user to make an instruction selection; and / or,
[0019] Step S34: If the system identifies that the recognition result is a human voice and the confidence level of the human voice is less than the preset confidence threshold, the voice announcement outputs a prompt tone with a questioning tone and waits for the user to repeat the previous instruction; and / or,
[0020] In step S35, if the recognition result identified by the system is non-human voice and the confidence level of the non-human voice is greater than or equal to the preset confidence threshold, the voice output plays a prompt tone to ask whether service is required, or performs a timed service, emits a prompt tone, and the system enters the state before waking up.
[0021] Further, in the said step S2,
[0022] Obtain the target voice command and perform noise reduction processing on the target voice;
[0023] Perform semantic understanding on the target voice command through speech recognition.
[0024] A voice interaction system that adjusts the voice output according to the confidence level provided by the present invention includes:
[0025] A voice receiving module, which receives external voice commands;
[0026] A feature recognition module, which performs the first speech recognition on the received voice command according to the features and outputs the recognition result and the corresponding confidence level;
[0027] A confidence level judgment and output module, which makes a judgment based on the recognition result and the corresponding confidence level, whether to output the corresponding action, and whether to select a voice output from the voice output library for broadcasting;
[0028] A loop recognition module, which judges whether to continue the second speech recognition. If so, continue the second speech recognition, and execute from the feature recognition module to the loop recognition module; otherwise, return to the voice receiving module, wait to be woken up, and enter the next module until the corresponding result is output;
[0029] Wherein, the confidence level is used to judge the correct probability of the semantics corresponding to the voice received by the system.
[0030] Further, in the said feature recognition module,
[0031] Judge the confidence level between the received voice command and the recognized semantics. When the target confidence level is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence level is lower than the preset threshold, the system selects an inquiry type voice output from the voice output library to interact with the user through voice.
[0032] Further, the said target confidence level types include: when the target confidence level is higher than the preset threshold, one type of result is output as non-human voice, and the other type of result is output as a single recognized semantics;
[0033] When the target confidence level is lower than the preset threshold, one type of result is output as a single recognized semantics, and the other type of result is output as multiple recognition results.
[0034] A voice interaction device that adjusts the broadcast voice according to the confidence level provided by the present invention includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the voice interaction method for adjusting the broadcast voice as described in this embodiment is implemented.
[0035] A computer-readable medium provided by the present invention stores a computer program. When the computer program is executed by one or more processors, the voice interaction method for adjusting the broadcast voice as described in this embodiment is implemented. Due to the above technical solutions, the present invention has the following advantages and positive effects compared with the prior art: In the voice recognition system of artificial intelligence, through the output of the identified neural network or expert system, not only the recognition result is output, but also the confidence level of this recognition result is output, or multiple recognition results and their corresponding confidence levels are output. Based on the high and low of these confidence levels, the subsequent reply text and the content of the instruction, and the conversion to the broadcast voice to broadcast the above steps are further processed to improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic flow chart of a voice interaction method for adjusting the broadcast voice according to the present invention;
[0037] Figure 2 It is a schematic overall flow chart of a voice interaction method for adjusting the broadcast voice according to the present invention;
[0038] Figure 3 It is a schematic specific flow chart of step S3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] As Figure 1 shown, a voice interaction method for adjusting the broadcast voice provided by this embodiment includes the following steps:
[0042] Step S1, receiving an external voice command;
[0043] Step S2, performing first voice recognition on the received voice command according to the features and outputting the recognition result and the corresponding confidence level;
[0044] Step S3: Based on the recognition result and the corresponding confidence level, determine whether to output the corresponding action and whether to select a broadcast voice from the broadcast audio library for broadcasting.
[0045] Step S4: Determine whether to continue with the second voice recognition and output the recognition result and the corresponding confidence level. If continuing with the second voice recognition, execute Steps S2 to S4; otherwise, return to Step S1, wait to be awakened, and proceed to the next operation until the corresponding result is output.
[0046] Among them, the confidence level is used to judge the correct probability of the semantics corresponding to the voice received by the system.
[0047] Those skilled in the art can understand that in the voice recognition system of artificial intelligence, through the output of the recognized neural network or expert system, not only the recognition result is output, but also the confidence level (Confidence level) of this recognition result is output, or multiple recognition results and their corresponding confidence levels are output. In this embodiment, the output of the neural network is often the possibility of a certain result, a word, a sentence, which is a combination of the outputs of many characters. The confidence level is the credibility of the result judged by the weighted algorithm or neural network. Based on the high or low of these confidence levels, further processing is performed on the subsequent reply text and the content of the instruction, and the step of converting to a broadcast voice for broadcasting, so as to provide a better user experience. The voice recognition technology of artificial intelligence is based on a neural network or an expert system to judge the received audio signal, and uses the most likely recognition result as the input for the next action to determine the content of the subsequent reply text and instruction, and convert it to a broadcast voice for broadcasting.
[0048] As Figure 2 shown, Step S1 includes: the system waits to be awakened, the user issues a voice command, and the system is awakened.
[0049] Further, in Step S2,
[0050] Judge the high or low confidence level between the received voice command and the recognized semantics. When the target confidence level is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence level is lower than the preset threshold, the system selects a broadcast voice of the inquiry type from the broadcast audio library to interact with the user in voice.
[0051] Further, in Step S2,
[0052] The types of the target confidence level include: when the target confidence level is higher than the preset threshold, one type of result is output as non-human voice, and the other type of result is output as a single recognized semantics.
[0053] When the target confidence level is lower than the preset threshold, one type of result is output as a single recognized semantics, and the other type of result is output as a multi-recognized result.
[0054] Further, the step S2 includes:
[0055] Step S21, obtaining the target voice command and performing noise reduction processing on the target voice;
[0056] Step S22, converting the target voice command into text through voice recognition, selecting the most likely text or semantics, performing semantic understanding, and then replying with text and commands and playing the voice of the replied text and commands.
[0057] Further, the step S2 further includes:
[0058] Step S23, determining whether the received voice command obtains a valid recognition value. If any of the voice commands obtains a valid recognition value, execute the corresponding command and output the corresponding command as the recognition result; if the recognition value is invalid, output the recognized human voice or the recognized external noise as the recognition result.
[0059] Further, the step S2 further includes:
[0060] Step S24, sampling various sound sources, determining the confidence threshold interval of various sound sources, comparing the target confidence level with the preset confidence threshold, and outputting the recognition result and the recognition result with the corresponding confidence level.
[0061] As Figure 3 shown, the step S3 includes:
[0062] Step S31, if the system recognizes a single recognition result and the confidence level of the single recognition result is greater than or equal to the preset confidence threshold, the voice broadcast directly executes the predetermined command action or outputs a short prompt tone with an affirmative tone and performs the command action;
[0063] Those skilled in the art can understand that the recognition result with the highest possibility recognized by the voice recognition system of artificial intelligence is also in a state with a high confidence level. The reply voice broadcast can use an affirmative tone, or even a short reply without repeating the recognized words, and directly execute the predetermined command action. For example, the voice recognition system of artificial intelligence recognizes that the user's voice is "turn on the air conditioner" and has a high confidence in this recognition. The device can directly turn on the air conditioner. Or be accompanied by a simple prompt tone to let the user know that the machine will perform the next operation. Or broadcast a voice broadcast with an affirmative tone, such as: "Okay, master, I'll turn on the air conditioner."
[0064] Step S32, if the system recognizes a single recognition result and the confidence level of the single recognition result is less than the preset confidence threshold, the voice broadcast outputs a prompt tone with a questioning tone and waits for the user's next command;
[0065] Those skilled in the art can understand that the recognition result with the highest possibility recognized by the speech recognition system of artificial intelligence is at the same time in a state with low confidence. The reply broadcast voice can use an interrogative tone to naturally guide the user to say it again. For example, the speech recognition system of artificial intelligence recognizes that the user's speech is most likely "turn on the air conditioner", but the confidence in this recognition is not high. The device can broadcast a broadcast voice with an interrogative tone, such as: "Excuse me, do you want to turn on the air conditioner?"
[0066] Step S33, if the system recognizes multiple recognition results and the confidence levels of the multiple recognition results are less than the preset confidence threshold, the broadcast voice outputs a prompt tone with an interrogative tone for the user to make an instruction selection;
[0067] Those skilled in the art can understand that the speech recognition system of artificial intelligence recognizes multiple results with similar and low confidence levels. The reply broadcast voice can use an interrogative tone for the user to choose. For example, the speech recognition system of artificial intelligence recognizes that the user's speech may be "turn on the air conditioner" or "turn on the console", and the confidence levels of both are not high and the difference is within a certain range. The device can broadcast a broadcast voice with an interrogative tone in the form of a multiple-choice question, such as: "Excuse me, do you want to turn on the air conditioner or turn on the console?"
[0068] Step S34, if the system recognizes that the recognition result is human voice and the confidence level of the human voice is less than the preset confidence threshold, the broadcast voice outputs a prompt tone with an interrogative tone and waits for the user to repeat the previous instruction; Those skilled in the art can understand that the speech recognition system of artificial intelligence recognizes human voice, but the confidence levels of all the recognized results are very low. The reply broadcast voice can use an interrogative tone to remind the user. For example, the speech recognition system of artificial intelligence recognizes that the user should have spoken to the device, but the confidence levels of all the recognition results are very low. The device can broadcast a reminder broadcast voice, such as: "I didn't hear clearly", or "Please speak louder", or "Please say it again".
[0069] Step S35, if the recognition result recognized by the system is non-human voice and the confidence level of the non-human voice is greater than or equal to the preset confidence threshold, the broadcast voice outputs a prompt tone to ask whether service is needed, or, perform timed service, emit a prompt tone, and the system enters the state before waking up.
[0070] Those skilled in the art can understand that for the speech recognition system of artificial intelligence, if it identifies that there is no human voice or has a high confidence in determining that the external environment is noise, the system can be designed in two possible ways. One way is to slightly issue a reminder phrase. For example, "Do you need my service?" or "You can use the following methods to ask me to serve you, such as asking me to turn on or off the air conditioner, or asking me to open or close the console." Another way is to start timing. After a period of time, a pre-ending reminder phrase is issued. For example, "If you don't need my service, I'll go and do something else." If no command word is heard or no human voice that requires a next-step response is detected within a certain period of time after that, the device enters the state before waking up.
[0071] Further, the recognition result includes words or sentences with a unique semantics recognized by the system, combinations of multiple groups of words or sentences with multiple semantics, sound sources such as human voices or non-human voices; the confidence level is used to judge the correct probability of the system recognizing semantics, human voices or non-human voice sound sources.
[0072] In this implementation, the broadcast voice is selected from the broadcast voice library. In addition to one-to-one selection, multiple possible broadcast voices can be pre-stored in specific states and selected in order or randomly, which increases the user-friendly features of the device and avoids user fatigue. Further, step S4 includes:
[0073] If continuous recognition is required, return to step S2, start recognition based on the features, output the recognition result and the corresponding confidence level, make a judgment based on the recognition result and the corresponding confidence level, and select a broadcast voice from the broadcast voice library for broadcasting; otherwise, return to step S1, wait to be woken up and perform the next operation until the result is output.
[0074] Those skilled in the art can understand that for broadcast voices in similar situations, several can be grouped together and selected alternately or randomly for use to make some changes and enhance the user experience.
[0075] Embodiment 2
[0076] A voice interaction system for adjusting broadcast voices based on confidence level provided in this embodiment includes:
[0077] A voice receiving module for receiving external voice commands;
[0078] A feature recognition module for performing the first voice recognition on the received voice command according to the features and outputting the recognition result and the corresponding confidence level;
[0079] A confidence level judgment and output module for making a judgment based on the recognition result and the corresponding confidence level, whether to output corresponding actions, and whether to select a broadcast voice from the broadcast voice library for broadcasting;
[0080] The loop recognition module determines whether to continue the second voice recognition. If the second voice recognition is to be continued, the feature recognition module to the loop recognition module are executed; otherwise, it returns to the voice receiving module, waits to be awakened, enters the next module until the corresponding result is output.
[0081] Among them, the confidence level is used to judge the correct probability of the semantics corresponding to the voice received by the system.
[0082] Furthermore, in the feature recognition module,
[0083] Judge the confidence level between the received voice command and the recognized semantics. When the target confidence level is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence level is lower than the preset threshold, the system selects a broadcast voice of the inquiry type from the broadcast sound library to interact with the user in voice.
[0084] Furthermore, the types of the target confidence level include: when the target confidence level is higher than the preset threshold, one type of result is output as non-human voice, and the other type of result is output as single recognized semantics;
[0085] When the target confidence level is lower than the preset threshold, one type of result is output as single recognized semantics, and the other type of result is output as multiple recognition results.
[0086] The present invention also provides a voice interaction device that adjusts the broadcast voice according to the confidence level, including a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the voice interaction method for adjusting the broadcast voice according to the confidence level as described in this embodiment is implemented.
[0087] The present invention also provides a computer-readable medium storing a computer program. When the computer program is executed by one or more processors, the voice interaction method for adjusting the broadcast voice according to the confidence level as described in this embodiment is implemented.
[0088] If the module in the second embodiment is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of software. This computer software is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0089] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A voice interaction method for adjusting the broadcast voice with confidence, characterized in that, It includes the following steps: Step S1, receiving an external voice command; Step S2, performing the first voice recognition on the received voice command according to features and outputting an identification result and corresponding confidence, including: Step S21, obtaining a target voice command and performing noise reduction processing on the target voice; Step S22, converting the target voice command into text through voice recognition, selecting the most likely text or semantics, performing semantic understanding, generating a reply text and command, and playing the sound of the reply text and command; Step S23, determining whether the received voice command obtains a valid identification value. If any of the voice commands obtains a valid identification value, execute the corresponding command and output the corresponding command as the identification result; if the identification value is invalid, output the recognized human voice or the recognized external noise as the identification result; Step S24, sampling various sound sources, determining the confidence threshold intervals of various sound sources, comparing the target confidence with the preset confidence threshold, and outputting the identification result and the recognition result of the corresponding confidence. The types of the target confidence include: when the target confidence is higher than the preset threshold, output one type of result as non-human voice and output another type of result as a single recognized semantics; when the target confidence is lower than the preset threshold, output one type of result as a single recognized semantics and output another type of result as multiple recognition results; Step S3, making a judgment based on the recognition result and the corresponding confidence, determining whether to output the corresponding action, and whether to select a broadcast sound from the broadcast sound library for broadcasting; Step S4, determining whether to continue the second voice recognition. If so, execute steps S2 to S4; otherwise, return to step S1, wait to be awakened, and perform the next operation until the corresponding result is output; Among them, the confidence is used to judge the correct probability of the semantics corresponding to the voice received by the system.
2. The voice interaction method for adjusting the broadcast voice with confidence as described in claim 1, wherein In step S2, judge the confidence level between the received voice command and the recognized semantics. When the target confidence is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence is lower than the preset threshold, the system selects a broadcast sound of an inquiry type from the broadcast sound library to interact with the user through voice.
3. The voice interaction method for adjusting the broadcast voice with confidence as claimed in claim 1 or 2, wherein In step S3, Step S31, if the system identifies a single identification result and the confidence of the single identification result is greater than or equal to the preset confidence threshold, the broadcast sound directly executes the predetermined command action or outputs a short prompt sound with an affirmative tone and performs the command action; and / or, Step S32, if the system identifies a single identification result and the confidence of the single identification result is less than the preset confidence threshold, the broadcast sound outputs a prompt sound with a questioning tone and waits for the user's next command; and / or, Step S33, if the system identifies multiple identification results and the confidence of the multiple identification results is less than the preset confidence threshold, the broadcast sound outputs a prompt sound with a questioning tone for the user to make a command selection; and / or, Step S34, if the system identifies that the identification result is a human voice and the confidence of the human voice is less than the preset confidence threshold, the broadcast sound outputs a prompt sound with a questioning tone and waits for the user to repeat the previous command; and / or, Step S35: If the recognition result identified by the system is non-human voice and the confidence level of the non-human voice is greater than or equal to the preset confidence threshold, the voice output module outputs a prompt tone, asks whether service is required, or performs a timed service, emits a prompt tone, and the system enters the state before waking up.
4. A voice interaction system for adjusting the broadcast voice with confidence, characterized in that, Including: A voice receiving module that receives external voice commands; A feature recognition module that performs first voice recognition on the received voice command according to features and outputs a recognition result and corresponding confidence level, including: Obtaining a target voice command and performing noise reduction processing on the target voice; Converting the target voice command into text through voice recognition, selecting the most likely text or semantics, performing semantic understanding, generating reply text and commands, and playing the sounds of the reply text and commands; Judging whether the received voice command obtains a valid recognition value. If any of the voice commands obtains a valid recognition value, execute the corresponding command and output the corresponding command as the recognition result; if the recognition value is invalid, output the recognized human voice or the recognized external noise as the recognition result; Sampling various sound sources, determining the confidence threshold interval of various sound sources, comparing the target confidence level with the preset confidence threshold, and outputting the recognition result and the recognition result of the corresponding confidence level. The types of the target confidence level include: when the target confidence level is higher than the preset threshold, output one type of result as non-human voice and output another type of result as a single recognized semantics; when the target confidence level is lower than the preset threshold, output one type of result as a single recognized semantics and output another type of result as multiple recognition results; A confidence level judgment and output module that makes a judgment based on the recognition result and the corresponding confidence level, whether to output the corresponding action, and whether to select a broadcast voice from the broadcast voice library for broadcasting; A loop recognition module that judges whether to continue the second voice recognition. If so, continue the second voice recognition, and execute the feature recognition module to the loop recognition module; otherwise, return to the voice receiving module, wait to be woken up, and enter the next module until the corresponding result is output; Among them, the confidence level is used to judge the correct probability of the semantics corresponding to the voice received by the system.
5. The voice interaction system for adjusting the broadcast voice with confidence as described in claim 4, characterized in that In the feature recognition module, Judge the high and low confidence level between the received voice command and the recognized semantics. When the target confidence level is higher than the preset threshold, the system directly executes the action corresponding to the recognized semantics; when the target confidence level is lower than the preset threshold, the system selects a broadcast voice of the inquiry type from the broadcast voice library and conducts a voice interaction with the user.
6. A voice interaction device that adjusts the broadcast voice with confidence, characterized in that, It includes a memory and a processor. Computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, it implements the voice interaction method for adjusting the broadcast voice with confidence level as described in any one of claims 1 to 3.
7. A computer-readable medium stores a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the voice interaction method for adjusting the broadcast voice with confidence level as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Voice interaction control method and device, electronic equipment, storage medium and system
CN111768783A