Digital human voice session question and answer interaction control and exception handling method
By setting system parameters and status, controlling the digital human voice conversation Q&A interaction process, and taking corresponding measures when an exception occurs, the problem of difficulty in handling abnormalities in the virtual digital human voice Q&A interaction process in the prior art is solved, and the system applicability and interactive experience are improved.
Patent Information
- Application Number
- CN202411993365.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
AI Technical Summary
During the existing virtual digital human voice Q&A interaction, it is easy to cause meaningless links and actions of the system due to abnormalities in speech recognition, answer generation, and voice playback, and it is difficult to deal with abnormalities that are difficult for the system to detect in a timely manner.
By setting system parameters and status, control the digital human voice conversation Q&A interaction process; when an exception occurs, judge the abnormal link and take measures such as voice-level resampling, conversation-level resampling, or voice broadcast start-stop switching, clear the abnormal results and recapture or generate the answer.
Effectively handle abnormalities in the voice conversation Q&A interaction process, avoid meaningless links and actions of the system, and improve the applicability and interactive experience of system applications.
Smart Images

Figure CN119964563A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice conversation technology, and more specifically, to a digital human voice conversation question-answer interaction control and exception handling method. Background Art
[0002] The voice question-and-answer interaction between users and virtual digital humans includes a series of links, including user question voice capture and sampling, user question voice recognition, large model return of answer text, answer text voice synthesis, answer voice broadcast, etc. These links are closely linked. If a problem occurs in the previous link, all the subsequent links are meaningless and the subsequent link process needs to be terminated in time. However, the system application itself is often not sure that an abnormal problem has occurred in a certain link, which requires timely human intervention to avoid meaningless links and actions of the system.
[0003] For example, the speech recognition process is inevitably affected by many factors, such as the language standard (Mandarin, dialect), voice loudness, speech speed, the stability, accuracy and recognition efficiency of the speech recognition service itself, and various anomalies often occur: inaccurate speech recognition results, speech recognition service anomalies (such as inability to access the speech recognition server, etc.). Speech recognition service anomalies will directly lead to the unavailability of the entire system service. If the user's speech recognition results are inaccurate, the subsequent large model's answer feedback and voice broadcast will be meaningless.
[0004] Another example: After the user asks a question, the big model will give the corresponding response feedback. The feedback results may be abnormal: irrelevant answers, inaccurate answers, big model service abnormalities, etc. The user may not be satisfied with the feedback results, and often needs to regenerate the response results one or more times based on the previous questions. Moreover, if the user is not satisfied with the previous feedback results, the subsequent voice broadcast will be meaningless and will appear superfluous.
[0005] In addition, the voice broadcast function is related to factors such as the specific scenario and the user's mentality at the time. Sometimes it is needed, and sometimes it is not. A user interface can be provided to start and stop the voice broadcast function at any time to respond to user scenario needs and increase user interaction experience. Summary of the invention
[0006] According to the present invention, a digital human voice conversation question-and-answer interaction control and exception handling method is provided to solve the technical problems generated in the existing virtual digital human voice question-and-answer interaction process.
[0007] According to the first aspect of the present invention, there is also provided a digital human voice conversation question-answer interaction control and exception handling, comprising:
[0008] Set the system parameters and their states, and control the digital human voice conversation question-and-answer interaction process according to the parameter states;
[0009] If an abnormality occurs during the digital human voice conversation question-and-answer interaction, determine at which stage of the voice recognition, answer generation, and voice playback the abnormality occurs;
[0010] In case of anomalies in speech recognition, voice-level resampling is adopted to clear the current speech recognition results and recapture the user's voice. In case of anomalies in answer generation, session-level resampling is adopted to clear the current answer generation results and re-input the text stream of the user's question into the large model service to generate the answer. In case of anomalies in voice playback, when the user presses the broadcast start / stop button, the voice broadcast is started or paused.
[0011] Optionally, set various system parameters and their states, including:
[0012] Set the system parameters as follows: system-level state SysState, business-level state AiState, and speech broadcast mode SpeechMode;
[0013] The parameter states of the system-level state SysState include: uninitialized INVALID, activated state ACTIVATED, and a question-and-answer session SESSION in a certain round;
[0014] The parameter states of the business-level state AiState include: uninitialized INVALID, capturing CAPTURE, speech recognition ASR, large model generation CHATGPT, speech synthesis, and speech broadcast TTS.
[0015] The parameter states of the speech broadcast mode SpeechMode include: uninitialized INVALID, start speech broadcast SPEECH_ON, and stop speech broadcast SPEECH_OFF.
[0016] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0017] Set the parameter state of the system-level state SysState to a question-and-answer session SESSION in a certain round, and determine the start of a new round of question-and-answer interaction session;
[0018] Capture user voice in real time. When the user starts speaking, send the captured user voice stream to the speech recognition service in real time to convert it into a text stream. Set the parameter state of the business-level state AiState to speech recognition ASR in progress.
[0019] The captured user voice stream is sent to the speech recognition service in real time to be converted into a text stream, and the parameter state of the business-level state AiState is set to speech recognition ASR in progress.
[0020] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0021] If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the question text segments generated during the speech recognition process are serially summarized in order to form a complete user question;
[0022] Push the complete user question to the big model, generate the corresponding text answer, and set the parameter state of the business-level state AiState to CHATGPT, which is generating the answer in the big model.
[0023] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0024] Perform speech synthesis on the text answer, convert the text stream into a speech stream and broadcast it in real time, use the voice stream to drive the AI digital human to perform corresponding actions, and set the parameter status of the business-level state AiState to speech synthesis and speech broadcast TTS.
[0025] Optionally, voice-level resampling is performed for abnormal speech recognition, the current speech recognition result is cleared, and the user's voice is recaptured, including:
[0026] During the speech recognition process, if the user's language standard, voice loudness, speech speed, or speech recognition service itself is wrong, the speech recognition is determined to be abnormal, and the parameter state of the current business-level state AiState is detected to be speech recognition ASR;
[0027] If the parameter state of the current service-level state AiState is in speech recognition ASR, clear the current speech recognition result, recapture the user's voice, send the captured user voice stream to the speech recognition service, perform speech recognition, and obtain the corresponding text result;
[0028] The text results of speech recognition are displayed in real time on the terminal interface.
[0029] Optionally, session-level resampling is performed to solve the problem of abnormal answer generation, clear the current answer generation result, and re-input the text stream of the user question into the large model service to generate the answer, including:
[0030] During the answer generation process, if the feedback answer is irrelevant, inaccurate, or the big model service is abnormal, the answer generation abnormality is determined, and the parameter state of the current business-level state AiState is detected to be the big model generated answer CHATGPT;
[0031] If the parameter state of the current business-level state AiState is the big model generating answer CHATGPT, clear the current big model answer result, re-input the text stream of the user question into the big model service, and the big model service returns the text stream of the corresponding answer result;
[0032] The text information of the response result is displayed in real time on the terminal interface, speech synthesis is performed, and the voice stream is used to drive the digital human's corresponding actions and voice broadcasts.
[0033] Optionally, in case of abnormal voice playback, when the user presses the broadcast start / stop button, the voice broadcast is started or paused, including:
[0034] During speech playback, when the user presses the speech start / stop button, the state of the current speech playback mode SpeechMode is detected;
[0035] If the current speech broadcast mode SpeechMode is equal to SPEECH_ON, it is turned off. The system detects whether the current business-level state AiState is in the state of speech synthesis or speech broadcast TTS. If so, it stops inputting the voice stream to the sound card and stops the speech broadcast function. If it is not in the state of speech synthesis or speech broadcast TTS, a round of conversation ends.
[0036] Optionally, in case of abnormal voice playback, when the user presses the broadcast start / stop button to start or pause the voice broadcast, the method further includes:
[0037] If SpeechMode is equal to SPEECH_ON, set SpeechMode to SPEECH_OFF, change the speech broadcast state from the original on state to the stopped state, and set the broadcast start and stop button to the style icon of the stopped state;
[0038] If SpeechMode is equal to SPEECH_OFF, set SpeechMode to SPEECH_ON, that is, the voice broadcast state is changed from the original off state to the on state, and the broadcast start and stop button is set to the style icon of the on state.
[0039] Optionally, it also includes:
[0040] During the digital human voice conversation question-and-answer interaction process, when the user presses a button to return to the initial state, the current service-level state AiState is detected to see whether it is in the voice recognition ASR, large model answer CHATGPT, speech synthesis or voice broadcast TTS service state;
[0041] If the current business-level state AiState is in the speech recognition ASR, large model answer CHATGPT, speech synthesis or speech broadcast TTS business state, a signal to return to the initial state is sent to the system. After the system captures the signal to return to the initial state, the current process link is immediately ended;
[0042] Returns the initial state of user voice capture and sampling and sets the initial state parameters.
[0043] According to another aspect of the present invention, a digital human voice conversation question-answering interaction method is provided, comprising:
[0044] Wake up the device based on user-specific audio information, enter user consultation mode after starting the service, capture the user's voice in real time, send the user's voice to the speech recognition service, and convert the voice stream of the user's voice into a text stream;
[0045] If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the text stream is summarized and semantically corrected to form a complete user question;
[0046] Push the complete user question to the big model, generate the corresponding text answer, synthesize the text answer and broadcast it in real time;
[0047] Display user questions and text answers on the terminal interface, and drive the AI digital human to take actions based on the real-time voice stream.
[0048] Therefore, for the important links in the voice conversation question-and-answer interaction process, it supports manual operation intervention and provides a humanized interactive response control mechanism. It can promptly handle anomalies that are difficult to detect in the system itself, avoid meaningless links and actions of the system, and respond to user scenario needs in a timely manner, greatly enhancing the applicability of the system application scenarios and the interactive experience of the conversation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0050] Figure 1 A flowchart of a digital human voice conversation question-answering interactive control and exception handling method described in this embodiment;
[0051] Figure 2 A schematic diagram of the flow of a digital human voice conversation question-answering interactive method described in this embodiment;
[0052] Figure 3 This is a schematic diagram of a flow chart of the digital human voice conversation question-answering interactive method described in this embodiment;
[0053] Figure 4This is a timing diagram of the standard question-and-answer interaction method for digital human conversation described in this embodiment;
[0054] Figure 5 This is a schematic diagram of the interface of the standard question-and-answer interactive method for digital human conversation described in this embodiment;
[0055] Figure 6 This is the overall flow chart of the standard question-answer interactive control process + exception handling method described in this implementation mode;
[0056] Figure 7 This is a schematic diagram of the voice-level resampling interface described in this implementation mode;
[0057] Fig. 8A and Figure 8B This is a schematic diagram of the session-level resampling interface described in this implementation mode;
[0058] Fig. 9 This is a schematic diagram of the voice broadcast start or stop state interface described in this implementation mode;
[0059] Fig.10 This is a schematic diagram of the interface for returning to the initial state with one key according to this embodiment. DETAILED DESCRIPTION
[0060] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0061] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0062] According to the first aspect of the present invention, there is also provided a digital human voice conversation question-answering interactive control and exception handling method 100, referring to Figure 1 As shown, the method 100 includes:
[0063] S101: Setting various system parameters and states of various parameters, and controlling the digital human voice conversation question-answering interactive process according to the states of the parameters;
[0064] S102: If an abnormality occurs during the digital human voice conversation question-and-answer interaction process, determine at which stage of voice recognition, answer generation, and voice playback the abnormality occurs;
[0065] S103: In case of abnormalities in speech recognition, voice-level resampling is performed, the current speech recognition result is cleared, and the user's voice is recaptured. In case of abnormalities in answer generation, session-level resampling is performed, the current answer generation result is cleared, and the text stream of the user's question is re-input into the large model service to generate an answer. In case of abnormalities in voice playback, when the user presses the broadcast start / stop button, the voice broadcast is started or paused.
[0066] Optionally, set various system parameters and their states, including:
[0067] Set the system parameters as follows: system-level state SysState, business-level state AiState, and speech broadcast mode SpeechMode;
[0068] The parameter states of the system-level state SysState include: uninitialized INVALID, activated state ACTIVATED, and a question-and-answer session SESSION in a certain round;
[0069] The parameter states of the business-level state AiState include: uninitialized INVALID, capturing CAPTURE, speech recognition ASR, large model generation CHATGPT, speech synthesis, and speech broadcast TTS.
[0070] The parameter states of the speech broadcast mode SpeechMode include: uninitialized INVALID, start speech broadcast SPEECH_ON, and stop speech broadcast SPEECH_OFF.
[0071] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0072] Set the parameter state of the system-level state SysState to a question-and-answer session SESSION in a certain round, and determine the start of a new round of question-and-answer interaction session;
[0073] Capture user voice in real time. When the user starts speaking, send the captured user voice stream to the speech recognition service in real time to convert it into a text stream. Set the parameter state of the business-level state AiState to speech recognition ASR in progress.
[0074] The captured user voice stream is sent to the speech recognition service in real time to be converted into a text stream, and the parameter state of the business-level state AiState is set to speech recognition ASR in progress.
[0075] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0076] If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the question text segments generated during the speech recognition process are serially summarized in order to form a complete user question;
[0077] Push the complete user question to the big model, generate the corresponding text answer, and set the parameter state of the business-level state AiState to CHATGPT, which is generating the answer in the big model.
[0078] Optionally, the digital human voice conversation question-answering interaction process is controlled according to the parameter state, including:
[0079] Perform speech synthesis on the text answer, convert the text stream into a speech stream and broadcast it in real time, use the voice stream to drive the AI digital human to perform corresponding actions, and set the parameter status of the business-level state AiState to speech synthesis and speech broadcast TTS.
[0080] Optionally, voice-level resampling is performed for abnormal speech recognition, the current speech recognition result is cleared, and the user's voice is recaptured, including:
[0081] During the speech recognition process, if the user's language standard, voice loudness, speech speed, or speech recognition service itself is wrong, the speech recognition is determined to be abnormal, and the parameter state of the current business-level state AiState is detected to be speech recognition ASR;
[0082] If the parameter state of the current service-level state AiState is in speech recognition ASR, clear the current speech recognition result, recapture the user's voice, send the captured user voice stream to the speech recognition service, perform speech recognition, and obtain the corresponding text result;
[0083] The text results of speech recognition are displayed in real time on the terminal interface.
[0084] Optionally, session-level resampling is performed to solve the problem of abnormal answer generation, clear the current answer generation result, and re-input the text stream of the user question into the large model service to generate the answer, including:
[0085] During the answer generation process, if the feedback answer is irrelevant, inaccurate, or the big model service is abnormal, the answer generation abnormality is determined, and the parameter state of the current business-level state AiState is detected to be the big model generated answer CHATGPT;
[0086] If the parameter state of the current business-level state AiState is the big model generating answer CHATGPT, clear the current big model answer result, re-input the text stream of the user question into the big model service, and the big model service returns the text stream of the corresponding answer result;
[0087] The text information of the response result is displayed in real time on the terminal interface, speech synthesis is performed, and the voice stream is used to drive the digital human's corresponding actions and voice broadcasts.
[0088] Optionally, in case of abnormal voice playback, when the user presses the broadcast start / stop button, the voice broadcast is started or paused, including:
[0089] During speech playback, when the user presses the speech start / stop button, the state of the current speech playback mode SpeechMode is detected;
[0090] If the current speech broadcast mode SpeechMode is equal to SPEECH_ON, it is turned off. The system detects whether the current business-level state AiState is in the state of speech synthesis or speech broadcast TTS. If so, it stops inputting the voice stream to the sound card and stops the speech broadcast function. If it is not in the state of speech synthesis or speech broadcast TTS, a round of conversation ends.
[0091] Optionally, in case of abnormal voice playback, when the user presses the broadcast start / stop button to start or pause the voice broadcast, the method further includes:
[0092] If SpeechMode is equal to SPEECH_ON, set SpeechMode to SPEECH_OFF, change the speech broadcast state from the original on state to the stopped state, and set the broadcast start and stop button to the style icon of the stopped state;
[0093] If SpeechMode is equal to SPEECH_OFF, set SpeechMode to SPEECH_ON, that is, the voice broadcast state is changed from the original off state to the on state, and the broadcast start and stop button is set to the style icon of the on state.
[0094] Optionally, it also includes:
[0095] During the digital human voice conversation question-and-answer interaction process, when the user presses a button to return to the initial state, the current service-level state AiState is detected to see whether it is in the voice recognition ASR, large model answer CHATGPT, speech synthesis or voice broadcast TTS service state;
[0096] If the current business-level state AiState is in the speech recognition ASR, large model answer CHATGPT, speech synthesis or speech broadcast TTS business state, a signal to return to the initial state is sent to the system. After the system captures the signal to return to the initial state, the current process link is immediately ended;
[0097] Returns the initial state of user voice capture and sampling and sets the initial state parameters.
[0098] According to another aspect of the present invention, a digital human voice conversation question-answering interaction method 200 is provided. Figure 2 As shown, the method 200 includes:
[0099] S201: waking up the device based on user-specific audio information, entering the user consultation mode after starting the service, capturing the user's voice in real time, sending the user's voice to the speech recognition service, and converting the voice stream of the user's voice into a text stream;
[0100] S202: If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the text stream is summarized and semantically corrected to form a complete user question;
[0101] S203: Push the complete user question to the big model, generate the corresponding text answer, perform speech synthesis on the text answer and broadcast it in real time;
[0102] S204: Display user questions and text answers on the terminal interface, and drive the AI digital human to take actions based on the real-time voice stream.
[0103] Specifically, participate in Figure 3 As shown, the user and the digital human conduct voice interaction and Q&A. The user asks a question, and the AI digital human gives a corresponding response based on the question. The response is pushed through text output and voice broadcast.
[0104] 1. The user wakes up the device, clicks the voice interaction button, the system enters the user consultation mode, and then asks questions;
[0105] The system can be waked up by voice. The system captures user-specific audio information, such as "Hello Xiaozhi", and the system enters the user consultation mode.
[0106] After the system application is started, the device completes self-checking and initialization (including self-checking of speech recognition service, self-checking of speech synthesis service, self-checking of large model service, initialization of speech capture function, etc.), and automatically enters the awakening state, that is, automatically enters the user consultation mode;
[0107] 2. The system captures the user's voice in real time, sends the voice stream to the speech recognition service, and converts the user's voice stream into a text stream;
[0108] 3. When the system detects that the user's voice has not been captured within a certain period of time (such as 2 seconds), it indicates that the user has finished stating the question. The system summarizes and semantically corrects the entire text stream of the speech recognition service results to form a complete user question;
[0109] 4. Push the question text stream to the big model to generate the corresponding answer;
[0110] 5. The system synthesizes the text answers into speech and announces the answers in speech;
[0111] 6. Synchronously, the text information of the question and answer is displayed in real time on the system terminal; and the digital human is driven to perform corresponding actions (facial expressions, lip shape, body posture, body movements, etc.) according to the real-time voice stream of the answer;
[0112] Note: How to drive the digital human to perform actions according to the voice stream is not the focus of the present invention. The industry already has mature technology and will not be elaborated on here.
[0113] 7. Go to step 2 for a new round of question and answer conversation.
[0114] Note: For the status of each link in the system process and the specific execution logic control, please refer to the "Control process and exception handling method of digital human conversation question and answer interaction" in Example 2.
[0115] Digital Human Conversation Standard Question and Answer Interaction Method Timing Diagram Reference Figure 4 As shown, the interface diagram of the digital human conversation standard question-answering interaction method is referenced Figure 5 shown.
[0116] =reference Figure 6 As shown in the figure, the standard question-answer interaction control method refers to: monitoring the status and executing logical control of the main process of question-answer interaction, including device wake-up, voice capture and sampling, voice recognition, large model answering, voice synthesis, voice broadcasting, terminal display and digital human action drive. Exception handling methods include: voice-level resampling, session-level resampling, voice broadcast start-stop switching, one-key return to initial state, etc.
[0117] The method of the present invention adopts a question-and-answer interaction control engine to monitor the state of the interaction process and control the execution logic to ensure the normal execution of each complex link of the question-and-answer interaction process. In the question-and-answer interaction control engine, the following system parameter mechanism is established for state monitoring and execution logic control, and supports the exception handling process, which greatly enhances the applicability of the system application.
[0118] Table 1 Parameter status
[0119]
[0120]
[0121] Refer to Table 1, which shows the states of various system parameters. Specific methods for monitoring the interactive state of conversation questions and answers and executing logic control
[0122] 1. After the system starts, set the status of each parameter:
[0123] System-level status SysState: INVALID;
[0124] Business-level state AiState: INVALID;
[0125] SpeechMode: INVALID;
[0126] 2. After the system completes self-check and initialization (including voice service ASR, TTS self-check, large model service CHATGPT self-check, voice capture function initialization, etc.), and there is no abnormality in the system, set the system initial state:
[0127] System-level state SysState: ACTIVATED;
[0128] Business-level state AiState: CAPTURE;
[0129] SpeechMode: INVALID;
[0130] Note: The above assumes that the system self-checks and initializes after startup, the system wakes up, and automatically enters the user session consultation mode.
[0131] 3. The system captures the sound. When the user starts speaking, the system captures the user's effective sound and attempts to start a new round of conversation.
[0132] Note: The user and the digital human are engaged in a question-and-answer interaction, where one question and one answer constitute a conversation round. You can conduct a single-round question-and-answer conversation (including one conversation round) or a multi-round conversation (including multiple conversation rounds).
[0133] 4. Before starting a new round of conversation, first check the system-level state SysState to see if it is: ACTIVATED. If it is, you can start a new round of Q&A conversation; otherwise, you cannot start a new round of Q&A conversation. For example, if SysState is SESSION, it means that a round of Q&A conversation is in progress and you cannot start a new round of conversation.
[0134] 5. Start a new round of voice question-and-answer interactive conversation. The specific process is as follows:
[0135] a) Setting parameters: System-level state SysState: SESSION, marking the start of a new round of question-and-answer interactive session;
[0136] b) The system sends the captured user voice stream to the speech recognition service (ASR) in real time to convert the user voice stream into a text stream.
[0137] At the same time, set the parameters: business level state AiState: ASR.
[0138] This business process is called an ASR process, and the result of each ASR process will generate a question text segment. There may be one or more ASR processes. If there are multiple ASR processes, the result will generate multiple corresponding question text segments.
[0139] During the ASR process, if an abnormality occurs, speech-level resampling can be performed. For details, see the subsequent methods.
[0140] c) When the system detects that the user's voice has not been captured within a certain period of time (for example, 2 seconds), it indicates that the user has finished telling the question. The system will sequentially connect and summarize all the question text segments generated by the above ASR process results to form a complete user question;
[0141] d) The system sends the above complete user question to the big model service to generate the corresponding answer. This business process is called CHATGPT process, and at the same time sets: business level state AiState: CHATGPT.
[0142] During the CHATGPT process, if an exception occurs, session-level resampling can be performed. For details, see the subsequent methods;
[0143] e) The system receives the answer text stream generated by the large model service in real time, and then sends it to the speech synthesis service to convert the answer text stream into a speech stream, and then broadcasts the speech stream to the system in real time, and drives the AI digital human to perform corresponding actions with the speech stream. This business process is called the TTS process, and at the same time sets: business-level state AiState: TTS.
[0144] When the AiState service-level status is not INVALID, the user can switch the start or stop status of the voice broadcast, that is, start or stop the voice broadcast. For details, please refer to the subsequent methods;
[0145] When the AiState service-level status is ASR, CHATGPT, or TTS, the user can return to the initial state with one click. For details, see the subsequent methods.
[0146] f) When the TTS process is finished, set:
[0147] System-level state SysState: ACTIVATED, which means that the session has changed from SESSION to ACTIVETED (the SysState state will be checked in step 4), and a new session can be started;
[0148] Business-level state AiState: CAPTURE;
[0149] g) Return to step 3.
[0150] Specific methods for handling exceptions in conversational question-and-answer interactions
[0151] The voice question-and-answer interaction between users and virtual digital humans includes a series of links, including user question voice capture and sampling, question voice recognition, large model return of answer text, answer text voice synthesis, answer voice broadcast, etc. These links are closely linked. If a problem occurs in the previous link, all the subsequent links are meaningless and the subsequent links need to be terminated in time. However, the system application itself is often not sure that an abnormal problem has occurred in a certain link, which requires timely human intervention to avoid meaningless links and actions of the system.
[0152] For example, the speech recognition process is inevitably affected by many factors such as the user's language standard, voice loudness, speaking speed, the stability, accuracy and recognition efficiency of the speech recognition service itself, and often has anomalies: inaccurate speech recognition results, speech recognition service anomalies (such as inability to access the speech recognition server, etc.). Speech recognition service anomalies will directly lead to the unavailability of the entire system service. If the user's speech recognition results are inaccurate, the subsequent large model's answer feedback and voice broadcast will be meaningless.
[0153] In the method of the present invention, the "speech level resampling" technology is adopted to address the abnormal situation of inaccurate speech recognition results. Specific method:
[0154] The current service-level state AiState detected by the system, whether it is in the speech recognition ASR process;
[0155] If not, the subsequent logic is not executed;
[0156] Clear the current speech recognition results;
[0157] Recapture the user's voice, send the captured user voice stream to the speech recognition service, perform speech recognition, and obtain the corresponding text result;
[0158] The text results of speech recognition are displayed in real time on the terminal interface.
[0159] During the ASR process, you can perform a "speech level resampling" interface diagram, refer to Figure 7 shown.
[0160] Another example: After the user asks a question, the big model will give the corresponding response feedback. The feedback results may be abnormal: irrelevant answers, inaccurate answers, big model service abnormalities, etc. The user may not be satisfied with the feedback results, and often needs to regenerate the response results one or more times based on the previous questions. Moreover, if the user is not satisfied with the previous feedback results, the subsequent voice broadcast will be meaningless and will appear superfluous.
[0161] The method of the present invention adopts the "session level resampling" technology to address this anomaly.
[0162] Specific methods:
[0163] 1. The system detects whether the current business-level state AiState is in the process of generating answers from the large model CHATGPT;
[0164] 2. If not, do not execute subsequent logic;
[0165] 3. Clear the current large model response result;
[0166] 4. Re-input the text stream of the user's question into the big model service, and the big model service returns the text stream of the corresponding answer result;
[0167] 5. Display the response result text information in real time on the terminal interface;
[0168] 6. Perform speech synthesis TTS and use the voice stream to drive the digital human's corresponding actions and voice broadcasts.
[0169] refer to Fig. 8A As shown in Figure AB, this is a schematic diagram of the session-level resampling interface.
[0170] In addition, the voice broadcast function is related to factors such as the specific scenario and the user's mentality at the time. Sometimes it is needed, and sometimes it is not. In the method of the present invention, the "voice broadcast start-stop switching" technology is adopted to address this problem, that is, a user interface is provided, and a main switch button for switching the voice broadcast start or stop state is provided, so that the voice broadcast function can be switched to start and stop at any time by clicking the button, so as to respond to the user's scenario needs and increase the user interaction experience.
[0171] Note: The necessary part of voice broadcast is speech synthesis TTS. The process of speech synthesis TTS is the process of converting the text stream of the answer into a speech stream. The system can transmit the speech stream output by the TTS process to the system sound card in real time to play the sound. In other words, when you need to start the broadcast, you only need to transmit the speech stream to the system sound card; when you need to stop the broadcast, you only need to stop inputting the speech stream to the sound card.
[0172] Specific method to switch the voice broadcast start or stop status:
[0173] 1. When the user presses the speech start / stop button, the system detects the status of the current speech broadcast mode SpeechMode;
[0174] 2. If SpeechMode is equal to SPEECH_ON, it means that voice broadcast has been turned on and needs to be turned off. The system detects whether the current business-level state AiState is in TTS state. If it is in TTS state, it means that the current system is in the voice broadcast stage in TTS state, then immediately stop inputting voice stream to the sound card, that is, truly stop the voice broadcast function. If it is not in TTS state, it means that the current system is not in the state of speech synthesis and voice broadcast, and the TTS process will no longer be needed in the business-level link. After the CHATGPT process is completed, a round of conversation ends, thus avoiding subsequent meaningless links and actions;
[0175] 3. If SpeechMode is equal to SPEECH_ON, set SpeechMode to SPEECH_OFF, that is, change the voice broadcast state from the original on state to the stopped state, and set the broadcast start and stop button to the style icon of the stopped state;
[0176] 4. If SpeechMode is equal to SPEECH_OFF, set SpeechMode to SPEECH_ON, that is, the speech broadcast state is changed from the original off state to the on state, and the broadcast start and stop button is set to the style icon of the on state;
[0177] 5. In each round of conversational Q&A interaction, when the big model generates an answer, it will detect SpeechMode. If SpeechMode is SPEECH_ON (voice broadcast needs to be turned on), the TTS speech synthesis process is required, and the system will output the TTS result, that is, the voice stream, to the system sound card in real time for sound playback. If SpeechMode is SPEECH_OFF (voice broadcast needs to be turned off), the TTS speech synthesis process is no longer required, and there is no generation of the voice stream, and there is no subsequent link and action of transmitting the voice stream to the sound card for voice broadcast.
[0178] refer to Fig. 9 The figure shows the interface diagram of voice broadcast start or stop status.
[0179] The present invention also provides a "one-key return to initial state" technology. The so-called "initial state" refers to the system activation state, the system is awake, can monitor, and can capture user voice and samples. The system-level state SysState is ACTIVATED, and the business-level state AiState is CAPTURE. The initial state is the prerequisite for the user and the digital human system to start a conversation and question-answering interaction.
[0180] Specific methods:
[0181] 1. When the user presses the "One-key return to initial state" button, the system detects whether the current business-level state AiState is in the business states of speech recognition ASR, large model answer CHATGPT, speech synthesis and speech broadcast TTS;
[0182] 2. If not, ignore the button click event and do not execute subsequent logic;
[0183] 3. Send a signal to the system to return to the initial state;
[0184] 4. After the system captures the signal, it immediately ends the current process;
[0185] 5. The system returns to the initial state of user voice capture and sampling, that is, setting parameters:
[0186] System-level state SysState: ACTIVATED;
[0187] Business-level state AiState: CAPTURE.
[0188] refer to Fig.10 The figure shows a schematic diagram of the interface for returning to the initial state with one click.
[0189] Optionally, after the system is started, set the status of various system parameters, including:
[0190] Set the status of each parameter to:
[0191] System-level state SysState: invalid, uninitialized INVALID;
[0192] Business-level state AiState: invalid, uninitialized INVALID;
[0193] Speech mode SpeechMode: invalid, uninitialized INVALID.
[0194] Therefore, for the important links in the voice conversation question-and-answer interaction process, it supports manual operation intervention and provides a humanized interactive response control mechanism. It can promptly handle anomalies that are difficult to detect in the system itself, avoid meaningless links and actions of the system, and respond to user scenario needs in a timely manner, greatly enhancing the applicability of the system application scenarios and the interactive experience of the conversation process.
[0195] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0196] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0197] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0198] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0199] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0200] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A digital human voice conversation question-answering interactive control and exception handling method, characterized in that: include: Set the system parameters and their states, and control the digital human voice conversation question-and-answer interaction process according to the parameter states; If an abnormality occurs during the digital human voice conversation question-and-answer interaction, determine at which stage of the voice recognition, answer generation, and voice playback the abnormality occurs; In case of anomalies in the speech recognition process, voice-level resampling is adopted to clear the current speech recognition results and recapture the user's voice. In case of anomalies in the answer generation process, session-level resampling is adopted to clear the current answer generation results and re-input the text stream of the user's question into the large model service to generate the answer. In case of anomalies in the voice playback process, when the user presses the broadcast start / stop button, the voice broadcast is started or paused.
2. The method according to claim 1, characterized in that Set various system parameters and their status, including: Set the system parameters as follows: system-level state SysState, business-level state AiState, and speech broadcast mode SpeechMode; The parameter states of the system-level state SysState include: uninitialized INVALID, activated state ACTIVATED, and a question-and-answer session SESSION in a certain round; The parameter states of the business-level state AiState include: uninitialized INVALID, capturing CAPTURE, speech recognition ASR, large model generation CHATGPT, speech synthesis, and speech broadcast TTS. The parameter states of the speech broadcast mode SpeechMode include: uninitialized INVALID, start speech broadcast SPEECH_ON, and stop speech broadcast SPEECH_OFF.
3. The method according to claim 2, characterized in that The digital human voice conversation question-answering interactive process is controlled according to the parameter status, including: Set the parameter state of the system-level state SysState to a question-and-answer session SESSION in a certain round, and determine the start of a new round of question-and-answer interaction session; Capture user voice in real time. When the user starts speaking, send the captured user voice stream to the speech recognition service in real time to convert it into a text stream. Set the parameter state of the business-level state AiState to speech recognition ASR in progress. The captured user voice stream is sent to the speech recognition service in real time to be converted into a text stream, and the parameter state of the business-level state AiState is set to speech recognition ASR in progress.
4. The method according to claim 2, characterized in that: The digital human voice conversation question-answering interactive process is controlled according to the parameter status, including: If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the question text segments generated during the speech recognition process are serially summarized in order to form a complete user question; Push the complete user question to the big model, generate the corresponding text answer, and set the parameter state of the business-level state AiState to CHATGPT, which is generating the answer in the big model.
5. The method according to claim 2, characterized in that: The digital human voice conversation question-answering interactive process is controlled according to the parameter status, including: Perform speech synthesis on the text answer, convert the text stream into a speech stream and broadcast it in real time, use the voice stream to drive the AI digital human to perform corresponding actions, and set the parameter status of the business-level state AiState to speech synthesis and speech broadcast TTS.
6. The method according to claim 2, characterized in that If speech recognition is abnormal, voice-level resampling is performed to clear the current speech recognition results and recapture the user's voice, including: During the speech recognition process, if the user's language standard, voice loudness, speech speed, or speech recognition service itself is wrong, the speech recognition is determined to be abnormal, and the parameter state of the current business-level state AiState is detected to be speech recognition ASR; If the parameter state of the current service-level state AiState is in speech recognition ASR, clear the current speech recognition result, recapture the user's voice, send the captured user voice stream to the speech recognition service, perform speech recognition, and obtain the corresponding text result; The text results of speech recognition are displayed in real time on the terminal interface.
7. The method according to claim 2, characterized in that: In case of anomalies in answer generation, session-level resampling is performed to clear the current answer generation results and re-input the text stream of the user's question into the large model service to generate answers, including: During the answer generation process, if the feedback answer is irrelevant, inaccurate, or the big model service is abnormal, the answer generation abnormality is determined, and the parameter state of the current business-level state AiState is detected to be the big model generated answer CHATGPT; If the parameter state of the current business-level state AiState is the big model generating answer CHATGPT, clear the current big model answer result, re-input the text stream of the user question into the big model service, and the big model service returns the text stream of the corresponding answer result; The text information of the response result is displayed in real time on the terminal interface, speech synthesis is performed, and the voice stream is used to drive the digital human's corresponding actions and voice broadcasts.
8. The method according to claim 2, characterized in that: In case of abnormal voice playback, when the user presses the start / stop button to start or pause the voice playback, the following events may occur: During speech playback, when the user presses the speech start / stop button, the state of the current speech playback mode SpeechMode is detected; If the current speech broadcast mode SpeechMode is equal to SPEECH_ON, it is turned off. The system detects whether the current business-level state AiState is in the state of speech synthesis or speech broadcast TTS. If so, it stops inputting voice stream to the sound card and stops the speech broadcast function. If it is not in the state of speech synthesis or speech broadcast TTS, a round of conversation ends.
9. The method according to claim 8, characterized in that In case of abnormal voice playback, when the user presses the start / stop button to start or pause voice playback, it also includes: If SpeechMode is equal to SPEECH_ON, set SpeechMode to SPEECH_OFF, change the speech broadcast state from the original on state to the stopped state, and set the broadcast start and stop button to the style icon of the stopped state; If SpeechMode is equal to SPEECH_OFF, set SpeechMode to SPEECH_ON, the speech broadcast state changes from the original off state to the on state, and the broadcast start and stop button is set to the style icon of the on state.
10. The method according to claim 1, characterized in that Also includes: During the digital human voice conversation question-and-answer interaction process, when the user presses a button to return to the initial state, the current service-level state AiState is detected to see whether it is in the voice recognition ASR, large model answer CHATGPT, speech synthesis or voice broadcast TTS service state; If the current business-level state AiState is in the speech recognition ASR, large model answer CHATGPT, speech synthesis or speech broadcast TTS business state, a signal to return to the initial state is sent to the system. After the system captures the signal to return to the initial state, the current process link is immediately ended; Returns the initial state of user voice capture and sampling and sets the initial state parameters.
11. A digital human voice conversation question-answering interactive method, characterized in that: include: Wake up the device based on user-specific audio information, enter user consultation mode after starting the service, capture the user's voice in real time, send the user's voice to the speech recognition service, and convert the voice stream of the user's voice into a text stream; If it is detected that the user's voice is not captured within the predetermined time, it is determined that the user's question is over, and the text stream is summarized and semantically corrected to form a complete user question; Push the complete user question to the big model, generate the corresponding text answer, synthesize the text answer and broadcast it in real time; Display user questions and text answers on the terminal interface, and drive the AI digital human to take actions based on the real-time voice stream.
Citation Information
Patent Citations
Man-machellone interaction processing method and device and electronic device
CN109215642A
System for realizing intelligent real-time interactive question and answer based on virtual digital human and processing method thereof
CN116229977A
Voice interaction method and device and storage medium
CN117456998A
Live broadcast monitoring processing method, system, equipment and medium
CN117692665A
An artificial intelligence-based chatbot conversation consultation system and method thereof
KR102653266B1