Question-answer method and related apparatus

By deploying large language models in terminal devices and cloud devices, detecting user instructions and cacheing unreported content, the problem of inability to continue broadcasting in the prior art is solved and the user experience is improved.

WO2025130156A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/116567
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-09-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing question-and-answer method cannot recognize the user instructions that continue to broadcast after the user interrupts the broadcast, resulting in the user needing to ask questions again to obtain the complete broadcast content, which has a poor user experience.

Method used

By deploying a large language model in terminal devices and cloud devices, user instructions are detected and interrupt broadcast requests are generated, and unreported content is cached so that the user instructions continue to broadcast when they appear.

Benefits of technology

It realizes that the user can continue to broadcast after interrupting the broadcast, improving the user experience, and users can continue to listen to unreported content at the interrupted location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116567_26062025_PF_FP_ABST
    Figure CN2024116567_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A question-answer method, applied to a terminal device which is deployed with a large language model. The method comprises: the large language model generating answer content on the basis of a first user instruction, and the terminal device broadcasting the answer content; in the answer content broadcasting process, detecting a second user instruction, interrupting the broadcast, sending a processing instruction to the large language model, on the basis of the processing instruction, the large language model instructing to interrupt the generation of the answer content or continuing to generate the answer content, and caching the continuously generated content; in response to a broadcast interruption request, interrupting the broadcast of the answer content; detecting a third user instruction to instruct to continue to broadcast, broadcasting unbroadcast content that is obtained by the large language model on the basis of the first user instruction and broadcast content, or is obtained by reading the cached continuously generated content. According to the present application, on the basis of the characteristics of edge generation and edge broadcast of a large language model, the broadcast of answer content can be continued after an interruption by controlling the generation process of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Question-answering method and related device

[0001] This application claims priority to Chinese patent application No. 202311764553X, filed on December 20, 2023, entitled “A Question-Answering Method and Related Devices,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a question-answering method and related devices. Background Art

[0003] With the development of artificial intelligence, question-and-answer technology is becoming increasingly mature. Existing question-and-answer methods can interrupt the broadcast of the current response based on user instructions, thereby starting a new conversation. However, when the user wants to continue broadcasting the interrupted response, because existing question-and-answer methods only associate adjacent conversations, they cannot recognize the user's instructions to continue broadcasting. This means that the user is required to ask the question again and regenerate the broadcast content, without being able to continue broadcasting at the interruption point, resulting in a poor user experience.

[0004] Summary of the Invention

[0005] In order to solve the above problems, an embodiment of the present application provides a question-answering method and related devices, which perform corresponding processing after the broadcast is interrupted, thereby achieving the purpose of continuing to broadcast the answer content after the interruption.

[0006] To this end, the following technical solutions are adopted in the embodiments of the present application:

[0007] In a first aspect, an embodiment of the present application provides a question-and-answer method, which is applied to a terminal device, which is deployed with a large language model. The method mainly includes: the large language model generates a response content based on a first user instruction, and the terminal device broadcasts the response content; in the process of broadcasting the response content, the terminal device detects a second user instruction, generates an interrupt broadcast request, and sends a processing instruction to the large language model, the large language model instructs to interrupt the generation of the response content based on the processing instruction, or continues to generate the response content based on the processing instruction, and the terminal device or cloud device caches the continued generated content; upon receiving the interrupt broadcast request, the terminal device interrupts the broadcast of the response content; a third user instruction is detected, and the third user instruction instructs to continue broadcasting the response content, and the terminal device broadcasts the unbroadcasted content in the response content, where the unbroadcasted content is generated by the large language model based on the first user instruction and the broadcasted content, or the unbroadcasted content is read from the cached continued generated content.

[0008] That is to say, the characteristics of generating response content based on the large language model: generating and broadcasting at the same time, the embodiment of the present application provides a question-and-answer method, so that the broadcast of the response content can be continued after an interruption. Analyzing from the user's perspective, the user sends a first user instruction to obtain the response content for broadcast; during the broadcast of the response content, the user sends a second user instruction, which requires interrupting the broadcast of the response content to execute the second user instruction or other instructions; after executing several user instructions including the second user instruction, the user can send a third user instruction to continue broadcasting the unbroadcasted part of the response content after the interruption.

[0009] In an embodiment of the present application, in order to achieve the purpose of continuing the broadcast, after receiving the second user instruction, an interruption broadcast request and a processing instruction are generated. First, the interruption broadcast request can control the interruption of the broadcast of the current answer content to respond to other instructions of the user. Second, the processing instruction can control the interruption of the generation of the answer content while interrupting the broadcast, and record the first user instruction and the broadcasted content, so that when the user instruction to continue the broadcast appears, the unbroadcasted content is determined based on the recorded first user instruction and the broadcasted content; or the generation of the answer content is not interrupted, that is, the answer content continues to be generated, and the unbroadcasted content is recorded from the interruption position, so that when the user instruction to continue the broadcast appears, the unbroadcasted content can be directly obtained. Therefore, the embodiment of the present application realizes the continuous broadcast of the answer content. Therefore, based on the method of the embodiment of the present application, the user can interrupt the broadcast during the broadcast of the answer content to execute other instructions of the user; after executing other instructions, the user can continue to broadcast the answer content from the position where the broadcast was interrupted. This application controls the generation process of the large language model based on the characteristics of simultaneous generation and broadcasting of the large language model, so that the broadcast of the response content can be continued after interruption.

[0010] In one possible implementation, when the large language model interrupts the generation of the response content, the first user instruction and / or the announced content are cached. When a third user instruction is detected, the cached first user instruction and / or the announced content are retrieved, and the unannounced content is obtained using the large language model.

[0011] In this implementation, a method for determining unbroadcasted content based on a first user instruction and the broadcasted content is provided. Exemplarily, the first user instruction and the broadcasted content of the response content are used as inputs of a large language model, and the unbroadcasted content of the response content is output; or the first user instruction is used as input of a voice model, and the entire content of the response content is output, and then based on the broadcasted content of the response content, the unbroadcasted content is determined from the entire content. That is, when the response content is interrupted from playing, only the first user instruction and the broadcasted content of the response content are cached. If there is a user instruction to continue broadcasting, the unbroadcasted content is generated by the voice model based on the cached first user instruction and the broadcasted content of the response content; if there is no user instruction to continue broadcasting, the voice model is not required to generate the response content. Therefore, this implementation method calls the large language model only when there is a user instruction to continue broadcasting, occupies less logical resources, and is highly efficient.

[0012] In another possible implementation, when the large language model interrupts the generation of the response content, the context of the first user instruction is also cached. When a third user instruction is detected, the cached first user instruction, the context of the first user instruction, and the announced content are retrieved, and the unannounced content is obtained using the large language model.

[0013] In this implementation, the context of the first user instruction also needs to be recorded in order to more accurately resume the broadcast. For example, the context can be the conversation turn at the time of interruption, whether the user's car is parked, the user's historical questions, the user's profile, etc. For example, after the user interrupts the broadcast, the user has multiple conversations, and the context may have changed at this time, such as the user's historical questions have changed or the user profile has been updated. Therefore, when continuing the broadcast, the broadcast content after the breakpoint can be generated based on the broadcast content and the context of the recorded first user instruction, so that the continued broadcast content is more consistent with the scene before the interruption, without being disturbed by the context change.

[0014] In another possible implementation, when a second user instruction is detected, an interruption broadcast request is generated and a processing instruction is sent to the large language model only if the user intention corresponding to the second user instruction meets a preset condition.

[0015] In this implementation, the streaming broadcast of the language model can be interrupted only when the user has a clear intention, thereby achieving the purpose of controllable interruption of the streaming broadcast.

[0016] In another possible implementation, determining whether the user intent corresponding to the second user instruction satisfies a preset condition is to use a preset whitelist. Specifically, by determining that the user intent corresponding to the second user instruction is within the preset whitelist, the second user instruction is determined to interrupt the broadcast of the current response content, thereby generating an interruption request and sending a processing instruction to the large language model.

[0017] In this implementation, the present application does not limit the method for determining whether a user has explicit intent. For example, a whitelist management method can be used. User intents within a preset whitelist are considered explicit intents, while user intents not within the preset whitelist are considered implicit intents. When an intent on the whitelist occurs, the language model can respond to the user's intent by performing actions such as interrupting the streaming broadcast or starting a new conversation turn.

[0018] In another possible implementation, one of the uses of the natural language understanding model is to perform user intent analysis. Therefore, the user intent corresponding to the second user instruction can be determined through the natural language understanding model.

[0019] In this implementation, the intention recognition function of the natural language understanding model can be used to understand and analyze user instructions and determine the user intention corresponding to the user instructions.

[0020] In another possible implementation, the embodiment of the present application does not limit the form of the user input content. When the user input content is in the form of audio, speech analysis is performed on the user input audio to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction; when the user input content is in the form of text, the user input text is processed based on preset processing rules to obtain the first user instruction and / or second user instruction and / or third user instruction.

[0021] In this implementation, in the embodiments of the present application, the form of user input is not limited. That is, the user input can be voice or text. When the user input is audio, the user input can be subjected to voice analysis to obtain the user instruction in the embodiments of the present application. When the user input is text, the user input can be processed based on preset processing rules to obtain the user instruction in the embodiments of the present application. The preset processing rules include rules such as generating prompt text based on the user input text.

[0022] On the second aspect, the present application provides a question-answering method, which is applied to a question-answering system. The question-answering system includes a terminal device and a cloud device. The terminal device is communicatively connected to the cloud device. The cloud device is deployed with a large language model. The terminal device detects user instructions and broadcasts the response content. The question-answering method includes: when the terminal detects a first user instruction, it sends the first user instruction to the cloud device; after the cloud device receives the first user instruction, it generates the response content through the large language model; the terminal device receives and broadcasts the response content, and in the process of broadcasting the response content, the terminal device detects a second user instruction and sends the second user instruction to the cloud device; after the cloud device receives the second user instruction, it generates an interruption broadcast request, sends the interruption broadcast request to the terminal device, and sends a processing instruction to the large language model, wherein the processing instruction instructs the large language model to interrupt the generation of the response content, or the processing instruction instructs the large language model to continue the response. Generation of content so that the continuously generated content is cached; after the terminal device receives the interrupt broadcast request, it interrupts the broadcast of the response content; the terminal device detects the third user instruction, sends the third user instruction to the cloud device, and the third user instruction instructs to continue broadcasting the response content; the cloud device receives the third user instruction, generates the unbroadcasted content in the response content based on the first user instruction and the broadcasted content, and determines the unbroadcasted content in the response content through the large language model, and sends the unbroadcasted content to the terminal device; or, the cloud device or the terminal device reads the unbroadcasted content in the response content from the cached continuously generated content; the terminal device broadcasts the obtained unbroadcasted content.

[0023] In one possible implementation, the method further includes: when the large language model interrupts generation of the response content, the cloud device or the terminal device caches the first user instruction and / or the broadcast content. Thus, the large language model can generate the unbroadcast content based on the cached first user instruction and the broadcast content.

[0024] In another possible implementation, the method further includes: when the large language model interrupts the generation of the response content, the cloud device or the terminal device caches the context of the first user instruction. Thus, the large language model can generate the unannounced content based on the cached first user instruction, the context of the first user instruction, and the announced content. This avoids the impact of context changes on the generation of the unannounced content.

[0025] In another possible implementation, the cloud device receives the second user instruction and generates a request to interrupt the broadcast after determining that the user intention corresponding to the second user instruction meets the preset conditions. In other words, the broadcast of the current response content will be interrupted only if the user intention corresponding to the second user instruction meets the preset conditions.

[0026] In another possible implementation, a preset whitelist is set to limit the user intentions that can interrupt the broadcast of the response content. In other words, the broadcast of the response content will be interrupted only if the user intention corresponding to the second user instruction is within the preset whitelist.

[0027] In another possible implementation, the user intention corresponding to the second user instruction is determined through a natural language understanding model.

[0028] In another possible implementation, the applicable scenario of the method is not limited, and the method can be applied to audio and text question-and-answer scenarios. The first user instruction and / or the second user instruction and / or the third user instruction can be obtained based on voice analysis of the audio input by the user, or can be obtained based on processing the text input by the user according to preset processing rules.

[0029] In a third aspect, the present application provides a question-and-answer method, which is applied to a cloud device, wherein the cloud device is deployed with a large language model. The question-and-answer method includes: using the large language model to generate corresponding answer content in response to a first user instruction of a user, and the generated answer content can be broadcast to the user; during the broadcast of the answer content, when a second user instruction is received, an interrupt broadcast request is generated to indicate the interruption of the broadcast of the answer content, and a processing instruction is sent to the large language model to cause the large language model to interrupt or continue the generation of the answer content, when the large language model is instructed to continue the generation of the answer content, the continued generated content is cached, and the interrupt broadcast request is used to interrupt the broadcast of the answer content; receiving a third user instruction, when the third user instruction indicates that the user intends to continue broadcasting the answer content, based on the first user instruction and the broadcast content generation, using the large language model, determining the unbroadcasted content in the answer content, or reading the cached continued generated content to obtain the unbroadcasted content, so that the unbroadcasted content is broadcast. Therefore, the purpose of interrupting and continuing the broadcast process of the answer content is achieved.

[0030] In another possible implementation, the method further includes: when the large language model interrupts generation of the response content, the first user instruction and / or the announced content are cached. Thus, the large language model can generate the unannounced content based on the cached first user instruction and the announced content.

[0031] In another possible implementation, the method further includes caching the context of the first user instruction when the large language model interrupts generation of the response content. Thus, the large language model can generate the unannounced content based on the first user instruction, the context of the first user instruction, and the announced content.

[0032] In another possible implementation, user instructions for interrupting the broadcast are restricted. Further, the user intention corresponding to the second user instruction is determined, and a request to interrupt the broadcast is generated only when the user intention meets a preset condition.

[0033] In another possible implementation, the preset condition is implemented by presetting a whitelist, determining the user intention corresponding to the second user instruction, and interrupting the broadcast of the response content when the user intention is within the preset whitelist.

[0034] In another possible implementation, a natural language understanding model is used to analyze and obtain the user intention corresponding to the second user instruction.

[0035] In another possible implementation, when the user input content is in the form of audio, speech analysis is performed on the user input audio to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction. When the user input content is in the form of text, text processing is performed based on preset processing rules to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction.

[0036] In a fourth aspect, the present application provides a question-and-answer method, which is applied to a terminal device, and the question-and-answer method includes: the terminal device detects a first user instruction, and sends the first user instruction to a cloud device, wherein the first user instruction is used to instruct the cloud device to generate response content based on a large language model; the terminal device receives and broadcasts the response content sent by the cloud device, and in the process of broadcasting the response content, the terminal device detects a second user instruction for instructing the cloud device to generate an interrupt broadcast request, and sends the second user instruction to the cloud device; thereafter, the terminal device receives the interrupt broadcast request sent by the cloud device, and interrupts the broadcast of the response content; the terminal device detects a third user instruction, and sends the third user instruction to the cloud device, and the third user instruction instructs to continue broadcasting the unbroadcasted content in the response content; finally, the terminal device receives the unbroadcasted content in the above-mentioned response content sent by the cloud device, and broadcasts the unbroadcasted content.

[0037] In another possible implementation, when the user input content is in the form of audio, speech analysis is performed on the user input audio to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction. When the user input content is in the form of text, text processing is performed based on preset processing rules to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction.

[0038] In a fifth aspect, the present application provides a terminal device. The terminal device is deployed with a large language model, and the terminal device is used to perform any method of the first aspect above. Alternatively, the terminal device is used to perform any method of the fourth aspect above.

[0039] In the sixth aspect, the present application provides a question-and-answer system, which includes a terminal device and a cloud device, the terminal device is communicatively connected to the cloud device, the cloud device is deployed with a large language model, the terminal device is used to execute any one of the methods of the above-mentioned fourth aspect, and / or the cloud device is used to execute any one of the methods of the above-mentioned third aspect, and / or the question-and-answer system is used to execute any one of the methods of the above-mentioned second aspect.

[0040] In the seventh aspect, the present application provides a cloud device, which is deployed with a large language model. The cloud device is communicatively connected with the terminal device, and the cloud device is used to execute any one of the methods of the third aspect above.

[0041] In an eighth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, enables any one of the above methods to be implemented.

[0042] In a ninth aspect, the present application provides a chip comprising a processing circuit and a storage medium, wherein the storage medium stores computer program code; when the computer program code is executed by the processing circuit, any one of the above methods is implemented.

[0043] In a tenth aspect, the present application provides a computer program product comprising program instructions, which, when executed by a computer, enables the computer to execute any one of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The following is a brief introduction to the drawings required for use in the embodiments or technical descriptions.

[0045] FIG1 is a schematic diagram of a scenario of an existing question-answering method;

[0046] FIG2 is a schematic diagram of a scenario of a question-answering method provided in an embodiment of the present application;

[0047] FIG3 is a schematic diagram of an example scenario of a user and a terminal device in a question-and-answer method provided in an embodiment of the present application;

[0048] FIG4 is a flow chart of a question-answering method provided in an embodiment of the present application;

[0049] FIG5 is a schematic diagram of a first example process of a question-and-answer method provided in an embodiment of the present application;

[0050] FIG6a is a schematic diagram of a question-answering system provided in an embodiment of the present application;

[0051] FIG6 b is a schematic diagram of another question-answering system provided in an embodiment of the present application;

[0052] FIG7 is a schematic diagram of a method flow of a first example of a question-answering system provided in an embodiment of the present application;

[0053] FIG8 is a flow chart of a second example of a question-and-answer method provided in an embodiment of the present application;

[0054] FIG9 is a schematic diagram of a second example method flow of a question-answering system provided in an embodiment of the present application;

[0055] FIG10 is a flow chart of another question-and-answer method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship. For example, A / B means either A or B.

[0057] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.

[0058] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0059] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0060] In order to facilitate understanding of the solution provided by the embodiment of the present application, a brief introduction to some of the terms involved in the solution is first given.

[0061] Single-round recording: After recording is completed once, the system needs to complete all operations of voice-to-text, text-to-intent, and intent-to-broadcast, and then exit the conversation. The user needs to wake up again to continue using voice conversation.

[0062] Multi-round recording: After recording is complete, and all operations of voice-to-text, text-to-intent, and intent-to-announcement are completed, the user can directly speak the next round of conversation without waking up, but cannot interrupt the previous round of conversation.

[0063] Full-duplex recording: Supports processing the previous round of voice-to-text, text-to-intent, and intent-to-broadcast while recording the next round of voice, and can interrupt the previous round of conversation at any time.

[0064] Large Language Models (LLMs) generally refer to neural network models with extremely large parameters (often exceeding one billion). They have the following characteristics: First, they have a massive parameter scale. Large language models contain billions of parameters, and their model size can reach hundreds of GB or even larger. This enormous model size provides them with powerful expressive and learning capabilities. Second, they can perform multi-task learning. Large language models typically learn multiple different NLP (Neuro-Linguistic Programming) tasks simultaneously, such as machine translation, text summarization, and question-answering systems. Multi-task learning enables models to acquire broader and more generalized language understanding capabilities. Third, they require powerful computing resources. Training large language models typically requires hundreds or even thousands of GPUs (graphics processing units) and a significant amount of time, typically weeks to months. Powerful computing resources can accelerate the training process while preserving the capabilities of large language models. Fourth, they require abundant data. Large language models require a large amount of data for training, as only a large amount of data can fully leverage the advantages of their large parameter scale. Furthermore, large language models are widely used in the field of natural language processing and are revolutionizing the state of NLP tasks, giving rise to more powerful and intelligent language technologies. Large language models are a key development direction in AI (artificial intelligence). They also excel in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, and translation. Large language models can be used in a variety of application areas, including automated writing, chatbots, virtual assistants, voice assistants, and automatic translation.

[0065] With the development of artificial intelligence, question-answering technology is becoming increasingly sophisticated. Taking large language models as an example, conversations between users and large speech models can move beyond simple question-and-answer sessions to more analytical and interpretive conversations, such as situation analysis, news analysis, and topic writing. Because these types of conversations are typically lengthy, users often interrupt the current broadcast to discuss other issues, resuming the broadcast after the discussion is complete. Generally speaking, the evolution from single-round recording, multi-round recording, to full-duplex recording has enabled interactive conversations. Especially with the advent of large language models, interruptions and resumptions are common.

[0066] Existing question-and-answer methods can interrupt the current broadcast based on user commands to start a new discussion. However, when the user wants to continue the interrupted broadcast, the system only associates the adjacent conversations, and therefore fails to recognize the user's instructions to continue the broadcast. This means that the user must ask the question again and regenerate the broadcast content, without being able to continue the broadcast at the interruption point, resulting in a poor user experience.

[0067] Please refer to Figure 1, which shows a scenario diagram of an existing question-and-answer method. As shown in Figure 1, in a scenario of the existing question-and-answer method, the user issues an instruction to "analyze the development trend of the Internet." Based on the instruction, the terminal device broadcasts "According to the current development situation, several possible trends in the development of the Internet include: First, the widespread application of artificial intelligence." During the broadcast, the user issues the instruction to "turn on the air conditioner." The terminal device interrupts the broadcast and, based on the instruction to "turn on the air conditioner," broadcasts the corresponding text "OK." The user issues the instruction to "continue analyzing, I'm listening," and wants to continue the broadcast of "analyzing Internet trends." However, the terminal device does not associate the interval conversation turns, cannot recognize the current instruction, and cannot continue the interrupted broadcast.

[0068] There is a solution that configures node configuration information similar to "jump out of this node and be automatically pulled back" in the original conversation flow, so that when the new conversation flow ends after the interruption of the broadcast, the original conversation flow will be pulled back and the broadcast will be carried out.

[0069] This solution pre-configures whether to rewind the conversation flow, which only mechanically rewinds the conversation, rather than automatically redirecting it based on the user's will. This means that when a user is talking to a machine, automatically rewinding certain conversations may conflict with the user's intent, thus reducing the user experience.

[0070] Another solution involves associating the first parameter of the first conversation between the user and the automated assistant with a third-party application. The system can retrieve the first parameter and, based on it, resume the first conversation between the user and the automated assistant. In other words, if the user wishes to continue where they left off, the first parameter can be used to reconstruct the lost conversation context, without forcing the user to restart the conversation from the beginning.

[0071] In this solution, the first parameter can be used to reconstruct the lost conversation context and continue the previously stopped topic. In the case of an interrupted broadcast, this solution can reconstruct the lost conversation context and find the interruption point by using the recorded first parameter, but it cannot know the subsequent broadcast content, making it impossible to continue the broadcast.

[0072] To address the existing technical issues, the present application provides another solution. When a user intends to interrupt the broadcast content, the user's instruction information and the already broadcast content are recorded, or the broadcast is interrupted without interrupting content generation, to obtain the unbroadcasted content after the interruption point. When the user resumes the broadcast content, the broadcast content is regenerated based on the recorded user instruction information and the already broadcast content. Based on the already broadcast content, the broadcast content after the interruption point is obtained for continuous broadcast, or the generated unbroadcasted content is obtained for continuous broadcast.

[0073] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings. In the embodiments of the present application, in order to concisely describe the specific implementation methods of the embodiments of the present application, the embodiments of the present application are described using voice interaction as an example. It should be further explained that the method or related device of the embodiments of the present application can also be applied to any other form of dialogue content, such as text dialogue, where the user interrupts the text display and executes other instructions. When a user instruction to continue display appears, the text display continues from the interruption position.

[0074] Please refer to Figure 2, which shows a scenario diagram of a question-answering method. As shown in Figure 2, a user can receive or send messages through the transceiver module, and the processing module processes the user's instructions and sends the processing results to the transceiver module to achieve interaction with the user.

[0075] Optionally, the voice transceiver module and the processing module can be integrated and deployed in the terminal device, or the terminal device can only deploy the transceiver module to pick up and play the voice, and the voice processing module can be deployed in the server.

[0076] As shown in FIG2 , the terminal device may be the terminal device 101, the terminal device 102, and the terminal device 103 shown in FIG2 , and the server may be the server 104 shown in FIG2 . As shown in FIG2 , when the voice processing module is deployed on the server, the terminal devices 101, the terminal devices 102, and the terminal device 103 are communicatively connected to the server 104 via a communication link. The communication link may include various connection types, such as a wired communication link, a wireless communication link, or an optical fiber cable.

[0077] Optionally, terminal device 101, terminal device 102, and terminal device 103 may be installed with various communication client applications, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc. Terminal device 101, terminal device 102, and terminal device 103 may be various electronic devices with display screens and supporting web browsing, including but not limited to mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc., and the embodiments of the present application do not impose any restrictions on this. Server 104 may be a server that provides various services, such as a background server that provides support for the pages displayed on terminal device 101, terminal device 102, and terminal device 103.

[0078] It should be noted that the question-and-answer method of the embodiment of the present application is generally executed by a terminal device or a server, and interaction with the user is achieved through the terminal device.

[0079] Please refer to Figure 3, which shows a schematic diagram of an example scenario of a user and a terminal device in a question-and-answer method. As shown in Figure 3, in this example, a user issues a first user instruction; the question-and-answer system generates a first response based on the first user instruction; during the broadcast of the first response, the user issues a second user instruction; the question-and-answer system interrupts the first response and generates a second response based on the second user instruction; the user issues a third user instruction, which instructs the user to continue with the first response; the question-and-answer system recognizes the third user instruction and continues broadcasting the unbroadcasted content of the first response.

[0080] In the embodiment of the present application, the Q&A system generates responses using a large language model, which are then broadcasted via the terminal device. That is, the large language model can be deployed on the terminal device, and optionally, the entire Q&A system can be deployed on the terminal device; or the large language model can be deployed on a cloud device, which generates responses and then broadcasts them via the terminal device.

[0081] In the embodiment of the present application, based on the evolution of a single round of recording, it supports processing and recording at the same time, and can associate the previous and next rounds of conversations. That is to say, the previous round of conversation can be interrupted by continuing to send instructions, and the interval conversation rounds can be associated at the same time, so that the streaming can be achieved without interruption and the streaming broadcast can be restored under full-duplex.

[0082] Optionally, the process from interrupting the first response content to continuing the unbroadcasted content is not limited to one user command interaction. That is, after interrupting the first response content, multiple user command interactions can be performed before continuing the unbroadcasted content.

[0083] For example, as shown in FIG3 , the first user instruction issued by the user may be “analyze the development trend of the Internet”; the first response content generated by the question-and-answer system based on the first user instruction may be “according to the current development situation, several possible trends in the development of the Internet include: first, the widespread application of artificial intelligence”, and start streaming broadcast.

[0084] For example, as shown in FIG3 , the second user instruction issued by the user may be a hardware-related instruction such as “turn on the air conditioner”; the question-answering system interrupts the first response content, and based on the second user instruction, recognizes the intention to turn on the air conditioner, makes a second response content of “OK”, and sends the relevant instruction to turn on the air conditioner to the control system. The embodiment of this application does not elaborate on the relevant operations of the control system. In addition, the second user instruction issued by the user may also be a generation-type instruction such as “analyze the development trend of the manufacturing industry”. The question-answering system interrupts the first response content and generates a second response content based on the second user instruction, such as “The development trend of the manufacturing industry can be analyzed from the following aspects”.

[0085] In the embodiment of the present application, the specific implementation method of the second user instruction that interrupts the current broadcast is not limited. For example, as shown in Figure 1, the second user instruction issued by the user can be a hardware-related instruction such as "turn on the air conditioner" or "open the car window"; the terminal device interrupts the first response content, and based on the second user instruction, recognizes the intention to turn on the air conditioner, performs the second response content "OK", and sends the relevant instructions for turning on the air conditioner to the control system. The embodiment of the present application will not go into details about the relevant operations of the control system. The second user instruction issued by the user can also be a generation-type instruction such as "analyze the development trend of the manufacturing industry". The terminal device interrupts the first response content, and based on the second user instruction, broadcasts the second response content "The development trend of the manufacturing industry can be analyzed from the following aspects" and so on.

[0086] For example, as shown in FIG3 , after one or more conversations, the user can issue a third user instruction, “Continue analyzing, I’m listening,” to continue the first response. Because the embodiment of the present application processes the interrupted first response, when the user indicates an intention to resume the conversation, such as “Continue analyzing,” the embodiment of the present application can resume the previously interrupted streaming broadcast and continue broadcasting the unfinished content, such as “Artificial intelligence will penetrate into various fields, including smart homes, smart transportation, and smart healthcare.”

[0087] In the embodiment of the present application, taking the large language model as an example, the streaming broadcast of the large language model can be interrupted at any time. At the same time, the streaming broadcast of the large language model can be interrupted only when there is a clear intention in the user's voice. The embodiment of the present application does not limit the method of judging whether the user has a clear intention. For example, it can be managed by a whitelist, and the intention of users in the whitelist is a clear intention, and the intention of users not in the whitelist is an unclear intention. When the intention under the whitelist appears, the large language model can respond to the user's intention and perform actions such as interrupting the streaming broadcast and starting a new conversation round. After the streaming broadcast of the large language model is interrupted, after determining the user's intention to resume the broadcast, the previous streaming broadcast can be restored and the broadcast can continue from the interrupted breakpoint.

[0088] For example, as shown in FIG3 , the user speaks “Baby, let’s go to the amusement park tomorrow.” When it is determined that the user intention of the voice content is not a clear intention, the streaming broadcast of the large language model is not interrupted.

[0089] Please refer to Figure 4, which shows a flowchart of a question-answering method. As shown in Figure 4, an embodiment of the present application provides a question-answering method, which is applied to a terminal device that is deployed with a large language model. The method mainly includes the following steps:

[0090] Step S401: broadcast the response content, where the response content is generated by the large language model based on the first user instruction.

[0091] In another possible implementation, in response to the user input content being in the form of audio, the method further includes: performing voice analysis on the user input audio to obtain a first user instruction and / or a second user instruction and / or a third user instruction; or, in response to the user input content being in the form of text, the method further includes: processing the user input text based on preset processing rules to obtain a first user instruction and / or a second user instruction and / or a third user instruction.

[0092] Step S402: During the broadcasting of the response content, when a second user instruction is detected, a request to interrupt the broadcasting is generated and a processing instruction is sent to the large language model; the processing instruction instructs the large language model to interrupt the generation of the response content, or the processing instruction instructs the large language model to continue the generation of the response content to cache the content that continues to be generated.

[0093] In one possible implementation, in response to the second user instruction, after determining that the user intention corresponding to the second user instruction meets a preset condition, an interruption broadcast request is generated and a processing instruction is sent to the large language model.

[0094] In another possible implementation, the preset condition is a preset whitelist. After determining that the user intention corresponding to the second user instruction is within the preset whitelist, an interruption broadcast request is generated and a processing instruction is sent to the large language model.

[0095] In another possible implementation, determining whether the user intention corresponding to the second user instruction meets the preset conditions includes: based on the second user instruction, determining the user intention corresponding to the second user instruction through a natural language understanding model; and judging whether the user intention meets the preset conditions.

[0096] Step S403: in response to the broadcast interruption request, interrupt the broadcast of the response content.

[0097] In step S404, a third user instruction is detected, and the unannounced content in the response content is broadcast. The third user instruction instructs to continue broadcasting the response content. The unannounced content is generated by the large language model based on the first user instruction and the announced content, or the unannounced content is read from the cached continuously generated content.

[0098] In one possible implementation, when the unannounced content is generated by the large language model based on the first user instruction and the announced content, the first user instruction and / or the announced content are cached when the large language model interrupts the generation of the response content.

[0099] In another possible implementation, the unannounced content is generated by the large language model based on the first user instruction, the context of the first user instruction, and the announced content, and the context of the first user instruction is cached when the large language model interrupts the generation of the response content.

[0100] Please refer to Figures 5 to 7, which show an example of a question-and-answer method. It should be noted that, in order to facilitate the description of the technical solution of the embodiment of the present application, in the embodiment of the present application, a question asked by a user and a corresponding answer of the terminal device are defined as a round. The focus of the embodiment of the present application is on the interruption and resumption of a certain broadcast. Therefore, the round in which the user sends an instruction to generate a broadcast and the terminal device generates a corresponding broadcast is defined as a generation round; the round in which the user sends an instruction to resume the broadcast and the terminal device resumes the broadcast is defined as a recovery round; and one or more dialogue rounds in between are defined as secondary rounds.

[0101] Please refer to Figure 5, which shows a flowchart of a first example of a question-answering method. As shown in Figure 5, it is assumed that some functional components used for information detection and information broadcasting in the terminal device are transceiver modules, and some functional components used for voice processing in the terminal device or cloud device are processing modules. The question-answering method of the first example of the embodiment of the present application mainly includes the following steps:

[0102] Step S501: The transceiver module detects a first user instruction, which may be in the form of audio such as voice or recording.

[0103] Step S502: The transceiver module sends a first user instruction to the processing module.

[0104] Step S503: The processing module performs speech processing and generates a first response content through the large language model. Optionally, when generating the first response content, the first user instruction and the generated portion of the first response content are recorded.

[0105] Step S504: The processing module sends the first response content to the transceiver module.

[0106] Step S505: The transceiver module broadcasts the first response content.

[0107] In the embodiment of the present application, steps S501 to S505 are also referred to as a generation round. In the generation round, the user generates a first user instruction, the transceiver module in the terminal device picks up the user's first user instruction, and the processing module performs voice processing to generate a first response content.

[0108] Step S506: The transceiver module detects a second user instruction, which may be in the form of audio such as voice or recording.

[0109] Step S507: The transceiver module sends a second user instruction to the processing module.

[0110] Optionally, in step S508, the processing module performs voice processing to determine the user's intention. If the user's intention is in the intention whitelist, the first response content is paused and the second user instruction is executed to generate the second response content.

[0111] In the embodiment of the present application, there is also a step of judging the user's intention in the generation round and the recovery round, the purpose of which is to judge whether the user's intention is a clear intention in the whitelist. When the user's intention is judged to be a clear intention, the transceiver module performs voice broadcast or stops broadcasting.

[0112] In step S509, the processing module sends a pause broadcast instruction to the transceiver module. Optionally, the processing module further sends a second response content to the transceiver module.

[0113] Step S510: The transceiver module pauses broadcasting the first response content. Optionally, the transceiver module broadcasts the second response content.

[0114] Optionally, in step S511, a suspension processing event is reported.

[0115] Step S512: suspending the generation of the first response content. Optionally, recording the first user instruction and the generated portion of the first response content.

[0116] Optionally, recording the first user instruction may be recording all the information or key information of the first user instruction; key information is relevant information that can be used to generate the first response content, such as "Internet", "development trend", etc. as shown in Figure 3. Recording the generated part of the first response content may be text information or attribute information of the generated part of the first response content; text information refers to the generated text content; attribute information refers to information such as the position, index, and number of words of the generated part. For example, as shown in Figure 3, the text information is "Based on the current development situation, several possible trends in the development of the Internet include: First, the widespread application of artificial intelligence"; the attribute information is: the stop position is "...widely applied", the position of the first trend analysis, thirty-six words have been generated, etc.

[0117] Optionally, based on the characteristic of streaming broadcasting that it is generated and broadcasted at the same time, the steps of generating, sending and broadcasting the first response content (i.e., steps S503 to S505) are performed simultaneously. Therefore, the ungenerated content and the unbroadcasted content in the first response content are the same. Furthermore, the generated content and the broadcasted content in the first response content are also the same. Additionally, the first user instruction and the generated part of the first response content can be recorded verbatim during the process of generating, sending and broadcasting the first response content; the first user instruction and the generated part of the first response content can also be recorded when reporting the pause processing event.

[0118] Optionally, steps S506 to S512 are also referred to as the secondary round. In this secondary round, the user generates a second user instruction. The transceiver module in the terminal device receives the second user instruction, and the processing module performs voice processing and determines the user's intention. If the user's intention is determined to be clear, the transceiver module performs voice broadcast or stops broadcasting.

[0119] Optionally, the secondary round is one or more dialogue rounds that interrupt the generation round. FIG3 illustrates an application scenario in which the secondary round is one dialogue round. In other scenarios, the secondary round may also be multiple dialogue rounds.

[0120] In the embodiment of the present application, the specific content of the second round is not limited, and it can be the instruction of "turn on the air conditioner" as shown in Figure 3, or any other instruction.

[0121] Step S513: The transceiver module detects a third user instruction.

[0122] Step S514: the transceiver module sends a third user instruction.

[0123] Optionally, in step S515, the processing module continues the speech processing and determines the user's intention. The determination of the user's intention is the same as that in step S508, except for the determination object and the determination result, which will not be described in detail here.

[0124] Step S516: When the user intends to continue broadcasting the first response content, the processing module obtains the ungenerated portion of the first response content based on the recorded first user instruction and the generated portion of the first response content.

[0125] For example, based on the recorded first user instruction, the processing module regenerates the entire first response content. Based on the generated portion of the first response content, the processing module can obtain the ungenerated portion of the first response content, that is, the content to be broadcasted in the recovery round.

[0126] Step S517: The processing module sends the ungenerated portion of the first response content to the transceiver module.

[0127] Step S518: The transceiver module continues to broadcast the ungenerated portion of the first response content.

[0128] In the embodiment of the present application, steps S513 to S518 are also referred to as a recovery round. In the recovery round, the user generates a third user instruction, the transceiver module in the terminal device picks up the user's third user instruction, and the processing module performs voice processing and determines the user's intention. When it is determined that the user's intention is to continue broadcasting the first answer content, the processing module obtains the ungenerated part of the first answer content based on the recorded first user instruction and the generated part of the first answer content, and the transceiver module broadcasts the ungenerated part of the first answer content. Among them, the user's intention to continue broadcasting the first answer content is a clear intention.

[0129] In this embodiment of the present application, the steps of the generation round, secondary round, and recovery round can be performed alternately. For example, when the user begins to generate a second user instruction, step S506 and step S505 are performed simultaneously, that is, the work of detecting the second user instruction and the work of broadcasting the first response content are performed simultaneously. In other words, the second user instruction is detected while the first response content is broadcast, and the broadcasting of the first response content is not stopped until the complete second user instruction is detected and the user's intention is determined to be clear.

[0130] Please refer to FIG6a and FIG6b, which show schematic diagrams of a question-answering system.

[0131] As shown in Figure 6a, in one specific implementation, both the transceiver module and the processing module are deployed on the terminal device. The transceiver module includes a voice assistant and an automatic speech recognition (ASR) module. The voice assistant has voice pickup and voice broadcast functions. The ASR module has a language processing function for processing user voice and obtaining user instructions. The processing module includes a dialogue manager (DM) module, a large language model, and a natural language understanding (NLU) module.

[0132] As shown in Figure 6b, in another specific implementation, the transceiver module is deployed in the terminal device, and the processing module is deployed in the cloud device. The implementation of the transceiver module and the processing module is the same as Figure 6a, and will not be repeated here. In addition, a transmission channel (Access Services / kit, AS / kit) module is also included between the transceiver module and the processing module. Among them, the AS / kit module is used for the channel of signal transmission between the terminal device and the cloud, which is suitable for the scenario where a large language model is deployed on the server. The AS module refers to the cloud input service, and the kit module refers to the toolkit of the terminal device, thereby realizing signal conversion and transmission between the cloud device and the terminal device.

[0133] For example, when the user says the wake-up word "Xiaoyi Xiaoyi" to wake up the voice assistant, after the user says the voice command "Analyze the development trend of the Internet?" or "Turn on the air conditioner", the voice command goes through the ASR module, AS / kit module, DM module, large language model or NLU module and other processes, and the voice assistant outputs the response content of the voice command. Among them, the DM module judges the voice command and inputs different voice commands into the large language model or NLU module. The large language model is used to generate the response content of the voice command, and the NLU module is used to judge the user's intention, etc. Optionally, in order to improve the efficiency of voice interaction, the voice command can be sent to the large language model and the NLU module at the same time, and the large language model or the NLU module will make a judgment and analysis based on the voice command and proceed to the next execution step.

[0134] Please refer to Figure 7, which shows a schematic diagram of the first example of a method flow in a question-answering system. As shown in Figure 7, it mainly includes a generation round, a secondary round, and a recovery round. In the generation round, the user's question and the generated content are recorded.

[0135] Optionally, in the next round, it is necessary to determine the user's intention to determine whether to interrupt or resume the broadcast. Among them, through whitelist management, when the user's intention is in the whitelist, the user's intention instruction is executed, and the pause event is reported to the cloud side at the same time, pausing the generation of the streaming broadcast. In the recovery round, the large language model is called to continuously generate subsequent follow-up content through the recorded user questions and the generated content. The large language model supports the ability to generate follow-up content through user questions and the generated content. It can be further understood that clarifying user intentions can also be done in other dialogue rounds.

[0136] In this embodiment of the present application, the DM module records the conversations during the generation round so that the recovery round can be broadcasted and restored based on the recorded conversations. For example, the module can record the user's question and the generated content. For example, the user's question might be "Analyze the development trends of the Internet," and the generated content might be "Based on current developments, several possible trends in Internet development include: First, the widespread application of artificial intelligence."

[0137] Optionally, the DM module also records the context of the generation round to more accurately achieve broadcast recovery. Exemplary context can be the conversation round, whether the user's car is parked, the user's historical questions, the user's profile, etc. For example, after the user interrupts the broadcast, the user has multiple conversations. At this time, the context may have changed, such as the user's historical questions have changed or the user profile has been updated. Therefore, when the recovery round resumes the conversation, the unannounced content after the breakpoint can be generated based on the broadcasted conversation content and the recorded generation round context, so that the continuous unannounced content is more in line with the user's expectations of the generation round, ensuring the consistency of the response content.

[0138] As shown in Figure 7, in the generation round, the user initiates a conversation by waking up the voice assistant. During the first conversation, the voice assistant can record the user's voice, obtain the first audio and the corresponding first context, and send the first audio and first context to the ASR module for speech recognition. The ASR module receives the first audio and first context from the voice assistant, performs speech recognition, and obtains the corresponding first user command. The ASR module sends the first user command and the first context to the AS / kit module. The ASR module also performs voice activity detection (VAD) and sends the first event text corresponding to the VAD event to the AS / kit module. The AS / kit module sends the first user command, the first event text, and the first context to the DM module. The DM module receives the first user command, the first event text, and the first context and sends the first user command to the large language model. The large language model generates the first response content based on the first user command and sends the first response content to the DM module. The characteristic of streaming generation is that the response content is generated and sent simultaneously. After the command is issued, the DM module sends the first response content to the voice assistant. The voice assistant streams the first response content. Streaming broadcasting means that the large language model generates the first response content while the voice assistant broadcasts the first response content.

[0139] As shown in Figure 7, in the second round, there are two scenarios: in the first scenario, the user starts one or more conversations after the first answer content is broadcast. This scenario does not involve the issue of interrupting the streaming broadcast, and the embodiments of this application will not be repeated; in the second scenario, the user starts one or more conversations before the first answer content is broadcast. At this time, the broadcast of the first answer content needs to be interrupted before a new conversation can be started.

[0140] Specifically, in the second round of VAD event processing, the voice assistant sends the VAD event's audio to the ASR module. The ASR module performs speech recognition on the VAD event's audio and sends the recognized text to both the device and the cloud. Specifically, the recognized text is sent to the DM module via the AS / kit module, interrupting the streaming generation of the large language model. Optionally, the voice assistant also sends the recognized text for display.

[0141] Optionally, in the second round, after the end and cloud send the message simultaneously, the AS / kit module sends the text of the VAD event to the DM module. The DM module sends the processed VAD event text to the NLU module. The NLU module determines the user intent based on the VAD event text and sends the user intent to the DM module. Based on the preset intent whitelist, the DM module determines whether the user intent of the VAD event is in the intent whitelist. When the user intent is in the intent whitelist, the DM module sends an instruction to interrupt the streaming broadcast to the voice assistant. The voice assistant pauses the broadcast of the first response content and reports the pause streaming generation event to the DM module. The DM module controls the large language model to pause the generation of the first response content. At this point, the DM module records the generated content in the first response content.

[0142] As shown in Figure 7, during the recovery round, the ASM module detects VAD events. The detailed process is omitted here. The NLU module determines the user's intent based on the VAD event. If it determines that the user intends to continue the interrupted content, it uses the large language model to continuously generate subsequent follow-up content based on the generated round of conversation, such as the recorded user question and the generated content. In other words, the large language model has the ability to generate follow-up content based on the user question and the generated content.

[0143] Please refer to Figures 8 and 9, which show another example of a question-and-answer method. It should be noted that this example implements the technical solution of this example based on the question-and-answer system shown in Figures 6a or 6b. The secondary round is the same as the examples in Figures 5 to 7, and the parts of the generation round and recovery round that are the same as the examples in Figures 5 to 7 are not repeated here. In other words, this example only describes and explains the parts of the generation round and recovery round that are different from the examples in Figures 5 to 7.

[0144] Please refer to Figure 8, which shows a flow chart of the second example of a question-and-answer method. As shown in Figure 8, compared with the first example, the processing module of this example does not record the first user instruction and the generated part of the first answer content, but continues to generate the unbroadcasted content of the first answer content and cache it after pausing the broadcast. Optionally, the unbroadcasted content that continues to be generated can be cached to the transceiver module, that is, the end side. When it is judged that the user intends to continue the interrupted broadcast content, the cached content can be broadcast. The question-and-answer method of the second example of the embodiment of the present application mainly includes the following steps:

[0145] Step S801: The transceiver module detects a first user instruction, which may be in the form of audio such as voice or recording.

[0146] Step S802: The transceiver module sends a first user instruction to the processing module.

[0147] Step S803: The processing module generates a first response content through the large language model.

[0148] Step S804: The processing module sends the first response content to the transceiver module.

[0149] Step S805: The transceiver module broadcasts the first response content.

[0150] Step S806: The transceiver module detects a second user instruction, which may be in the form of audio such as voice or recording.

[0151] Step S807: The transceiver module sends a second user instruction to the processing module.

[0152] Optionally, in step S808, the processing module performs voice processing to determine the user's intention. If the user's intention is in the intention whitelist, the first response content is paused and the second user instruction is executed to generate the second response content.

[0153] In step S809, the processing module sends a pause broadcast instruction to the transceiver module. Optionally, the processing module further sends a second response content to the transceiver module.

[0154] Step S810: The transceiver module pauses broadcasting the first response content. Optionally, the transceiver module also broadcasts the second response content.

[0155] In step S811, the large language model generates the first response content without interruption, sends the continuously generated content, and caches the continuously generated content. That is, step S803 continues to be executed, the generation of the first response content is not interrupted, and the unbroadcasted content of the first response content is cached. The embodiment of the present application does not limit the cache location. For example, the unbroadcasted content can be stored on the terminal side so that when the unbroadcasted content is continued to be broadcast, the unbroadcasted content can be read directly from the terminal side, reducing the signal transmission path, improving the working efficiency of the terminal device, and improving the user experience.

[0156] Step S812: The transceiver module detects a third user instruction.

[0157] Step S813: The transceiver module sends a third user instruction.

[0158] Optionally, in step S814, the processing module continues voice processing to determine the user's intention.

[0159] In step S815 , when the user intends to continue broadcasting the first response content, the processing module obtains the unbroadcasted content of the first response content cached on the terminal side, and sends an instruction to continue broadcasting to the transceiver module.

[0160] Step S816: The transceiver module continues to broadcast the unbroadcasted content of the first response content.

[0161] That is, during the generation round, the context, questions, and generated content for that round are no longer cached separately. Instead, when the broadcast is interrupted, the generation round continues to generate the remaining content to be broadcast and caches the remaining content. Therefore, when the third user instruction for the resumption round is sent, the client directly broadcasts the remaining cached content.

[0162] Please refer to Figure 9, which shows a schematic diagram of the method flow of the second example under a question-answering system. As shown in Figure 9, the steps within the box corresponding to label A represent the main differences between this example and the first example. The processing module does not interrupt the generation of the first content by the large language model, but obtains and caches the unbroadcast content after the first content is completely generated. The embodiment of the present application does not limit the cache location. Exemplarily, the unbroadcast content is sent to the terminal device, and the unbroadcast content is cached in the terminal device. When in the recovery round, the user sends a user instruction indicating the resumption of the interrupted broadcast content, the unbroadcast content cached by the terminal device is read for broadcast.

[0163] Based on the same concept as the aforementioned embodiment, another question-and-answer method is provided in the embodiment of the present application.

[0164] Referring to FIG. 10 , this application provides another question-answering method, which is applied to a question-answering system. The question-answering system includes a terminal device and a cloud device, wherein the terminal device and the cloud device are in communication with each other, and a large language model is deployed on the cloud device. The method mainly includes the following steps:

[0165] Step S1001: The terminal device detects a first user instruction and sends the first user instruction to the cloud device.

[0166] In one possible implementation, in response to the user input content being in the form of audio, the method further includes: the terminal device performing voice analysis on the user input audio to obtain a first user instruction and / or a second user instruction and / or a third user instruction; or, in response to the user input content being in the form of text, the method further includes: based on preset processing rules, the terminal device processes the text input by the user to obtain a first user instruction and / or a second user instruction and / or a third user instruction.

[0167] In step S1002, after receiving the first user instruction, the cloud device generates response content through a large language model.

[0168] Step S1003: The terminal device receives and broadcasts the response content. During the broadcasting of the response content, the terminal device detects a second user instruction and sends the second user instruction to the cloud device.

[0169] Step S1004: After receiving the second user instruction, the cloud device generates an interruption broadcast request and sends the interruption broadcast request to the terminal device.

[0170] In one possible implementation, the cloud device receives the second user instruction and generates the interrupt broadcast request, which includes: the cloud device receives the second user instruction, and generates the interrupt broadcast request after determining that the user intention corresponding to the second user instruction meets a preset condition.

[0171] In another possible implementation, the preset condition is a preset whitelist, and determining whether the user intention corresponding to the second user instruction meets the preset condition includes: determining that the user intention corresponding to the second user instruction is an intention within the preset whitelist.

[0172] In another possible implementation, determining whether the user intention corresponding to the second user instruction meets the preset conditions includes: based on the second user instruction, determining the user intention corresponding to the second user instruction through a natural language understanding model; and judging whether the user intention meets the preset conditions.

[0173] Step S1005 , sending a processing instruction to the large language model, the processing instruction instructing the large language model to interrupt the generation of the response content, or the processing instruction instructing the large language model to continue the generation of the response content, so that the terminal device or cloud device caches the content that continues to be generated.

[0174] Step S1006: The terminal device receives the request to interrupt the broadcast, and interrupts the broadcast of the response content.

[0175] In step S1007, the terminal device detects the third user instruction and sends the third user instruction to the cloud device, where the third user instruction instructs to continue broadcasting the response content.

[0176] In step S1008-1, the cloud device receives a third user instruction and, based on the first user instruction and the generated reported content, determines the unreported content in the response content using a large language model, or retrieves the unreported content in the response content from the cached continuously generated content. The unreported content is then sent to the terminal device.

[0177] Step S1008-2: The terminal device reads the unbroadcasted content in the response content from the cached continuously generated content.

[0178] Among them, steps S1008-1 and S1008-2 provide two methods for determining the unannounced content. That is, when a third user instruction to continue to broadcast the response content is detected, step S1008-1 or step S1008-2 is executed to determine the unannounced content.

[0179] Step S1009: The terminal device broadcasts the unbroadcasted content.

[0180] In one possible implementation, the method further includes: when the large language model interrupts the generation of response content, the terminal device or cloud device caches the first user instruction and / or the broadcast content, so that the large language model generates the unbroadcast content based on the first user instruction and the broadcast content.

[0181] In another possible implementation, the method further includes: when the large language model interrupts the generation of response content, the terminal device or cloud device caches the context of the first user instruction, so that the large language model generates the unannounced content based on the first user instruction, the context of the first user instruction and the announced content.

[0182] Based on the same concept as the aforementioned embodiment, the present application also provides a question-and-answer method, which is applied to a cloud device, and the cloud device is deployed with a large language model. The question-and-answer method includes: using the large language model, based on the user's first user instruction, generating corresponding answer content, and the generated answer content can be broadcast to the user; in the process of the answer content being broadcast, when a second user instruction is received, an interruption broadcast request is generated to indicate the interruption of the broadcast of the answer content, and a processing instruction is sent to the large language model to cause the large language model to interrupt or continue the generation of the answer content. When the large language model is instructed to continue the generation of the answer content, the continued generated content is cached, and the interruption broadcast request is used to interrupt the broadcast of the answer content; receiving a third user instruction, when the third user instruction indicates that the user intends to continue broadcasting the answer content, based on the first user instruction and the broadcast content generation, using the large language model, determining the unbroadcasted content in the answer content, or reading the cached continued generated content to obtain the unbroadcasted content, so that the unbroadcasted content is broadcast. Therefore, the purpose of interrupting and continuing the broadcast process of the answer content is achieved.

[0183] In another possible implementation, the method further includes: when the large language model interrupts generation of the response content, the first user instruction and / or the announced content are cached. Thus, the large language model can generate the unannounced content based on the cached first user instruction and the announced content.

[0184] In another possible implementation, the method further includes caching the context of the first user instruction when the large language model interrupts generation of the response content. Thus, the large language model can generate the unannounced content based on the first user instruction, the context of the first user instruction, and the announced content.

[0185] In another possible implementation, user instructions for interrupting the broadcast are restricted. Further, the user intention corresponding to the second user instruction is determined, and the broadcast interruption request is generated only when the user intention meets a preset condition.

[0186] In another possible implementation, the preset condition is implemented by presetting a whitelist, determining the user intention corresponding to the second user instruction, and interrupting the broadcast of the response content when the user intention is within the preset whitelist.

[0187] In another possible implementation, a natural language understanding model is used to analyze and obtain the user intention corresponding to the second user instruction.

[0188] In another possible implementation, when the user input content is in the form of audio, speech analysis is performed on the user input audio to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction. When the user input content is in the form of text, text processing is performed based on preset processing rules to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction.

[0189] Based on the same concept as the aforementioned embodiment, the present application also provides a question-and-answer method, which is applied to a terminal device, and the question-and-answer method includes: the terminal device detects a first user instruction, and sends the first user instruction to the cloud device, wherein the first user instruction is used to instruct the cloud device to generate response content based on the large language model; the terminal device receives and broadcasts the response content sent by the cloud device, and in the process of broadcasting the response content, the terminal device detects a second user instruction for instructing the cloud device to generate an interrupt broadcast request, and sends the second user instruction to the cloud device; thereafter, the terminal device receives the interrupt broadcast request sent by the cloud device, and interrupts the broadcast of the response content; the terminal device detects a third user instruction, and sends the third user instruction to the cloud device, and the third user instruction instructs to continue broadcasting the unbroadcasted content in the response content; finally, the terminal device receives the unbroadcasted content in the above-mentioned response content sent by the cloud device, and broadcasts the unbroadcasted content.

[0190] In another possible implementation, when the user input content is in the form of audio, speech analysis is performed on the user input audio to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction. When the user input content is in the form of text, text processing is performed based on preset processing rules to obtain the corresponding first user instruction and / or second user instruction and / or third user instruction.

[0191] Based on the same concept as the aforementioned embodiments, the present application provides a terminal device, which is used to execute the method provided in the embodiments of the present application.

[0192] Based on the same concept as the aforementioned embodiment, the present application provides a cloud device, which is deployed with a large language model. The cloud device is communicatively connected to the terminal device, and the cloud device is used to execute the method provided in the embodiment of the present application.

[0193] Based on the same concept as the aforementioned embodiments, the present application provides a question-and-answer system, which includes a terminal device and a cloud device. The terminal device is communicatively connected to the cloud device, and the cloud device is deployed with a large language model. The terminal device is used to execute the method provided in the embodiments of the present application, and / or the cloud device is used to execute the method provided in the embodiments of the present application, and / or the question-and-answer system is used to execute the method provided in the embodiments of the present application.

[0194] Based on the same concept as the aforementioned embodiments, this application provides a chip comprising a processing circuit and a storage medium, wherein the storage medium stores computer program code; when the computer program code is executed by the processing circuit, the method provided in the embodiment of the application is implemented. Based on the same concept as the aforementioned embodiments, this application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiment of the application is implemented.

[0195] Based on the same concept as the aforementioned embodiments, the present application provides a computer program product, which includes program instructions. When the program instructions are executed by a computer, the computer executes the method provided by the embodiments of the present application.

[0196] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0197] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0198] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.

Claims

1. A question-answering method, characterized in that: Applied to a terminal device, the terminal device is deployed with a large language model, and the method includes: reporting a response content, wherein the response content is generated by the large language model based on the first user instruction; In the process of broadcasting the response content, when a second user instruction is detected, a request for interrupting the broadcasting is generated and a processing instruction is sent to the large language model, wherein the processing instruction instructs the large language model to interrupt the generation of the response content, or the processing instruction instructs the large language model to continue the generation of the response content to cache the content that continues to be generated; In response to the interruption broadcast request, interrupting the broadcast of the response content; A third user instruction is detected, and unannounced content in the response content is broadcast, wherein the third user instruction instructs to continue broadcasting the response content, and the unannounced content is generated by the large language model based on the first user instruction and the announced content, or the unannounced content is read from the cached content that continues to be generated.

2. The method according to claim 1, characterized in that When the unannounced content is generated by the large language model based on the first user instruction and the announced content, the first user instruction and / or the announced content are obtained after being cached when the large language model interrupts the generation of the response content.

3. The method according to claim 1 or 2, characterized in that: The unannounced content is generated by the large language model based on the first user instruction, the context of the first user instruction and the announced content, and the context of the first user instruction is obtained after being cached when the large language model interrupts the generation of the response content.

4. The method according to any one of claims 1 to 3, characterized in that: When the second user instruction is detected, generating an interruption broadcast request and sending a processing instruction to the large language model includes: When the second user instruction is detected, in response to the second user instruction, after determining that the user intention corresponding to the second user instruction meets a preset condition, the interrupt broadcast request is generated and the processing instruction is sent to the large language model.

5. The method according to claim 4, characterized in that The preset condition is a preset whitelist. After determining that the user intention corresponding to the second user instruction is an intention within the preset whitelist, the interrupt broadcast request is generated and a processing instruction is sent to the large language model.

6. The method according to claim 4 or 5, characterized in that: The determining that the user intention corresponding to the second user instruction satisfies a preset condition includes: Based on the second user instruction, determining the user intention corresponding to the second user instruction through a natural language understanding model; Determine whether the user intention meets the preset condition.

7. The method according to any one of claims 1 to 6, characterized in that: The first user instruction and / or the second user instruction and / or the third user instruction are obtained based on speech analysis of audio input by the user, or based on processing of text input by the user according to preset processing rules.

8. A question-answering method, characterized in that: Applied to a question-answering system, the question-answering system includes a terminal device and a cloud device, the terminal device is in communication connection with the cloud device, and the cloud device is deployed with a large language model, the method includes: The terminal device detects a first user instruction and sends the first user instruction to the cloud device; After receiving the first user instruction, the cloud device generates response content through the large language model; The terminal device receives and broadcasts the response content. During the broadcasting of the response content, the terminal device detects a second user instruction and sends the second user instruction to the cloud device; After receiving the second user instruction, the cloud device generates an interrupt broadcast request, sends the interrupt broadcast request to the terminal device, and sends a processing instruction to the large language model, wherein the processing instruction instructs the large language model to interrupt the generation of the response content, or the processing instruction instructs the large language model to continue the generation of the response content so that the cloud device or the terminal device caches the content that continues to be generated; The terminal device receives the interruption broadcast request and interrupts the broadcast of the response content; The terminal device detects a third user instruction, and sends the third user instruction to the cloud device, wherein the third user instruction instructs to continue broadcasting the response content; The cloud device receives the third user instruction, determines the unannounced content in the response content based on the first user instruction and the reported content through the large language model, and sends the unannounced content to the terminal device; or the cloud device or the terminal device reads the unannounced content in the response content from the cached continuously generated content; The terminal device broadcasts the unbroadcasted content.

9. The method according to claim 8, characterized in that Also includes: When the large language model interrupts the generation of the response content, the cloud device or the terminal device caches the first user instruction and / or the broadcast content, so that the large language model generates the unbroadcast content based on the first user instruction and the broadcast content.

10. The method according to claim 8 or 9, characterized in that: Also includes: When the large language model interrupts the generation of the response content, the cloud device or the terminal device caches the context of the first user instruction, so that the large language model generates the unannounced content based on the first user instruction, the context of the first user instruction and the announced content.

11. The method according to any one of claims 8 to 10, characterized in that: After receiving the second user instruction, the cloud device generates an interrupt broadcast request including: The cloud device receives the second user instruction, and after determining that the user intention corresponding to the second user instruction meets a preset condition, generates the interrupt broadcast request.

12. The method according to claim 11, characterized in that The preset condition is a preset whitelist, and determining that the user intention corresponding to the second user instruction meets the preset condition includes: Determine that the user intention corresponding to the second user instruction is an intention within the preset whitelist.

13. The method according to claim 11 or 12, characterized in that: Determining that the user intention corresponding to the second user instruction meets a preset condition includes: Based on the second user instruction, determining a user intent corresponding to the second user instruction through a natural language understanding model; Determine whether the user intention meets a preset condition.

14. The method according to any one of claims 8 to 13, characterized in that: The first user instruction and / or the second user instruction and / or the third user instruction are obtained based on speech analysis of audio input by the user, or based on processing of text input by the user according to preset processing rules.

15. A question-answering method, characterized in that: Applied to a cloud device, the cloud device is deployed with a large language model, and the method includes: In response to the first user instruction, generating response content through the large language model; In the process of the answer content being broadcasted, in response to a second user instruction, generating an interruption broadcast request, and sending a processing instruction to the large language model, wherein the processing instruction instructs the large language model to interrupt the generation of the answer content, or the processing instruction instructs the large language model to continue the generation of the answer content, so that the continuously generated content is cached, and the interruption broadcast request is used to interrupt the broadcasting of the answer content; In response to a third user instruction, based on the first user instruction and the reported content generation, the unreported content in the response content is determined through the large language model, or the unreported content in the response content is read from the cached continuously generated content so that the unreported content is reported, wherein the third user instruction instructs to continue to broadcast the response content.

16. The method according to claim 15, characterized in that Also includes: When the large language model interrupts the generation of the response content, the first user instruction and / or the announced content are cached, so that the large language model generates the unannounced content based on the first user instruction and the announced content.

17. The method according to claim 15 or 16, characterized in that Also includes: When the large language model interrupts the generation of the response content, the context of the first user instruction is cached, so that the large language model generates the unannounced content based on the first user instruction, the context of the first user instruction and the announced content.

18. The method according to any one of claims 15 to 17, characterized in that: The generating of the interrupt broadcast request in response to the second user instruction comprises: In response to the second user instruction, after determining that the user intention corresponding to the second user instruction meets a preset condition, the interruption broadcast request is generated.

19. The method according to claim 18, characterized in that The preset condition is a preset whitelist, and determining that the user intention corresponding to the second user instruction meets the preset condition includes: Determine that the user intention corresponding to the second user instruction is an intention within the preset whitelist.

20. The method according to claim 18 or 19, characterized in that Determining that the user intention corresponding to the second user instruction meets a preset condition includes: Based on the second user instruction, determining a user intent corresponding to the second user instruction through a natural language understanding model; Determine whether the user intention meets a preset condition.

21. The method according to any one of claims 15 to 20, characterized in that: The first user instruction and / or the second user instruction and / or the third user instruction are obtained based on speech analysis of audio input by the user, or based on processing of text input by the user according to preset processing rules.

22. A question-answering method, characterized in that: Applied to a terminal device, the method comprises: The terminal device detects a first user instruction, and sends the first user instruction to the cloud device, wherein the first user instruction is used to instruct the cloud device to generate response content based on the large language model; The terminal device receives and broadcasts the response content sent by the cloud device. During the process of broadcasting the response content, the terminal device detects a second user instruction and sends the second user instruction to the cloud device, wherein the second user instruction is used to instruct the cloud device to generate an interruption broadcast request; The terminal device receives the interruption broadcast request sent by the cloud device, and interrupts the broadcast of the response content; The terminal device detects a third user instruction, and sends the third user instruction to the cloud device, wherein the third user instruction instructs to continue to broadcast the unbroadcasted content in the response content; The terminal device receives the unannounced content in the response content sent by the cloud device, and announces the unannounced content.

23. The method according to claim 22, characterized in that The first user instruction and / or the second user instruction and / or the third user instruction are obtained based on speech analysis of audio input by the user, or based on processing of text input by the user according to preset processing rules.

24. A terminal device, characterized in that: The terminal device is used to execute the method according to any one of claims 1-7, or to execute the method according to any one of claims 22-23.

25. A question-answering system, characterized in that: The question-and-answer system includes a terminal device and a cloud device, the terminal device is communicatively connected to the cloud device, the cloud device is deployed with a large language model, the terminal device is used to execute the method as described in any one of claims 22-23, and / or the cloud device is used to execute the method as described in any one of claims 15-21, and / or the question-and-answer system is used to execute the method as described in any one of claims 8-14.

26. A cloud device, characterized in that: The cloud device is deployed with a large language model, the cloud device is in communication connection with the terminal device, and the cloud device is used to execute the method as described in any one of claims 15-21.

27. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 23 is implemented.

28. A chip, characterized in that: The chip includes a processing circuit and a storage medium, wherein the storage medium stores a computer program code; when the computer program code is executed by the processing circuit, the method according to any one of claims 1 to 23 is implemented.

29. A computer program product, characterized in that The computer program product comprises program instructions, and when the program instructions are executed by a computer, the computer is enabled to perform the method according to any one of claims 1 to 23.

Citation Information

Patent Citations

  • Question and answer method and related device

    CN120216621A

  • Voice interaction method and system, electronic equipment and storage medium

    CN115148205A

  • Voice control method and device, electronic equipment, vehicle and storage medium

    CN115440214A

  • Voice interaction system and method, electronic equipment and storage medium

    CN116417003A

  • Voice broadcast interruption processing method and device, equipment and storage medium

    CN116975242A

Cited By

  • Low-delay streaming voice interaction system with interruption processing function

    CN120636409A

  • A low-delay streaming voice interaction system with break handling functionality

    CN120636409B