Question answering method, apparatus and system, device, medium, vehicle and program product
By processing user interaction input data in the Q&A system to generate a collection of content information and determining the target service in the Q&A scenario, the existing system's accuracy problem of identifying Q&A instructions in multiple user instructions is solved, and higher response accuracy and user experience are achieved.
Patent Information
- Application Number
- PCT/CN2024/140304
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-26
AI Technical Summary
The existing question and answer system is difficult to accurately identify question and answer instructions when processing multiple user instructions, resulting in low response accuracy and poor presentation of response data.
By receiving the user's interactive input data, the process obtains the generated content information set, including user intention description information, response data and interaction description information, determines the user interaction scenario, and determines the target service in the question and answer scenario to perform subsequent actions.
It improves the accuracy of responding to interactive input data and improves the user's reading experience of response data.
Smart Images

Figure CN2024140304_26062025_PF_FP_ABST
Abstract
Description
Question-answering method, device, system, equipment, medium, vehicle, and program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to Chinese patent application number 2023117439689, filed on December 18, 2023, entitled “Question-answering method, device, system and vehicle,” Chinese patent application number 2023117452999, filed on December 18, 2023, entitled “Question-answering method, device, equipment, storage medium and vehicle,” Chinese patent application number 2023117462929, filed on December 18, 2023, entitled “Question-answering method, device, equipment, storage medium and vehicle,” and Chinese patent application number 2023117462952, filed on December 18, 2023, entitled “Question-answering method, device, equipment, storage medium and vehicle,” and the applicants are all Beijing Rockwell Technology Co., Ltd., and the full texts of the above applications are incorporated into this disclosure by reference. Technical Field
[0003] The present disclosure belongs to the field of intelligent question-answering technology, and in particular relates to a question-answering method, apparatus, system, equipment, medium, vehicle, and program product. Background Art
[0004] With the development of artificial intelligence, question-answering systems are becoming more and more widely used. Generally, question-answering systems can respond to user commands and output and display the response data corresponding to the user commands.
[0005] Currently, question-and-answer systems typically respond to user commands after receiving them. However, in reality, a question-and-answer system may receive multiple user commands at once, and these multiple user commands may include both question-and-answer commands and control commands. In this case, the question-and-answer system may not be able to accurately identify the question-and-answer command from among the multiple user commands, resulting in lower accuracy in responding to user commands. Furthermore, the response data currently displayed is typically plain text, failing to highlight key information, resulting in poor question-and-answer presentation of the response data. Summary of the Invention
[0006] The embodiments of the present disclosure provide a question-answering method, apparatus, system, device, medium, vehicle, and program product. This solution can not only improve the accuracy of responses to interactive input data, but also enhance the user's reading experience of the response data.
[0007] The present disclosure provides a question-answering method, which includes:
[0008] Receive interactive input data from the user;
[0009] Processing the interactive input data to obtain a generated content information set; the generated content information set includes user intention description information, response data and interaction description information;
[0010] Determining a user interaction scenario corresponding to the user intention description information;
[0011] In the case where the user interaction scenario is a question-and-answer scenario, a target service corresponding to the question-and-answer scenario is determined from a plurality of services; the service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interaction input data.
[0012] The present disclosure also provides another question-and-answer method, including:
[0013] Receive user interaction input data; wherein the interaction input data includes multimodal feature data and instruction interaction data, the multimodal feature data represents the input form of the interaction input data, and the instruction interaction data represents the user intention corresponding to the interaction input data;
[0014] Inputting the interactive input data into the artificial intelligence big model so that the artificial intelligence big model recognizes and responds to the interactive input data, thereby obtaining and streaming out a plurality of response data, wherein a later response data includes a previous response data;
[0015] During the process of streaming the plurality of response data, performing a dialogue interaction on the first response data, the dialogue interaction including voice interaction;
[0016] Starting from the second response data, determining the response data other than the previous response data in the response data as the first target response data;
[0017] A dialog interaction is performed on the first target response data until the response data becomes the last one of the multiple response data.
[0018] The present disclosure also provides another question-answering method, including:
[0019] Receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing an input form of the interaction input data, and the instruction interaction data representing a user intention corresponding to the interaction input data;
[0020] Inputting the interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and outputs a plurality of response data;
[0021] In the process of the artificial intelligence large model outputting the plurality of response data, determining a target paragraph based on the plurality of output response data;
[0022] Inputting the target paragraph into the artificial intelligence model to obtain card information corresponding to the target paragraph through the artificial intelligence model;
[0023] Render and display the card information corresponding to the target paragraph.
[0024] The present disclosure also provides a question-answering device, which includes:
[0025] A receiving module configured to receive interactive input data from a user;
[0026] A first processing module is configured to process the interactive input data using an artificial intelligence large model to obtain a generated content information set, wherein the generated content information set includes user intention description information, response data, and interaction description information;
[0027] A first determining module is configured to determine a user interaction scenario corresponding to the user intention description information;
[0028] The second determination module is configured to determine a target service corresponding to the question-and-answer scenario from multiple services when the user interaction scenario is a question-and-answer scenario. The service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interaction input data.
[0029] The present disclosure also provides another question-answering device, including:
[0030] a receiving module configured to receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing an input form of the interaction input data, and the instruction interaction data representing a user intention corresponding to the interaction input data;
[0031] An input module is configured to input the interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, obtains and streams out a plurality of response data, wherein a later response data includes a previous response data;
[0032] A first interaction module is configured to perform a dialogue interaction on a first piece of the response data during the process of streaming the plurality of response data, wherein the dialogue interaction includes a voice interaction;
[0033] a determining module configured to, starting from the second response data, determine the response data other than the previous response data in the response data as the first target response data;
[0034] The second interaction module is configured to perform a dialog interaction on the first target response data until the response data becomes the last one of the multiple response data.
[0035] The present disclosure also provides another question-and-answer device, including:
[0036] a receiving module configured to receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing an input form of the interaction input data, and the instruction interaction data representing a user intention corresponding to the interaction input data;
[0037] A first input module is configured to input the interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and outputs a plurality of response data;
[0038] A determination module is configured to determine a target paragraph based on the outputted plurality of response data during the process of the artificial intelligence large model outputting the plurality of response data;
[0039] A second input module is configured to input the target paragraph into the artificial intelligence model to obtain card information corresponding to the target paragraph through the artificial intelligence model;
[0040] The rendering module is configured to render and display the card information corresponding to the target paragraph.
[0041] The present disclosure also provides a question-answering system, which includes:
[0042] A main interaction module is configured to receive user interaction input data;
[0043] a large model module configured to process the interactive input data through an artificial intelligence large model to obtain a generated content information set, wherein the generated content information set includes user intention description information, response data, and interaction description information;
[0044] a dialogue management module configured to determine a user interaction scenario corresponding to the user intent description information, and to send the user interaction scenario to the service management module, and to determine a scenario identifier corresponding to the response data, and to send the scenario identifier and the corresponding response data to the service management module;
[0045] The service management module is configured to determine a target service assistant corresponding to the user interaction scenario sent by the dialogue management module based on a correspondence between a service assistant and a registration intent, and a correspondence between the registration intent and the user interaction scenario, and to route the response data corresponding to the scenario identifier to the target service assistant based on a correspondence between the scenario identifier and the target service assistant;
[0046] The service assistant is configured to register the user interaction scenario of interest to the service management module, obtain the registration intention, and also to register the user interaction scenario, obtain the scene identifier, and process the response data corresponding to the scene identifier. The service assistant includes the target service assistant.
[0047] An embodiment of the present disclosure further provides an electronic device, the electronic device comprising: a processor and a memory storing computer program instructions;
[0048] When the processor executes the computer program instructions, the question-answering method as described above is implemented.
[0049] An embodiment of the present disclosure further provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, any of the above-described question-and-answer methods is implemented.
[0050] The present disclosure also provides a vehicle, which includes at least one of the following:
[0051] The question-and-answer device as described above;
[0052] As mentioned above, the question answering system;
[0053] Electronic devices as previously mentioned;
[0054] The computer-readable storage medium as described above.
[0055] An embodiment of the present disclosure further provides a computer program, which includes a computer-readable code. When the computer-readable code runs in an electronic device, the processor of the electronic device executes the code to implement any of the above-described question-and-answer methods.
[0056] An embodiment of the present disclosure also provides a computer program product, which includes a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device implements any of the above-described question-and-answer methods when executing the code.
[0057] In the question-answering method of the embodiment of the present disclosure, the artificial intelligence large model can process the interactive input data to obtain a generated content information set including user intention description information, response data and interaction description information. Based on this, since the service corresponds to the user interaction scenario, by determining the user interaction scenario corresponding to the user intention description information, and then executing the subsequent action corresponding to the interactive input data through the service corresponding to the user interaction scenario, the professionalism of executing the subsequent action can be improved, and the accuracy of executing the subsequent action can be improved; by determining the target service corresponding to the question-answering scenario among multiple services when the user interaction scenario is a question-answering scenario, and executing the subsequent action corresponding to the interactive input data through the target service, the accuracy of responding to the interactive input data can be improved. In this way, through the embodiment of the present disclosure, the accuracy of responding to the interactive input data can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0059] FIG1 is a schematic diagram of the structure of a question-answering system provided by an embodiment of the present disclosure;
[0060] FIG2 is a flow chart of a question-answering method provided by an embodiment of the present disclosure;
[0061] FIG3 is a schematic diagram showing a comparison between plain text and TTS structured data provided by an embodiment of the present disclosure;
[0062] FIG4 is a schematic diagram of a process for determining a target paragraph through a dynamic parsing engine according to an embodiment of the present disclosure;
[0063] FIG5 is a schematic diagram of card information provided by an embodiment of the present disclosure;
[0064] FIG6 is a schematic diagram of a voice interaction interface including a TTS display area and a card information display area provided by an embodiment of the present disclosure;
[0065] FIG7 is a schematic diagram of a correspondence between template types and card types provided by an embodiment of the present disclosure;
[0066] FIG8 is a schematic diagram of a process of performing a dialog interaction on sub-response data through a TTS structured engine according to an embodiment of the present disclosure;
[0067] FIG9 is an interactive diagram of a question-and-answer method provided by an embodiment of the present disclosure;
[0068] FIG10 is a flow chart of another question-answering method provided by an embodiment of the present disclosure;
[0069] FIG11 is a schematic diagram of a dynamic parsing process provided by an embodiment of the present disclosure;
[0070] FIG12 is a schematic structural diagram of a question-answering device provided by an embodiment of the present disclosure;
[0071] FIG13 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0072] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0073] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.
[0075] As described in the background technology section, in order to solve the problems of the prior art, the embodiments of the present disclosure provide a question-and-answer method, apparatus, system, device, medium, vehicle, and program product. The question-and-answer system can be set in a terminal server, a cloud server, or a distributed server, without limitation. That is, the execution entity of the question-and-answer method can be a server. The server can be a terminal server, a cloud server, or a distributed server, without limitation. The terminal server can include, for example, servers corresponding to mobile devices such as mobile phones, computers, and tablets, servers corresponding to smart wearable devices such as smart glasses and smart watches, and vehicle servers. Furthermore, the question-and-answer method can be specifically applied to voice question-and-answer scenarios in an in-vehicle environment. That is, the terminal server can be a vehicle server. Furthermore, the execution entity of the question-and-answer method can also include both a terminal server and a cloud server. That is, the question-and-answer system can also have some functional modules set in the terminal server and other functional modules set in the cloud server. The terminal server and the cloud server can communicate via a network.
[0076] The following is an introduction to the question-answering system provided by the embodiments of the present disclosure.
[0077] Figure 1 shows a structural diagram of a question-and-answer system 100 provided in an embodiment of the present disclosure. As shown in Figure 1, the question-and-answer system 100 provided in an embodiment of the present disclosure may include a main interaction module 11, a dialogue management module 12, a large model module 13, a service management module 14, a service assistant 15 and a generative user interface service 16. Among them, the service assistant 15 may include a task-based service assistant 151 and an artificial intelligence (AI) service assistant 152. The AI-based service assistant may be an abbreviation of the AI-based service assistant. The AI-based service assistant 152 may include a text-to-speech (TTS) broadcast engine 1521, a TTS structuring engine 1522, a rich media structuring engine 1523 and a dynamic parsing engine 1524.
[0078] First, the large model module 13 may include an AI large model; wherein, an AI large model refers to a high-performance artificial intelligence model built through huge training samples and computing resources. They can learn a large amount of language knowledge, image features and speech patterns, and can reason and generate outputs similar to humans. They have wide applications in natural language processing, image recognition, speech recognition and other fields. AI large models may include, for example, large language models (LLM), ChatGPT (Chat Generative Pre-trained Transformer), multimodal large models and multimodal cognitive large models. Specifically, the AI large model can process the interactive input data to obtain a generated content information set. The generated content information set may include user intent description information, response data and interaction description information. The response data may include multiple sub-response data output in a streaming manner, wherein the latter sub-response data may include the previous sub-response data. The interaction description information can be used to perform structured processing on the response data.
[0079] In addition, the large model module 13 may also include TaskFormer. TaskFormer can be seen as an improved and expanded model based on GPT (Generative Pre-trained Transformer) for task-based natural language processing problems. It can better handle a variety of tasks and has high flexibility and versatility. In the embodiment of the present disclosure, TaskFormer can be equivalent to the central control of the artificial intelligence large model, and can pre-process and post-process the input and output content of the artificial intelligence large model.
[0080] In addition, the large model module 13 can be set on a terminal server, a cloud server, or a distributed server, which is not limited here.
[0081] The main interaction module 11 may be configured to receive user interaction input data, determine user intention description information corresponding to the interaction input data, and send the user intention description information to the dialogue management module 12 .
[0082] Among them, the interactive input data can be multimodal command interaction data input by the user. That is, the interactive input data can include multimodal feature data and command interaction data. Among them, the multimodal feature data can characterize the input form of the interactive input data. The input form of the interactive input data can include, for example, voice input, text input, touch input, and gesture input. In addition, the command interaction data can characterize the user intention corresponding to the interactive input data. The command interaction data can, for example, include control instructions related to vehicle control and vehicle settings, can include question instructions related to vehicle Q&A and general Q&A, and can also include control instructions and question instructions at the same time, which are not limited here.
[0083] As an example, the main interaction module 11 may have an Automatic Speech Recognition (ASR) function. If the interactive input data is in the form of voice input, the main interaction module may use the ASR function to convert the voice command into a command text and send the command text to the large model module 13. After receiving the command text, the large model module 13 may perform semantic analysis on the command text using the AI large model to obtain a description of the user's intention and return the user's intention description to the main interaction module 11.
[0084] In addition, the question-and-answer system 100 may also include a voice user interface (VUI) corresponding to the main interaction module 11. The voice interaction interface may include a text input box. The user may enter text instructions in the text input box. The main interaction module 11 may determine user intent description information corresponding to the interactive input data based on the correspondence between the text instructions and the user intent.
[0085] In addition, the voice interaction interface may also include multiple interactive controls, which may include preset instructions related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. The user can perform touch commands by clicking on the interactive controls. The main interaction module 11 may determine user intent description information corresponding to the interactive input data based on the correspondence between the touch commands and the user intent.
[0086] In addition, if the voice interaction interface supports gesture air touch, the user can select any preset instruction by gesture air control interaction control. The main interaction module 11 can determine the user intention description information corresponding to the interaction input data based on the correspondence between the gesture instruction and the user intention.
[0087] In addition, the main interaction module 11 is also responsible for building basic voice capabilities, including dialogue types, dialogue execution, interaction modes, user modes, access to the Software Development Kit (SDK) for large models and free dialogues, and maintenance of voice images. The SDK can serve as a general capability for in-vehicle smart cockpits and is widely used in scenarios such as large models, input methods, voice assistants, and gesture recognition.
[0088] The dialogue management module 12 can receive the user intent description information sent by the main interaction module 11 via the SDK. After receiving the user intent description information, the dialogue management module 12 can determine the user interaction scenario corresponding to the user intent description information based on the arbitration strategy and send the user interaction scenario to the service management module 14. User interaction scenarios can include question-and-answer scenarios and interactive control scenarios. The arbitration strategy can be a predefined correspondence between user intent description information and user interaction scenarios.
[0089] The service management module 14 may pre-store multiple correspondences between service assistants 15 and registration intents. Registration intents may correspond one-to-one to user interaction scenarios. A registration intent may represent a user interaction scenario that a service assistant can handle. After receiving a user interaction scenario, the service management module 14 may first determine a target registration intent corresponding to the received user interaction scenario based on the correspondence between the user interaction scenario and the registration intent, and then determine a service assistant corresponding to the target registration intent based on the correspondence between the registration intent and the service assistant.
[0090] It should be noted that, in order for the service management module 14 to determine the correspondence between the service assistant and the registration intention, each service assistant 15 may register the user interaction scenario it is interested in with the service management module 14 in advance to obtain the registration intention.
[0091] In addition, after determining the target service assistant, the service management module 14 can notify the target service assistant to register its corresponding user interaction scenario and obtain a scenario identifier. The service management module 14 can record the correspondence between the target service assistant and the scenario identifier and send the scenario identifier to the dialogue management module 12. After receiving the scenario identifier, the dialogue management module 12 can initiate a streaming data request corresponding to the scenario identifier to the large model module 13. The streaming data request can be used to request the artificial intelligence large model to stream out multiple sub-response data. After receiving the streaming data request, the large model module 13 can stream out multiple sub-response data (i.e., response data) to the dialogue management module 12 through the main interaction module 11. After receiving the response data, the dialogue management module 12 can determine the scenario identifier corresponding to the response data and send the scenario identifier and its corresponding response data to the service management module 14. After receiving the scenario identifier and its corresponding response data, the service management module 14 can route the response data corresponding to the scenario identifier to the target service assistant. After receiving the response data, the target service assistant can process the response data (the response data corresponding to the scenario identifier registered by the target service assistant).
[0092] In addition, the dialogue management module 12 can also be responsible for skill operation, service operation, exception handling operation and exception handling, including skill whitelist configuration, script configuration, expression configuration and exception monitoring.
[0093] Specifically, the service assistant 15 may include a task-based service assistant 151 and an AI-based service assistant 152. Among them, the task-based service assistant 151 can be configured to process tasks corresponding to interactive control scenarios. The specific process of the task-based service assistant 151 processing tasks may include: determining the target application corresponding to the user instruction, and sending the user instruction to the target application so that the target application executes the user instruction. In addition, the AI-based service assistant 152 can be configured to process tasks corresponding to question-and-answer scenarios. The AI-based service assistant can be an auxiliary tool integrated into the voice system that can help users complete various questions and answers. The specific process of the AI-based service assistant 152 processing tasks may include: sending a structured request to the large model module 13. Among them, the structured request can be used to request the artificial intelligence large model to obtain interaction description information corresponding to the response data. After receiving the interaction description information, the AI-based service assistant 152 can perform structured processing on the response data based on the interaction description information, and obtain and display the structured data.
[0094] For the AI-type service assistant 152, the TTS broadcast engine 1521 can be configured to perform dialogue interaction on multiple sub-response data output in a streaming manner. The TTS broadcast engine 1521 can generate speech frame by frame according to the text based on the TTS interrupt broadcast strategy and the additional broadcast strategy, so that speech synthesis becomes a streaming process, which can output sound in real time and play it simultaneously with the generation of speech. This method can make the synthesized speech smoother and reduce delays; the TTS structuring engine 1522 can be configured to perform TTS structuring processing on the complete target response data corresponding to the interactive input data to obtain TTS structured data. The rich media structuring engine 1523 can be configured to perform rich media structuring processing on the target paragraph corresponding to the interactive input data to obtain card information. The dynamic parsing engine 1524 can be configured to dynamically parse multiple sub-response data, target response data and multiple target paragraphs in the process of streaming multiple sub-response data, performing dialogue interaction, performing TTS structuring processing and performing rich media structuring processing.
[0095] The generative user interface service 16 can dynamically display the generated TTS structured data and card information. The generative user interface service 16 can include a card service, a page service, a widget, and a webview container. The webview container is a component commonly used in mobile application development and can be configured to display web content. The webview container provides a browser view that can be embedded in an application, allowing developers to load and display web pages, HTML files, JavaScript applications, etc. through the webview. Furthermore, generative user interface (UI) technology can be composed of a container + template concept. After the large model outputs a display template corresponding to the response data, the container can be used to implement interface rendering. Generative UI technology can be implemented based on converged technologies such as Native, React Native, and HTML5 (H5). H5 is a standard technology for building and presenting web content. It consists of Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), and JavaScript. The emergence and development of H5 technology has brought a richer and more interactive internet experience. The advantages of generative UI lie in its flexibility and scalability. By separating the UI from the data, different UIs can be dynamically generated at runtime based on data changes, which is very suitable for handling dynamic content and variable structure interfaces.
[0096] In addition, the Q&A system can also include a graphical user interface (GUI) control center. The GUI control center can be responsible for overall GUI scheduling, including card management, card creation and destruction, widget data updates, and webview container creation and destruction.
[0097] Based on this, in order to improve the response efficiency of the question-answering system 100. In some embodiments, the main interaction module 11, the dialogue management module 12, the service management module 14, the service assistant 15, and the generative user interface service 16 can be set in the terminal server. The terminal server can include, for example, servers corresponding to mobile devices such as mobile phones, computers, and tablets, servers corresponding to smart wearable devices such as smart glasses and smart watches, and vehicle servers. The large model module 13 can be set in the cloud server. In addition, the large model module 13 can also be deployed in the cloud server, and an offline large model module can be set in the terminal server.
[0098] If the large model module 13 is located on a cloud server, the question-answering system 100 may also include a cloud-based central control / dialogue system. The cloud-based central control / dialogue system may be responsible for managing cloud-based dialogues. The cloud-based central control / dialogue system may be a cloud computing-based service that receives and processes user voice input commands, enabling voice control and dialogue functions.
[0099] Specifically, cloud-based central control can refer to sending user voice-entered commands to a remote cloud server for processing and analysis. The cloud server can use speech recognition technology to convert speech into text, and then understand the user's commands and intentions through technologies such as semantic parsing and intent recognition. Voice dialogue can refer to the ability to communicate and converse through voice. Cloud-based central control can use natural language processing technology to convert user voice commands into dialogue, providing an interactive experience similar to human-computer dialogue. By converting speech into text and utilizing technologies such as semantic understanding and language generation, the cloud-based central control can understand the user's questions or needs and provide corresponding answers or responses.
[0100] The advantage of cloud-based control / dialogue lies in its ability to leverage the powerful computing capabilities of cloud computing and rich voice processing technologies to provide a high-quality and intelligent voice interaction experience. Users can control devices and obtain information through voice input, or communicate with intelligent systems through voice dialogue.
[0101] If the large model module 13 is set up in the cloud server, the main interaction module 11 can also send the user's instruction data to the cloud central control through the SDK.
[0102] The following introduces the question-answering method provided by the embodiment of the present disclosure.
[0103] FIG2 shows a flow chart of a question-answering method provided by an embodiment of the present disclosure. The question-answering method can be executed by a processor in a vehicle. As shown in FIG2 , the question-answering method provided by an embodiment of the present disclosure includes the following steps:
[0104] S210, receiving interactive input data from the user;
[0105] S220: Process the interactive input data to obtain a generated content information set, where the generated content information set includes user intention description information, response data, and interaction description information;
[0106] S230: Determine a user interaction scenario corresponding to the user intention description information;
[0107] S240: When the user interaction scenario is a question-and-answer scenario, determine a target service corresponding to the question-and-answer scenario from multiple services. The service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interaction input data.
[0108] In the question-and-answer method of the embodiment of the present disclosure, the artificial intelligence big model can process the interactive input data to obtain a generated content information set including user intention description information, response data and interaction description information. Based on this, since the service corresponds to the user interaction scenario, by determining the user interaction scenario corresponding to the user intention description information, and then executing the subsequent actions corresponding to the interactive input data through the service corresponding to the user interaction scenario, the professionalism of executing the subsequent actions can be guaranteed, thereby improving the accuracy of executing the subsequent actions. By determining the target service corresponding to the question-and-answer scenario among multiple services when the user interaction scenario is a question-and-answer scenario, and executing the subsequent actions corresponding to the interactive input data through the target service, the accuracy of responding to the interactive input data can be improved. In this way, through the embodiment of the present disclosure, the accuracy of responding to the interactive input data can be improved.
[0109] The specific implementation methods of the above steps are introduced below.
[0110] In some embodiments, in S210, the interactive input data may be multimodal instruction interaction data input by the user. That is, the interactive input data may include multimodal feature data and instruction interaction data. Among them, the multimodal feature data may characterize the input form of the interactive input data. The input form of the interactive input data may include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interaction data may characterize the user intention corresponding to the interactive input data. The instruction interaction data may include, for example, control instructions related to vehicle control and vehicle settings, question instructions related to vehicle Q&A and general Q&A, and may also include both control instructions and question instructions, which are not limited here.
[0111] As an example, the question-answering system may include a main interaction module and a large model module. Among them, the main interaction module may have an ASR function. If the input form of the interactive input data is voice input, the main interaction module may convert the voice command into a command text through the ASR function and send the command text to the large model module. After receiving the command text, the large model module may input the interactive input data into the artificial intelligence large model, perform semantic analysis on the command text through the artificial intelligence large model, obtain user intent description information (i.e., command interaction data), and return the user intent description information to the main interaction module.
[0112] In addition, the question-and-answer system may also include a VUI corresponding to the main interaction module. The voice interaction interface may include a text input box. Users can enter text commands in the text input box. The main interaction module may determine the command interaction data corresponding to the interaction input data based on the correspondence between the text commands and the command interaction data.
[0113] In addition, the voice interaction interface may also include multiple interactive controls, which may include preset commands related to vehicle control, vehicle settings, vehicle Q&A, and general Q&A. Users can perform touch commands by clicking on the interactive controls. The main interaction module may determine the command interaction data corresponding to the interactive input data based on the correspondence between the touch commands and the command interaction data.
[0114] In addition, if the voice interaction interface supports gesture air touch, the user can select any preset command by gesture air control interaction control. The main interaction module can determine the command interaction data corresponding to the interaction input data based on the correspondence between the gesture command and the command interaction data.
[0115] In some embodiments, in S220, if the input form of the interactive input data is voice input, then after receiving the user's interactive input data, the main interactive module can send the interactive input data to the large model module. The large model module can input the interactive input data into the artificial intelligence large model, and the artificial intelligence large model can recognize and respond to the interactive input data. In particular, recognizing the interactive input data can generate user intent description information. Responding to the interactive input data can generate response data corresponding to the interactive input data. If the input form of the interactive input data is non-voice input (including text input, touch input, and gesture input), then after receiving the user's interactive input data, the main interactive module can determine the instruction interaction data corresponding to the interactive input data and send the instruction interaction data to the large model module. The large model module can input the instruction interaction data into the artificial intelligence large model, and the artificial intelligence large model can recognize and respond to the instruction interaction data. In particular, recognizing the instruction interaction data can generate user intent description information. Responding to the instruction interaction data can generate response data corresponding to the instruction interaction data.
[0116] As an example, if the large model module is located on a cloud server, and the main interaction module and dialogue management module are located on a terminal server, then after the terminal server sends interactive input data to the cloud server via the main interaction module, it can continue to send an asynchronous generation interface to the cloud server, causing the large model module to return user intent description information to the terminal server. After receiving the user intent description information, the main interaction module in the terminal server can send the user intent description information to the dialogue management module.
[0117] In addition, after returning the user intention description information to the terminal server, the large model module can continue to obtain response data corresponding to the interactive input data (including instruction interaction data) and cache the obtained response data. Among them, obtaining response data corresponding to the interactive input data includes searching for response data corresponding to the interactive input data in a pre-set question and answer library, searching for response data corresponding to the interactive input data in the network, and generating response data corresponding to the interactive input data. In the process of the artificial intelligence large model generating response data corresponding to the interactive input data, when predicting the next word or character, the previously generated content will be taken into account. This context-aware generation method enables the artificial intelligence large model to generate text with coherence and a certain logic.
[0118] In some embodiments, at S230, after receiving the user intent description information, the dialogue management module may determine the user interaction scenario corresponding to the user intent description information based on an arbitration strategy. User interaction scenarios may include question-and-answer scenarios and interactive control scenarios. Specifically, question-and-answer scenarios may include general question-and-answer scenarios and vehicle question-and-answer scenarios. Interactive control scenarios may include vehicle control scenarios. Furthermore, the arbitration strategy may be a pre-set correspondence between user intent description information and user interaction scenarios.
[0119] As an example, when the dialogue management module determines that the user interaction scenario is a question-and-answer scenario, the dialogue management module may send a streaming data request to the large model module. The streaming data request may be used to request the large model module to stream-out multiple sub-response data corresponding to the interactive input data. That is, the response data may include multiple sub-response data that are streamed out, and after receiving the streaming data request, the large model module may respond to the streaming data request and stream-out multiple sub-response data. The latter sub-response data may include the previous sub-response data. For example, if the interactive input data is "Introduce Zhou xx," the multiple sub-response data that are streamed out may include: the first sub-response data is "Zhou xx," the second sub-response data is "Zhou xx, from xxx," and the third sub-response data is "Zhou xx, from xxx, won x award on x year x month x day..."
[0120] As a more specific example, if the large model module is located in a cloud server, and the main interaction module and the dialogue management module are located in a terminal server, the dialogue management module can send a streaming data request to the large model module through the main interaction module.
[0121] As another example, after determining the user interaction scenario, the dialogue management module may also send the user interaction scenario to the service management module so that the service management module selects the service corresponding to the user interaction scenario from multiple services and performs subsequent actions corresponding to the response data.
[0122] In some embodiments, in S240, the service may correspond to a user interaction scenario and be used to execute subsequent actions corresponding to the interactive input data. For example, if the user interaction scenario is a question-and-answer scenario, the service corresponding to the question-and-answer scenario may be used to conduct a dialogue interaction on a plurality of sub-response data output in a streamed manner, and to structure the response data. Among them, the dialogue interaction may include voice interaction and data display corresponding to the voice interaction. For another example, if the user interaction scenario is an interactive control scenario, the service corresponding to the interactive control scenario may be used to send the user intention description information to the corresponding application, so that the corresponding application executes the control instruction corresponding to the user intention description information. In addition, the target service may be a service corresponding to the question-and-answer scenario.
[0123] As an example, different services can be executed by different service assistants. That is, there can be a one-to-one correspondence between service assistants and services. For example, a target service corresponding to a question-and-answer scenario can be executed by an AI-based service assistant corresponding to the question-and-answer scenario.
[0124] Based on this, in some embodiments, the above S240 may specifically include:
[0125] Obtain registration intents corresponding to multiple services respectively; determine the target service corresponding to the question-and-answer scenario based on the correspondence between the registration intents and the user interaction scenarios.
[0126] Here, the correspondence between multiple service assistants (i.e., services) and registration intentions can be pre-stored in the service management module. Among them, there can be a one-to-one correspondence between registration intent and user interaction scenarios. Registration intent can represent the user interaction scenarios that the service assistant can handle. After receiving the user interaction scenario (which is a question-and-answer scenario), the service management module can first determine the target registration intent corresponding to the question-and-answer scenario based on the correspondence between the user interaction scenario and the registration intent, and then determine the service assistant corresponding to the target registration intent (i.e., AI-type service assistant) based on the correspondence between the registration intent and the service assistant. After determining the AI-type service assistant, the target service corresponding to the AI-type service assistant can be determined.
[0127] For example, after S240 is completed, the following operations may be performed:
[0128] S250: Structuring the response data based on the interaction description information through the target service to obtain and display structured data.
[0129] Through the target service, the response data is structured based on the interaction description information to obtain and display structured data. Compared with displaying plain text, key information can be highlighted, thereby improving the user's reading experience of the response data.
[0130] In some embodiments, in S250, after determining the target service (i.e., the AI-type service assistant), a structured request can be sent to the large model module through the AI-type service assistant. The structured request can be used to request the artificial intelligence large model to obtain the interaction description information corresponding to the response data. After receiving the structured request, the large model module can input the response data into the artificial intelligence large model, and generate the interaction description information corresponding to the response data through the artificial intelligence large model. Afterwards, the large model module can return the interaction description information to the AI-type service assistant. After receiving the interaction description information, the AI-type service assistant can structure the response data based on the interaction description information through the target service to obtain and display the structured data.
[0131] Based on this, in order to implement structured processing of the response data based on the interaction description information through the target service, in some embodiments, the above S250 may specifically include:
[0132] The question-and-answer scenario is registered through the target service to obtain a scenario identifier; the response data corresponding to the scenario identifier is routed to the target service; and the response data is structured based on the interaction description information through the target service to obtain and display structured data.
[0133] Here, after determining the target service (i.e., the AI-based service assistant), the service management module can notify the target service to register the question-and-answer scenario and obtain a scenario identifier. The service management module can record the correspondence between the target service and the scenario identifier and send the scenario identifier to the dialogue management module. After receiving the scenario identifier, the dialogue management module can initiate a streaming data request corresponding to the scenario identifier to the large model module. After receiving the streaming data request, the large model module can stream multiple sub-response data (i.e., response data) to the dialogue management module through the main interaction module. After receiving the response data, the dialogue management module can determine the scenario identifier corresponding to the response data and send the scenario identifier and its corresponding response data to the service management module. After receiving the scenario identifier and its corresponding response data, the service management module can route the response data corresponding to the scenario identifier to the target service. After receiving the response data, the target service can structure the response data based on the interaction description information to obtain and display the structured data.
[0134] In this way, by routing the response data corresponding to the scenario identifier to the target service, the response data can be structured based on the interaction description information through the target service.
[0135] Based on this, in order to improve the display efficiency of structured data and enhance the user experience, in some embodiments, the above-mentioned structured processing of the response data based on the interaction description information to obtain and display the structured data may specifically include:
[0136] Determine first sub-response data from multiple sub-response data; use an artificial intelligence big model to obtain interaction description information corresponding to the first sub-response data; perform structured processing on the first sub-response data based on the interaction description information to obtain and display structured data.
[0137] The first sub-response data is response data used for structured processing.
[0138] Here, the structured processing may include TTS structured processing and rich media structured processing. If the first sub-response data is used for TTS structured processing, the first sub-response data may be complete target response data corresponding to the interactive input data. The interaction description information corresponding to the target response data may be the first interaction description information. The first interaction description information may include a title, subtitles, and preset keywords.
[0139] If the first sub-response data is used for rich media structured processing, the first sub-response data may be the target paragraph corresponding to the interactive input data. The interactive description information corresponding to the target paragraph may be the second interactive description information. There may be multiple target paragraphs. If the complete target response data includes three major paragraphs, the target paragraph may be any one of the three major paragraphs. In addition, the second interactive description information may include a title, subtitle, summary, and preset keywords, etc. In addition, the second interactive description information may also include picture information. In addition, if the first sub-response data is a target paragraph related to comparison, the second interactive description information may also include a comparison object, comparison item, comparison content, etc.
[0140] Based on this, in some embodiments, determining the first sub-response data from the plurality of sub-response data may specifically include:
[0141] When the sub-response data includes an end marker, the sub-response data is determined to be complete target response data.
[0142] Here, the sub-response data may include an end marker. If the sub-response data includes an end marker, the sub-response data may be the last sub-response data among the multiple sub-response data. Since the subsequent sub-response data includes the previous sub-response data, if the sub-response data includes an end marker, the sub-response data may be the complete target response data corresponding to the interactive input data.
[0143] Based on this, in order to improve the user's reading experience of the target response data, in some embodiments, the first sub-response data includes the complete target response data corresponding to the interactive input data; the interactive description information includes the first description information corresponding to the target response data.
[0144] Accordingly, the above-mentioned structuring of the first sub-response data based on the interaction description information to obtain and display structured data may specifically include:
[0145] The first interaction description information in the target response data is marked to obtain TTS structured data; and the TTS structured data is displayed in the TTS display area.
[0146] Here, the marking process may include highlighting, bolding, italics, underlining, increasing font size, etc. In addition, the TTS structured data may include interactive input data and response data after TTS structured processing is performed on target response data of the interactive input data.
[0147] In addition, the TTS display area can be a preset display area of the voice interaction interface. The TTS display area can be used to display TTS structured data.
[0148] As an example, the TTS structuring engine may markup the first interaction description information based on Markdown syntax to convert the target response data into TTS structured data. After obtaining the TTS structured data, the TTS structured data may be displayed in the TTS display area.
[0149] In some implementations, the response data includes a session identifier, and the method may further include:
[0150] Starting from the second sub-response data, determine whether the first session identifier in the sub-response data is the same as the second session identifier in the previous sub-response data; if the first session identifier is the same as the second session identifier, determine the sub-response data in the sub-response data except the previous sub-response data as the first target response data.
[0151] It should be noted that one response data may include multiple sub-response data, or, when multiple response data are obtained, the response data in the multiple response data may have the same meaning as the sub-response data in one response data.
[0152] Based on this, in order to ensure flexibility in responding to user instructions and improve user experience, in some embodiments, after determining whether the first session identifier in the response data is the same as the second session identifier in the previous response data, the following steps may also be included:
[0153] In the case that the first session identifier is different from the second session identifier, the response data is determined to be the first sub-response data corresponding to the second session identifier, and the dialog interaction for the first sub-response data is performed again.
[0154] Based on this, in order to accurately determine the complete response data for the dialogue interaction, in some embodiments, the above method may further include:
[0155] During the dialog interaction with the first target response data, it is determined whether the response data includes an end identifier; if the response data includes an end identifier, the response data is determined to be the last one of the multiple sub-response data.
[0156] As an example, a schematic diagram of a comparison between plain text and TTS structured data provided by an embodiment of the present disclosure may be shown in FIG3 .
[0157] In this way, by displaying TTS structured data instead of a large paragraph of plain text, the key information can be highlighted, improving the user's reading experience of the target response data.
[0158] In addition, in some embodiments, in the streaming output of multiple sub-response data, the next sub-response data includes the previous sub-response data; and determining the first sub-response data in the multiple sub-response data may include:
[0159] For each sub-response data, determine whether the sub-response data includes a preset separator; if the sub-response data includes one preset separator, determine the sub-response data as the target paragraph; if the sub-response data includes multiple preset separators, determine the sub-response data between the last two adjacent preset separators among the multiple preset separators as the target paragraph.
[0160] The first sub-response data includes the target paragraph.
[0161] Here, determining whether the sub-response data includes a preset delimiter can be implemented by a dynamic parsing engine, wherein the dynamic parsing engine can include a delimiter interceptor. The delimiter interceptor can determine whether the sub-response data includes a preset delimiter. If the sub-response data does not include a preset delimiter, the next sub-response data of the sub-response data can be further determined to determine whether the preset delimiter is included. If the sub-response data includes a preset delimiter, the target paragraph for rich media structuring can be determined based on the preset delimiter.
[0162] Specifically, the preset delimiter can be represented as <|br / >, for example. If the sub-response data is "AAA.BBB.<|br / >," the response data includes one pre-delimiter, and the target paragraph can be "AAA.BBB."; if the sub-response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >," the sub-response data includes two preset delimiters, and the target paragraph can be "CCC.DDD."; if the sub-response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >EEE.FFF.<|br / >," the sub-response data includes three preset delimiters, and the target paragraph can be "EEE.FFF."
[0163] It should be noted that, in the embodiments of the present disclosure, the first process and the second process can be performed simultaneously without affecting each other. Specifically, the first process can be a process of determining target response data from multiple sub-response data and performing TTS structuring based on the target response data. The second process can be a process of determining target paragraphs from multiple sub-response data and performing rich media structuring based on the target paragraphs.
[0164] Based on this, in order to reasonably arrange the sub-response data and ensure the necessity and effectiveness of structured processing, in some embodiments, before determining the target paragraph in the multiple sub-response data, a dynamic parsing engine can be used to determine whether the target paragraph needs to be determined. The dynamic parsing engine can include a command interceptor, a template interceptor, a word count interceptor, a delimiter interceptor, and a structured interceptor.
[0165] As an example, as shown in FIG4 , a flowchart of determining a target paragraph through a dynamic parsing engine provided by an embodiment of the present disclosure may include the following steps:
[0166] S41, determine whether the i-th sub-response data is in the instruction whitelist, if so, execute S42, if not, execute S45;
[0167] S42, intercepting the control instruction in the sub-response data by an instruction interceptor;
[0168] S43, determine whether the control instruction is intercepted, if so, execute S44, if not, execute S45;
[0169] S44, executing control instructions;
[0170] S45. Determine whether the sub-response data is in the structured whitelist. If so, execute S46. If not, set i=i+1 and return to execute S41.
[0171] S46, intercepting and recording key information such as template type, word count threshold, and source in the sub-response data through the template interceptor;
[0172] S47, obtaining the number of words in the sub-response data through a word count interceptor;
[0173] S48, judging whether the word count is greater than the word count threshold by the word count interceptor, if so, executing S49, if not, setting i=i+1, and returning to executing S41;
[0174] S49, using the delimiter interceptor to determine whether the sub-response data includes a preset delimiter, if so, executing S410, if not, setting i=i+1, and returning to executing S41;
[0175] S410. Determine the target paragraph based on the preset delimiter through the structured interceptor, and initiate a rich media structured request corresponding to the target paragraph to the large model module. The rich media structured request is used to request the large model module to obtain the second interaction description information corresponding to the target paragraph.
[0176] In Figure 4, on the one hand, by determining the target paragraph when the sub-response data meets the above-mentioned multiple conditions, the sub-response data can be reasonably arranged to ensure the necessity and effectiveness of structured processing; on the other hand, by intercepting and executing the control instructions in the sub-response data for each sub-response data when the sub-response data is in the instruction whitelist, the timeliness of executing user instructions can be improved and the user experience can be enhanced.
[0177] In addition, in some embodiments, the inputting of the first sub-response data into the artificial intelligence big model to obtain, through the artificial intelligence big model, interaction description information corresponding to the first sub-response data may specifically include:
[0178] The target paragraph is input into the AI model. The AI model first obtains the corresponding textual description information, such as the title, subheadings, abstract, and preset keywords, and then searches for images corresponding to the textual description information. Images may or may not be found, and this is not a limitation. Furthermore, the image information may include, for example, links to the images.
[0179] Based on this, in order to further enhance the user's visual experience, in some embodiments, the first sub-response data includes a target paragraph, the interaction description information includes second interaction description information corresponding to the target paragraph, and the second interaction description information includes image information; the above-mentioned structuring of the first sub-response data based on the interaction description information to obtain and display structured data may further include:
[0180] The second interaction description information is rendered to generate card information; and the card information is displayed in the card information display area.
[0181] Here, the second interaction description information may also include the target paragraph's number in the complete target response data, as well as a paragraph identifier. After the artificial intelligence model returns the second interaction description information, the rich media structuring engine may render the second interaction description information and generate and display card information.
[0182] The card information may be in a rich media format. Rich media card information may include audio, text, links, interactive controls, images, and the like corresponding to the target paragraph. Furthermore, the card information may also include interactive input data and structured response data. For example, a schematic diagram of a card information provided in an embodiment of the present disclosure may be shown in FIG5 .
[0183] As an example, after the AI model returns the second interaction description information, the rich media structuring engine can pull up a Hypertext Markup Language interface (e.g., HTML 5, H5) and transparently transmit the second interaction description information to H5 via a JS bridge. The client can load the H5 and the second interaction description information through a browser (e.g., a webview container) to generate and display the card information.
[0184] In addition, the card information display area can be a preset display area of the voice interaction interface. The card information display area can be adjacent to the TTS display area and located to the right of the TTS display area. Of course, the card information display area can also be located to the left of the TTS display area, which is not limited here. A schematic diagram of a voice interaction interface including a TTS display area and a card information display area provided by an embodiment of the present disclosure can be shown in Figure 6. In Figure 6, the left side can be the TTS display area. The right side can be the card information display area.
[0185] In this way, the user's visual experience can be further enhanced by displaying card information.
[0186] In addition, as mentioned above, there can be multiple target paragraphs. Based on this, in order to ensure the timeliness of card information display and further enhance the user's visual experience, in some embodiments, the card information displayed in the card information display area may specifically include:
[0187] According to the generation order of the multiple target paragraphs, the displayed card information is updated based on the multiple target paragraphs; the continuously updated card information is displayed in the card information display area until the target paragraph includes an end marker, thereby obtaining complete card information.
[0188] Here, each time a target paragraph is determined, the card information corresponding to the target paragraph can be generated and displayed once. In the process of displaying the card information corresponding to the first target paragraph, the dynamic parsing engine can synchronously determine the second target paragraph from the multiple sub-response data streamed after the first target paragraph. After the second target paragraph is determined, the card information being displayed can be updated. For example, as shown in Figure 6, the card information display area can update the chart. In addition, the TTS display area can display a prompt message "Generating chart..." to remind the user that the card information has not yet been displayed.
[0189] If the sub-response data includes an end marker, it can be determined that the sub-response data is the complete target response data. Therefore, the target paragraph determined based on the sub-response data may also include an end marker. If the target paragraph includes an end marker, it can be determined that the target paragraph is the last target paragraph. After generating and displaying the card information corresponding to the last target paragraph, the card information can be completely updated to obtain the complete card information.
[0190] In this way, by displaying a portion of the card information each time a target paragraph is obtained, the timeliness of the card information display can be ensured, further improving the user's visual experience.
[0191] In addition, the sub-response data output by the large model module may also include display template information. The display template information included in multiple sub-response data may be the same. In addition, the display template information may be template type information corresponding to the display template of the sub-response data. The template types may include description templates, general classification templates, comparison templates, timeline templates, and service expert templates. In addition, different template types may correspond to different card types. The card types may include description cards, general classification cards, and service expert cards. Among description cards, general classification cards, and service expert cards, the cards may include cards with pictures and cards without pictures. The correspondence between template types and card types may be shown in Figure 7.
[0192] As an example, after receiving the interactive input data, the large model module may also determine a reply type corresponding to the interactive input data, and determine a presentation template type of the response data corresponding to the interactive input data based on the reply type.
[0193] Based on this, in order to further improve the display effect of card information and enhance the user's visual experience, in some embodiments, the above-mentioned rendering of the second interaction description information to generate card information may specifically include:
[0194] The second interaction description information is rendered according to the display template information to generate card information.
[0195] Here, the dynamic parsing engine may include a template interceptor, which can intercept a record of the display template information in the sub-response data. If the dynamic parsing engine is able to intercept the display template information, the rich media structuring engine can render the second interaction description information according to the display template information to generate card information. Specifically, the rich media structuring engine can pull up a hypertext markup language interface (such as HTML 5, H5) interface, and pass the second interaction description information to H5 through the JS bridge. The client can pull up the H5 interface corresponding to the template type through a browser (such as a webview container), load the H5 template and the second interaction description information, and generate card information.
[0196] In this way, by using templates and data to dynamically generate the user interface (i.e., card information), the structure and style of the user interface can be separated from the data, so that different user interfaces can be dynamically generated according to different data, further improving the display effect of the card information and enhancing the user's visual experience.
[0197] In addition, the sub-response data may also include source information for the sub-response data. The source information can indicate how the AI model obtained the response data. For example, if the response data corresponds to the interactive input data and is obtained by the AI model by searching a pre-set question-and-answer library corresponding to the vehicle, the source information for the response data (including the multiple sub-response data streamed out) can be the question-and-answer library corresponding to the vehicle.
[0198] As an example, when the artificial intelligence large model obtains response data, it can synchronously determine the source information corresponding to the response data and mark the source information in the response data (including multiple sub-response data output in streaming format).
[0199] Based on this, in order to further enhance the user experience of the question-and-answer system, in some embodiments, rendering the second interaction description information according to the display template information to generate card information may specifically include:
[0200] In the case where the source information of the response data is the question and answer library corresponding to the vehicle, the display template information is modified to the service expert template, and the service expert template is the display template corresponding to the vehicle question and answer;
[0201] The second interaction description information is rendered according to the service expert template to generate card information.
[0202] Here, the service expert template can be a display template corresponding to the vehicle Q&A. The text description in the service expert template can be the operation steps related to the vehicle's use service. The card information generated based on the service expert template can include the following content, for example: "The steps for using the child safety lock are as follows: 1. Click "Vehicle" in the central control screen settings, select "Locks," and select the option under "Child Lock Button" to set the child lock button; 2. After selecting "Left" or "Right", when the child safety lock is turned on, the corresponding side rear door will not be able to be opened and the window cannot be controlled from inside the car."
[0203] In this way, by rendering the second interaction description information according to the service expert template corresponding to the vehicle question and answer when the source information of the response data is the question and answer library corresponding to the vehicle, and generating card information, the professionalism of the card information can be improved, and the user experience of the question and answer model can be further enhanced.
[0204] In addition, in order to improve the timeliness of answer data reporting and display and enhance user experience, in some embodiments, after determining the target service corresponding to the question-answering scenario from multiple services, the following steps may also be included:
[0205] A conversational interaction is performed on the first output sub-response data through the target service, wherein the conversational interaction includes voice interaction and data display associated with the voice interaction; starting from the second output sub-response data, the sub-response data other than the previous sub-response data in the sub-response data are determined as the second sub-response data; a conversational interaction is performed on the second sub-response data until the sub-response data is the last one of the multiple sub-response data.
[0206] Here, voice interaction may include voice broadcasting of sub-response data. Data display associated with voice interaction may include on-screen display of text information during the voice interaction. The data may be displayed in a TTS display area. The displayed data may include text data, image data, video data, etc.
[0207] As an example, if the question-and-answer system is an in-vehicle question-and-answer system, and the large model module in the in-vehicle question-and-answer system is located in the cloud server, and the remaining modules other than the large model module are located in the terminal server, then each time the large model module outputs a sub-response data, the sub-response data can be sent to the terminal server via the SDK. The terminal server can receive the sub-response data through the main interaction module. After the main interaction module receives the sub-response data, it can send the sub-response data to the dialogue management module, which in turn sends the sub-response data to the TTS structuring engine through the dialogue management module. Each time the TTS structuring engine receives a sub-response data, it can process the sub-response data based on the additional broadcasting strategy to obtain a second sub-response data for voice broadcast and on-screen display. After obtaining the second sub-response data, the TTS structuring engine can, on the one hand, send the second sub-response data to the audio module via the TTS service, and the audio module will broadcast the second sub-response data in voice. On the other hand, the TTS structuring engine can send the second sub-response data to the generative user interface service via the TTS service, and the generative user interface service will stream the second sub-response data on the screen.
[0208] In addition, the TTS structured engine can also determine whether to perform voice broadcast and screen display on the second sub-response data based on the interruption broadcast strategy, and manage the TTS broadcast status. Among them, the TTS broadcast status can include which sub-response data is broadcast. In addition, the interruption broadcast strategy can be a broadcast strategy executed when the session identifier of the subsequent sub-response data is different from that of the previous sub-response data. The additional broadcast strategy can be a broadcast strategy executed when the session identifier of the subsequent sub-response data is the same as that of the previous sub-response data.
[0209] As an example, the sub-response data may include a session identifier. If the session identifier of the subsequent sub-response data is the same as that of the previous sub-response data, then the two sub-response data may be sub-response data corresponding to the same interactive input data, and the second sub-response data corresponding to the session identifier can continue to be voice broadcast and displayed on the screen; if the session identifier of the subsequent sub-response data is different from that of the previous sub-response data, then the two sub-response data are not sub-response data corresponding to the same interactive input data. That is, if another interactive input data is received during the process of responding to the current interactive input data, the response to the current interactive input data can be stopped and the other interactive input data can be responded to instead.
[0210] As a more specific example, as shown in FIG8 , a flow chart of performing a dialog interaction on sub-response data by a TTS structuring engine may include the following steps:
[0211] S81. Receive the first sub-response data streamed from the large model module, and perform voice broadcast and on-screen display of the first sub-response data.
[0212] Exemplarily, by voice broadcasting and on-screen display of the first sub-response data, a prerequisite for dialogue interaction with the first response data is provided.
[0213] S82, receiving the i-th (i≥2) sub-response data streamed from the large model module;
[0214] S83. Determine whether the session identifiers of the i-th sub-response data and the (i-1)-th sub-response data are the same. If so, execute S84; if not, execute S86.
[0215] S84. Record the index value of the i-th sub-response data;
[0216] S85: Based on the additional announcement strategy, determine the portion of the i-th sub-answer data excluding the (i-1)-th sub-answer data as the second sub-answer data, and perform voice announcement and on-screen display on the second sub-answer data;
[0217] S86. Reset the index value and session identifier of the i-th sub-response data;
[0218] S87, based on the interruption broadcast strategy, voice broadcast and screen display of the i-th sub-response data;
[0219] S88: Set i=i+1, and return to execute S82 until the sub-response data includes the end marker.
[0220] In this way, by generating speech frame by frame according to the text, speech synthesis becomes a streaming process, which can output sound in real time and play it simultaneously with the generation of speech. This method can make the synthesized speech smoother and reduce latency. In this way, by streaming the output of multiple sub-response data to the large model module for streaming broadcast and streaming on-screen display, each time a part of the sub-response data is generated, that is, a part of the sub-response data can be broadcast and displayed, thereby improving the timeliness of the broadcast and display of the response data and enhancing the user experience.
[0221] Based on this, since the first sub-response data and at least one second sub-response data have been streamed onto the screen and displayed in the TTS display area during the process of streaming output of multiple sub-response data by the large model module, in order to improve the timeliness of the response data display while also improving the user's reading experience of the response data and further enhancing the user experience, in some embodiments, the display of TTS structured data in the TTS display area may specifically include:
[0222] The displayed first sub-answer data and at least one second sub-answer data are replaced with TTS structured data.
[0223] For example, an interactive schematic diagram of a question-and-answer method provided by an embodiment of the present disclosure may be shown in FIG9 .
[0224] In Figure 9, the question-and-answer method can be applied to an in-vehicle question-and-answer system. The large model module can be located on a cloud server. The main interaction module, dialogue management module, service management module, and AI-based service assistant can be located on a terminal server. Specifically, the main interaction module on the terminal server can receive interactive input data from a user. If the interactive input data is input in a non-voice format, the main interaction module can determine user intent description information corresponding to the interactive input data. If the interactive input data is input in a voice format, the main interaction module can send the interactive input data to the large model module. The large model module can input the interactive input data into the artificial intelligence large model, which can then recognize the interactive input data using the large voice model and generate user intent description information. After determining the user intent description information, the large model module can return the user intent description information to the main interaction module. After receiving the user intent description information, the main interaction module can send the user intent description information to the dialogue management module. Furthermore, after determining the user intent description information, the large model module can continue to respond to the interactive input data based on the user intent description information, generating and caching response data.
[0225] After receiving the user intention description information, the dialogue management module can determine the user interaction scenario corresponding to the user intention description information according to the arbitration strategy, and send the user interaction scenario to the service management module. After receiving the user interaction scenario, the service management module can determine the target service assistant corresponding to the user interaction scenario based on the correspondence between the user interaction scenario and the registration intent, and the correspondence between the registration intent and the service assistant. In the case where the user interaction scenario is a question-and-answer scenario, the service management module can determine that the target service assistant is an AI-type service assistant, and notify the AI-type service assistant to register the scene and obtain a scene identifier. The service management module can record the correspondence between the scene identifier and the AI-type service assistant, and send the scene identifier to the dialogue management module.
[0226] After receiving the scene identifier, the dialogue management module can send a streaming data request corresponding to the scene identifier to the large model module through the main interaction module. After receiving the streaming data request, the large model module can stream the response data corresponding to the scene identifier. That is, the large model module can stream multiple sub-response data (including the scene identifier) to the main interaction module. Each time the main interaction module receives a sub-response data, it can send a sub-response data to the dialogue management module. Each time the dialogue management module receives a sub-response data, it can send a sub-response data to the service management module. Each time the service management module receives a sub-response data, it can route the sub-response data to the AI-type service assistant corresponding to the scene identifier based on the correspondence between the scene identifier and the service assistant.
[0227] After receiving the sub-response data, the AI-type service assistant can firstly perform TTS voice broadcast and streaming to the screen based on multiple sub-response data through the TTS broadcast engine; secondly, the AI-type service assistant can determine the target paragraph based on multiple sub-response data through the rich media structuring engine, and send a rich media structuring request to the large model module, as well as generate and display card information based on the second interaction description information returned by the large model module. In addition, the AI-type service assistant can determine the complete target response data based on multiple sub-response data through the TTS structuring engine, send a TTS structuring request to the large model module, and perform tagging based on the first interaction description information returned by the large model module to obtain and display TTS structured data. Among them, displaying TTS structured data can be replacing the first sub-response data and multiple second sub-response data displayed on the screen with TTS structured data.
[0228] In addition, the large model module can also be set up on the terminal server, which is not limited here.
[0229] Therefore, through the embodiments of the present disclosure, by generating a part of sub-response data each time, that is, broadcasting and displaying a part of sub-response data, the timeliness of the broadcasting and display of the response data can be guaranteed, and the user experience can be improved; by performing TTS structuring on the complete target response data, generating and displaying TTS structured data, compared with displaying a large paragraph of plain text, the user's reading experience of the target response data can be improved; by performing rich media structuring on the target paragraph, generating card information, and displaying the card information based on generative UI technology, the user's visual experience can be further improved; by performing TTS voice broadcasting, TTS streaming to the screen, TTS structuring and rich media structuring and other processes simultaneously based on the dynamic parsing engine in the process of streaming output of multiple sub-response data, the response speed to user instructions (that is, interactive input data) can be improved, and the user experience can be further improved.
[0230] In one embodiment, after the response data is determined to be the last one of the multiple sub-response data, the following operations may be performed:
[0231] The second target response data is input into the artificial intelligence big model, so as to extract the interaction description information from the second target response data through the artificial intelligence big model, mark the interaction description information in the second target response data, and obtain and display TTS structured data.
[0232] The second target response data is the last one of the multiple response data, and the interaction description information is used to perform structured processing on the second target response data.
[0233] Based on this, in order to ensure the necessity and effectiveness of structured processing, in some embodiments, the second target response data includes a scenario identifier. Before the second target response data is input into the artificial intelligence model, the following may also be included:
[0234] It is determined whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist, where the structured whitelist includes multiple user interaction scenarios.
[0235] Based on this, the above-mentioned inputting of the second target response data into the artificial intelligence model may specifically include:
[0236] When the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, the second target response data is input into the artificial intelligence big model.
[0237] In this way, by performing TTS structuring processing on the second target response data when the second target response data is in the structured whitelist, the necessity and effectiveness of the structuring processing can be guaranteed.
[0238] In addition, the process of streaming multiple sub-response data from a large AI model may also include:
[0239] For each sub-response data, determine whether the sub-response data is in the instruction whitelist; if the sub-response data is in the instruction whitelist, intercept the control instruction included in the sub-response data; if the control instruction is intercepted, execute the control instruction.
[0240] In this way, by intercepting and executing the control instructions in the response data for each sub-response data when the sub-response data is in the instruction whitelist, the timeliness of executing user instructions can be guaranteed and the user experience can be improved.
[0241] In one embodiment, the interactive input data includes multimodal feature data and instruction interaction data, wherein the multimodal feature data represents the input form of the interactive input data; and the instruction interaction data represents the user intention corresponding to the interactive input data.
[0242] The command interaction data may represent the user intent corresponding to the interactive input data. For example, the command interaction data may include control commands related to vehicle control and vehicle settings, question commands related to vehicle Q&A and general Q&A, or both control commands and question commands, without limitation.
[0243] Through the above-mentioned method, the data range of the interactive input data is expanded, thereby improving the diversity of the interactive input data.
[0244] In some embodiments, the technical solutions provided by the present disclosure may further perform the following operations:
[0245] The interactive input data is input into the artificial intelligence big model to identify the interactive input data through the artificial intelligence big model, obtain the reply type corresponding to the interactive input data, and determine the display template information corresponding to the reply type; and the artificial intelligence big model responds to the interactive input data to obtain and output multiple response data including display template information.
[0246] Based on this, in some embodiments, the target paragraph includes display template information; inputting the target paragraph into the artificial intelligence big model to obtain card information corresponding to the target paragraph through the artificial intelligence big model may specifically include:
[0247] The target paragraph is input into the artificial intelligence big model to determine the text interaction description information corresponding to the display template information through the artificial intelligence big model, extract the text interaction description information from the target paragraph, and search for multimedia interaction description information corresponding to the text interaction description information; the interaction description information and the multimedia description information are determined as card information.
[0248] Based on this, in order to further improve the display effect of card information and enhance the user's visual experience, in some embodiments, rendering and displaying the card information corresponding to the target paragraph may specifically include:
[0249] Fill the valid information into the display template corresponding to the display template information, obtain the card information corresponding to the target paragraph, and display the card information in the card information display area.
[0250] The valid information includes at least one of text interaction description information and multimedia interaction description information.
[0251] In this way, by using templates and data to dynamically generate the user interface (i.e., card information), the structure and style of the user interface can be separated from the data, so that different user interfaces can be dynamically generated according to different data, further improving the display effect of the card information and enhancing the user's visual experience.
[0252] In addition, the response data may also include source information of the response data. The source information may characterize the manner in which the artificial intelligence large model obtains the response data. As described above, obtaining response data corresponding to the interactive input data includes searching for response data corresponding to the interactive input data in a pre-set question and answer library, searching for response data corresponding to the interactive input data on the network, and generating response data corresponding to the interactive input data. Based on this, for example, if the response data is the response data corresponding to the interactive input data obtained by the artificial intelligence large model by searching in a pre-set question and answer library corresponding to the vehicle, then the source information of the response data may be the question and answer library corresponding to the vehicle.
[0253] As an example, when obtaining response data, the artificial intelligence large model can simultaneously determine the source information corresponding to the response data and mark the source information in the response data.
[0254] Based on this, in order to further enhance the user experience of the question-answering system, in some embodiments, the interaction description information is rendered according to the display template information to generate card information corresponding to the target paragraph, which may specifically include:
[0255] When the source information of the response data is the question and answer library corresponding to the vehicle, the display template information is modified to the service expert template, and the service expert template is the display template corresponding to the vehicle question and answer; the interaction description information is rendered according to the service expert template to generate card information corresponding to the target paragraph.
[0256] Here, the service expert template can be a display template corresponding to the vehicle Q&A. The service expert template can include two display areas, upper and lower. Among them, the display area in the upper part can display multimedia interactive description information, and the display area in the lower part can display text interactive description information. The text interactive description information in the service expert template can be the operation steps related to the use of the vehicle service. The card information generated based on the service expert template may include the following content, for example: "The steps for using the child safety lock are as follows: 1. Click "Vehicle" in the central control screen settings, select "Locks," and select the option under "Child Lock Button" to set the child lock button; 2. After selecting "Left" or "Right", when the child safety lock is turned on, the corresponding side rear door will not be able to be opened and the windows cannot be controlled from inside the car."
[0257] In this way, by rendering the interactive description information according to the service expert template corresponding to the vehicle question and answer when the source information of the response data is the question and answer library corresponding to the vehicle, and generating card information, the professionalism of the card information can be improved, and the user experience of the question and answer model can be further enhanced.
[0258] Based on this, in order to ensure the timeliness of executing user instructions and improve user experience, in some embodiments, before determining whether the sub-response data is in the structured whitelist, the following steps may also be performed:
[0259] Determine whether the response data is in the instruction whitelist; if the response data is in the instruction whitelist, intercept the control instruction included in the response data; if the control instruction is intercepted, execute the control instruction.
[0260] In this way, by intercepting and executing the control instructions in each response data when the response data is in the instruction whitelist, the timeliness of executing user instructions can be guaranteed and the user experience can be improved.
[0261] Based on this, in order to ensure the timeliness of card information display and improve user experience, in some embodiments, the above method may further include:
[0262] When the sub-response data is in the structured whitelist, it is determined whether the sub-response data includes a preset separator.
[0263] Here, the dynamic parsing module may also include a delimiter interceptor. The delimiter interceptor can determine whether the response data includes a preset delimiter. If the sub-response data does not include a preset delimiter, it can continue to determine whether the next sub-response data includes a preset delimiter. If the sub-response data includes a preset delimiter, the target paragraph for rich media structuring can be determined based on the preset delimiter. In addition, the card information can be card information in a rich media format. The card information in a rich media format may include audio, text, links, interactive controls, pictures, etc. corresponding to the target paragraph.
[0264] Based on this, in order to further improve the user experience, in some embodiments, when the number of words in the sub-response data is greater than the word count threshold, and before determining whether the sub-response data includes a preset separator, the following steps may also be included:
[0265] Intercept the network status in the sub-response data; if the network status is abnormal, display the target control, which is used to regenerate the card information.
[0266] The dynamic parsing module may also include a network anomaly interceptor. This interceptor can intercept the network status in the response data. If the network status indicates an anomaly, a target control for regenerating card information can be displayed in the voice interaction interface. After seeing the target control, the user can determine that the network anomaly exists and choose whether to regenerate the card information.
[0267] In this way, by providing users with an interactive method to regenerate and display card information when card information display fails, the user experience can be further improved.
[0268] The present disclosure also provides another question-answering method, which may include the following steps:
[0269] Step A1: Receive user interaction input data; wherein the interaction input data includes multimodal feature data and instruction interaction data, the multimodal feature data representing the input form of the interaction input data, and the instruction interaction data representing the user intent corresponding to the interaction input data.
[0270] Step A2: Input the interactive input data into the artificial intelligence big model so that the artificial intelligence big model can identify and respond to the interactive input data to obtain multiple response data.
[0271] Step A3: During the process of streaming out multiple response data, a conversation interaction is performed on the first response data, wherein the conversation interaction includes voice interaction; starting from the second response data, the response data other than the previous response data is determined as the first target response data; a conversation interaction is performed on the first target response data until the response data is the last one among the multiple response data.
[0272] The question-answering method of the embodiment of the present application inputs interactive input data into an artificial intelligence big model, uses the artificial intelligence big model to identify and respond to the interactive input data, obtains and streams out multiple response data, on the one hand, can use the natural language generation technology of the artificial intelligence big model to generate response data corresponding to the interactive input data, thereby ensuring the accuracy of the response data. On the other hand, the multiple response data output in a streamed manner can be multiple response data obtained after the artificial intelligence big model extracts key information from the complete response data. Therefore, by conducting a dialogue interaction on the first output response data during the process of the artificial intelligence big model streaming out multiple response data, and starting from the second response data, sequentially determining multiple first target response data from the multiple response data, and conducting dialogue interaction on the multiple first target response data respectively, it is possible to broadcast the relatively important response data corresponding to the interactive input data, which can ensure the validity of the response data compared to broadcasting the complete response data at one time, so that the user can promptly determine the information they want from the broadcast response data, thereby improving the user experience.
[0273] In some embodiments, conversational interaction may include voice interaction and data display associated with the voice interaction. Voice interaction may include, upon receiving user interaction input data, the voice broadcast of response data corresponding to the interaction input data. Data display associated with voice interaction may include the on-screen display of data during the voice interaction process. The data may be displayed in a text-to-speech (TTS) display area. The TTS display area may be a data display area pre-set in the voice interaction interface. The displayed data may include text data, image data, video data, etc.
[0274] As an example, the question-answering system may include a main interaction module and a large model module. Among them, the main interaction module may have an Automatic Speech Recognition (ASR) function. If the input form of the interactive input data is voice input, the main interaction module can convert the voice instruction into instruction text through the ASR function and send the instruction text to the large model module. After receiving the instruction text, the large model module can perform semantic analysis on the instruction text through the artificial intelligence large model to obtain instruction interaction data (i.e., determine the user's intention), and return the instruction interaction data to the main interaction module.
[0275] In addition, the question-and-answer system may also include a voice user interface (VUI) corresponding to the main interaction module. The voice interaction interface may include a text input box. Users can enter text instructions in the text input box. The main interaction module may determine the instruction interaction data corresponding to the interaction input data based on the correspondence between the text instructions and the instruction interaction data.
[0276] In addition, the voice interaction interface may also include multiple interactive controls, which may include preset commands related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. Users can perform touch commands by clicking on the interactive controls. The main interaction module may determine the command interaction data corresponding to the interactive input data based on the correspondence between the touch commands and the command interaction data.
[0277] In addition, if the voice interaction interface supports gesture air touch, the user can select any preset command by gesture air control interaction control. The main interaction module can determine the command interaction data corresponding to the interaction input data based on the correspondence between the gesture command and the command interaction data.
[0278] Artificial Intelligence (AI) large models refer to high-performance AI models built using vast amounts of training samples and computing resources. They can learn extensive language knowledge, image features, and speech patterns, and reason and generate human-like outputs. They have broad applications in natural language processing, image recognition, and speech recognition. Examples of these large AI models include large language models (LLMs), Chat Generative Pre-trained Transformers (ChatGPTs), multimodal large models, and multimodal cognitive large models.
[0279] As an example, if the input form of the interactive input data is voice input, the main interactive module can send the interactive input data to the large model module after receiving the user's interactive input data. The large model module can input the interactive input data into the artificial intelligence large model, and the artificial intelligence large model recognizes and responds to the interactive input data. Among them, recognizing the interactive input data can generate instruction interaction data. Responding to the interactive input data can generate response data corresponding to the instruction interaction data. If the input form of the interactive input data is non-voice input (including text input, touch input and gesture input), the main interactive module can determine the instruction interaction data corresponding to the interactive input data after receiving the user's interactive input data, and send the instruction interaction data to the large model module. The large model module can input the instruction interaction data into the artificial intelligence large model, and respond to the instruction interaction data through the artificial intelligence large model. Among them, responding to the instruction interaction data can generate response data corresponding to the instruction interaction data.
[0280] As an example, if the large model module is located on the cloud server, and the main interaction module and the dialogue management module are located on the terminal server, then after the terminal server sends the interactive input data to the cloud server through the main interaction module, it can continue to send an asynchronous generation interface to the cloud server so that the large model module returns the instruction interaction data to the terminal server. After the large model module returns the instruction interaction data to the terminal server, it can continue to obtain the response data corresponding to the interactive input data (including the instruction interaction data) and cache the obtained response data. Among them, obtaining the response data corresponding to the interactive input data includes searching for the response data corresponding to the interactive input data in a pre-set question and answer library, searching for the response data corresponding to the interactive input data in the network, and generating the response data corresponding to the interactive input data. In the process of the artificial intelligence large model generating the response data corresponding to the interactive input data, when predicting the next word or character, the previously generated content will be considered. This context-aware generation method enables the artificial intelligence large model to generate text with coherence and a certain logic.
[0281] Furthermore, after receiving the command interaction data, the main interaction module in the terminal server may send the command interaction data to the dialogue management module. After receiving the command interaction data, the dialogue management module may determine the user interaction scenario corresponding to the command interaction data based on an arbitration strategy. User interaction scenarios may include question-and-answer scenarios and interactive control scenarios. Specifically, question-and-answer scenarios may include general question-and-answer scenarios and vehicle question-and-answer scenarios. Interactive control scenarios may include vehicle control scenarios. Furthermore, the arbitration strategy may be a pre-set correspondence between command interaction data and user interaction scenarios.
[0282] If the dialogue management module determines that the user interaction scenario is a target Q&A scenario, it may send the target Q&A scenario to the service management module. Specifically, the target Q&A scenario can be any of a number of Q&A scenarios, such as a general Q&A scenario, a vehicle Q&A scenario, an encyclopedia Q&A scenario, a comparison Q&A scenario, or a plan-making scenario. After receiving the target Q&A scenario, the service management module may determine a target service assistant corresponding to the target Q&A scenario based on the correspondence between the user interaction scenario and the service assistant. The service assistant may correspond to the user interaction scenario and be responsible for executing subsequent actions corresponding to the response data. The target service assistant may be a service assistant corresponding to the target Q&A scenario. For example, the target service assistant may be an AI-based service assistant. An AI-based service assistant may specifically include a dynamic parsing engine and a rich media structuring engine. After determining the target service assistant, the service management module may notify the target service assistant to register the target Q&A scenario and obtain a scenario identifier. The target service assistant may then send the scenario identifier to the service management module. Upon receiving the scenario identifier, the service management module may record the correspondence between the target service assistant and the scenario identifier and send the scenario identifier to the dialogue management module. The dialogue management module can send a streaming data request corresponding to the scene identifier to the large model module. The streaming data request can be used to request the large model module to stream-output multiple response data corresponding to the interactive input data. After receiving the streaming data request, the large model module can respond to the streaming data request and stream-output multiple response data. The latter response data may include the previous response data. For example, if the interactive input data is "Introduce Zhou xx," the multiple response data streamed out may include: the first response data is "Zhou xx," the second response data is "Zhou xx, from xxx," and the third response data is "Zhou xx, from xxx, won x award on x year x month x day..." In addition, the response data may include a scene identifier.
[0283] As a more specific example, if the large model module is located in a cloud server, and the main interaction module and the dialogue management module are located in a terminal server, the dialogue management module can send a streaming data request to the large model module through the main interaction module.
[0284] In some embodiments, in A3, conversational interaction may include voice interaction and data display associated with the voice interaction. Voice interaction may include, upon receiving user interaction input data, the voice broadcast of response data corresponding to the interaction input data. Data display associated with voice interaction may include the on-screen display of data during the voice interaction process. The data may be displayed in a text-to-text (TTS) display area. The TTS display area may be a data display area pre-set in the voice interaction interface. The displayed data may include text data, image data, video data, etc.
[0285] As an example, the dialogue management module can receive multiple response data streamed from the large model module via the main interaction module. After receiving the response data, the dialogue management module can send the response data, including the scenario identifier, to the service management module. The service management module can route the response data to the target service assistant corresponding to the scenario identifier based on the correspondence between the scenario identifier and the target service assistant. After receiving the response data, the target service assistant can perform subsequent processing on the response data.
[0286] As a more specific example, the question-and-answer system may also include a generative user interface service. The target service assistant may include a TTS broadcast engine, a TTS service, and an audio module. If the large model module in the question-and-answer system is located in the cloud server, and the remaining modules except the large model module are located in the terminal server, then after obtaining the first response data, the large model module can send the first response data to the terminal server through the software development kit (SDK). The terminal server can receive the response data through the main interaction module. After the main interaction module receives the response data, it can send the first response data to the dialogue management module, and the dialogue management module can send the first response data to the TTS broadcast engine in the target service assistant. After receiving the first response data, the TTS broadcast engine can, on the one hand, send the first response data to the audio module through the TTS service, and perform voice interaction on the first response data through the audio module. On the other hand, the TTS broadcast engine can send the first response data to the generative user interface service through the TTS service, and the generative user interface service can perform data display on the first response data.
[0287] In some embodiments, in A4, since the latter response data can include the previous response data, for the i-th (i≥2) response data, the response data that has been voice interacted with may no longer be repeatedly broadcast, and the response data that has been data displayed may no longer be repeatedly displayed. Therefore, starting from the second output response data, the response data can be processed by the TTS broadcast engine based on the additional broadcast strategy to obtain the first target response data for voice interaction and data display. The additional broadcast strategy can be to perform additional broadcast and additional display on the response data that has not been broadcast and displayed in the i-th response data.
[0288] As an example, after receiving the i-th response data, the TTS broadcast engine can determine the response data in the i-th response data except the i-1-th response data as the first target response data. Let i = i + 1 and repeat the above steps to sequentially determine multiple first target response data until the i-th response data is the last response data streamed output by the artificial intelligence large model, thus obtaining all the first target response data corresponding to the interactive input data.
[0289] In addition, the response data may include a session identifier. The session identifier can be used to uniquely identify a session. For example, if a response to first interactive input data results in multiple first response data, the session identifiers of the multiple first response data may be the same. If a response to second interactive input data results in multiple second response data, the session identifiers of the multiple second response data may be the same. Furthermore, the session identifiers of the second response data may be different from the session identifiers of the first response data.
[0290] Based on this, in order to ensure that the response data other than the (i-1)th response data in the i-th response data can be determined as the first target response data and to ensure the accuracy of the response to the interactive input data, in some embodiments, the above S140 may specifically include:
[0291] Starting from the second response data, determining whether the first session identifier in the response data is the same as the second session identifier in the previous response data;
[0292] In a case where the first session identifier is the same as the second session identifier, the response data other than the previous response data in the response data is determined as the first target response data.
[0293] Here, after receiving the i-th (i≥2) response data, the TTS broadcast engine can first determine whether the first session identifier in the i-th response data is the same as the second session identifier in the i-1-th response data. If the first session identifier is the same as the second session identifier, it can be determined that the i-th response data and the i-1-th response data are response data corresponding to the same interactive input data, and further, it can be determined that the i-th response data includes the i-1-th response data. In this way, it can be ensured that the response data in the i-th response data except the i-1-th response data can be determined as the first target response data, thereby ensuring the accuracy of the response to the interactive input data.
[0294] Based on this, in order to ensure flexibility in responding to user instructions and improve user experience, in some embodiments, after determining whether the first session identifier in the response data is the same as the second session identifier in the previous response data, the following steps may also be performed:
[0295] In the case that the first session identifier is different from the second session identifier, the response data is determined to be the first response data corresponding to the second session identifier, and the dialog interaction for the first response data is performed again.
[0296] Here, if the first session identifier is different from the second session identifier, it can be determined that the i-th (i≥2) response data and the i-1-th response data are not response data corresponding to the same interactive input data. That is, the i-th response data may be the first response data corresponding to another interactive input data. That is, the i-th response data may not include the i-1-th response data. Therefore, the TTS broadcast engine can stop the response data corresponding to the previous interactive input data that is currently being broadcast and displayed based on the interruption broadcast strategy, and start voice interaction and data display for the response data corresponding to the next interactive input data. That is, if the first session identifier is different from the second session identifier, voice interaction and data display can be performed directly on the i-th response data (that is, the first response data corresponding to another interactive input data).
[0297] In this way, by stopping the current dialogue interaction and starting the dialogue interaction with the response data corresponding to the current interaction input data if another interaction input data is received, the flexibility of responding to user instructions can be guaranteed and the user experience can be improved.
[0298] In addition, the TTS broadcast engine can also manage the TTS broadcast status, wherein the TTS broadcast status may include the number of response data to be broadcast.
[0299] In some embodiments, each time a first target response data is generated, a conversational interaction can be performed on the first target response data. As described above, if the response data is the last response data, multiple first target response data can be determined according to the above-mentioned additional broadcast strategy. By first performing voice interaction and data display on the first response data, and then performing voice interaction and streaming data display on the multiple first target response data in the order in which the multiple first target response data are generated, the complete response data can be broadcast and displayed.
[0300] As an example, after determining the first target response data, the TTS broadcast engine can, for each first target response data, on the one hand, send the first target response data to the audio module through the TTS service, and perform voice broadcasting of the first target response data through the audio module. On the other hand, the TTS broadcast engine can send the first target response data to the generative user interface service through the TTS service, and perform data display of the first target response data through the generative user interface service.
[0301] Based on this, in order to accurately determine the complete response data for the dialogue interaction, in some embodiments, the above S150 may specifically include:
[0302] During the dialog interaction with the first target response data, determining whether the response data includes an end identifier;
[0303] When the response data includes an end marker, the response data is determined to be the last one of the plurality of response data.
[0304] Here, the response data may include an end marker. If the response data includes an end marker, the response data may be the last response data among multiple response data. Since the subsequent response data includes the previous response data, when the response data includes an end marker, the response data may be the complete second target response data corresponding to the interactive input data (i.e., the complete response data used for the dialogue interaction).
[0305] As an example, the target service assistant may further include a dynamic parsing engine, which may be used to determine complete second target response data from multiple response data.
[0306] Based on this, as an example, as shown in FIG10 , the question-answering method may include the following steps:
[0307] SA1: Receive the first response data streamed from the AI model and conduct a conversational interaction with the first response data.
[0308] SA2, receiving the i-th (i≥2) response data streamed from the artificial intelligence large model;
[0309] SA3. Determine whether the session identifiers of the i-th response data and the (i-1)-th response data are the same. If so, execute S24; if not, execute S26.
[0310] SA4. Record the index value of the i-th response data;
[0311] SA5. Based on the additional broadcast strategy, determine the portion of the i-th response data excluding the (i-1)-th response data as the first target response data, and perform a dialog interaction on the first target response data;
[0312] SA6. Reset the index value and session identifier of the i-th response data;
[0313] SA7. Based on the interruption broadcast strategy, conduct a dialogue interaction with the i-th response data;
[0314] SA8. Set i=i+1 and return to execute S22 until the response data includes the end marker.
[0315] By generating speech frame-by-frame from the text, speech synthesis becomes a streaming process, outputting sound in real time and playing it back as the speech is generated. This approach makes synthesized speech smoother and reduces latency.
[0316] Thus, in the case where the complete response data includes multiple response data streamed out by the artificial intelligence large model, by conducting a dialogue interaction with the first output response data during the process of the artificial intelligence large model streaming out multiple response data, and starting from the second response data, based on the additional broadcast strategy, a plurality of first target response data are sequentially determined in the multiple response data, and dialogue interactions are conducted with the multiple first target response data respectively, that is, each time a portion of response data is generated, a dialogue interaction is conducted with a portion of response data, which can ensure the timeliness of dialogue interaction with the response data (i.e., responding to user instructions). In addition, by sequentially conducting dialogue interaction with multiple first target response data until the response data is the last one among the multiple response data, it is possible to conduct dialogue interaction with the complete response data. In this way, through the embodiment of the present disclosure, it is possible to conduct dialogue interaction with the complete response data and ensure the timeliness of dialogue interaction (i.e., responding to user instructions), thereby improving the user experience.
[0317] Based on this, in order to improve the user's reading experience of the second target response data, in some embodiments, after determining the response data as the last one of the multiple response data, the following steps may also be included:
[0318] The second target response data is input into the artificial intelligence big model, so as to extract the interaction description information from the second target response data through the artificial intelligence big model, mark the interaction description information in the second target response data, and obtain and display TTS structured data.
[0319] The second target response data is the last one of the multiple response data, and the interaction description information is used to perform structured processing on the second target response data.
[0320] As an example, a schematic diagram of a comparison between plain text and TTS structured data provided by an embodiment of the present disclosure may be shown in FIG3 .
[0321] In this way, by displaying TTS structured data instead of a large paragraph of plain text, the key information can be highlighted, improving the user's reading experience of the second target response data.
[0322] In addition, since the second target response data is the last one among the multiple response data, the second target response data may include a scene identifier. Based on this, in order to ensure the necessity and effectiveness of structured processing, in some embodiments, before the second target response data is input into the artificial intelligence model, the following may also be included:
[0323] It is determined whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist, where the structured whitelist includes multiple user interaction scenarios.
[0324] Based on this, the above-mentioned inputting of the second target response data into the artificial intelligence model may specifically include:
[0325] When the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, the second target response data is input into the artificial intelligence big model.
[0326] In addition, as described above, during the process of the AI model streaming out multiple response data, the first response data and at least one first target response data are streamed onto the screen and displayed in the TTS display area. Therefore, in order to ensure both the timeliness of the response data display and the user's reading experience of the response data, and further enhance the user experience, in some embodiments, the above-mentioned display of TTS structured data may specifically include:
[0327] The displayed first response data and at least one target response data are replaced with TTS structured data.
[0328] Here, the process of streaming output of multiple response data, the process of streaming response data to the screen, and the process of TTS structuring can be carried out simultaneously without affecting each other. That is, after the artificial intelligence large model outputs the last response data (i.e., the second target response data), the streaming of the response data to the screen may not be over yet. On the one hand, the response data can continue to be streamed to the screen. On the other hand, the TTS structuring engine can perform TTS structuring on the second target response data. After obtaining the TTS structured data, the streaming of the response data to the screen may or may not have ended, which is not limited here. If the streaming to the screen has ended, a large section of plain text displayed on the screen can be replaced with TTS structured data. If the streaming to the screen has not ended, the first response data and at least one first target response data displayed on the screen can be replaced with TTS structured data.
[0329] By displaying each portion of response data as it is generated, the timeliness of response data display can be ensured. By replacing the displayed streaming data with TTS structured data after generating it, the user's reading experience of the response data can be enhanced. This ensures both the timeliness of response data display and the user's reading experience, further enhancing the user experience.
[0330] In addition, the process of streaming multiple response data from a large AI model may also include:
[0331] For each response data, determine whether the response data is in the instruction whitelist; if the response data is in the instruction whitelist, intercept the control instruction included in the response data; if the control instruction is intercepted, execute the control instruction.
[0332] The present disclosure also provides another question-answering method, which may include the following steps:
[0333] Step B1: Receive user interaction input data, where the interaction input data includes multimodal feature data and instruction interaction data.
[0334] Step B2: Input the interactive input data into the artificial intelligence big model so that the artificial intelligence big model can identify and respond to the interactive input data and obtain and output multiple response data.
[0335] Step B3: In the process of the artificial intelligence big model outputting multiple response data, determine the target paragraph based on the output multiple response data; input the target paragraph into the artificial intelligence big model to obtain the card information corresponding to the target paragraph through the artificial intelligence big model; render and display the card information corresponding to the target paragraph.
[0336] The question-answering method of the embodiment of the present application determines the target paragraph based on the output multiple response data during the process of the artificial intelligence big model outputting multiple response data, and then inputs the target paragraph into the artificial intelligence big model, and uses the artificial intelligence big model to obtain the card information corresponding to the target paragraph, and can determine the key display information corresponding to the target paragraph. In this way, by rendering and displaying the card information corresponding to the target paragraph, the key display information can be displayed. Compared with displaying a large paragraph of plain text and pictures, it can reduce the user's reading burden on the response data, allowing the user to clearly and accurately obtain the desired information, thereby improving the user's reading experience of the response data.
[0337] The specific implementation methods of the above steps are introduced below.
[0338] In some embodiments, in B1, the interactive input data may be multimodal instruction interaction data input by the user. That is, the interactive input data may include multimodal feature data and instruction interaction data. Among them, the multimodal feature data may characterize the input form of the interactive input data. The input form of the interactive input data may include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interaction data may characterize the user intention corresponding to the interactive input data. The instruction interaction data may include, for example, control instructions related to vehicle control and vehicle settings, question instructions related to vehicle Q&A and general Q&A, and may also include both control instructions and question instructions, which are not limited here.
[0339] As an example, the question-answering system may include a main interaction module and a large model module. Among them, the main interaction module may have an Automatic Speech Recognition (ASR) function. If the input form of the interactive input data is voice input, the main interaction module can convert the voice instruction into instruction text through the ASR function and send the instruction text to the large model module. After receiving the instruction text, the large model module can perform semantic analysis on the instruction text through the artificial intelligence large model to obtain instruction interaction data (i.e., determine the user's intention), and return the instruction interaction data to the main interaction module.
[0340] In addition, the question-and-answer system may also include a voice user interface (VUI) corresponding to the main interaction module. The voice interaction interface may include a text input box. Users can enter text instructions in the text input box. The main interaction module may determine the instruction interaction data corresponding to the interaction input data based on the correspondence between the text instructions and the instruction interaction data.
[0341] In addition, the voice interaction interface may also include multiple interactive controls, which may include preset commands related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. Users can perform touch commands by clicking on the interactive controls. The main interaction module may determine the command interaction data corresponding to the interactive input data based on the correspondence between the touch commands and the command interaction data.
[0342] In addition, if the voice interaction interface supports gesture air touch, the user can select any preset command by gesture air control interaction control. The main interaction module can determine the command interaction data corresponding to the interaction input data based on the correspondence between the gesture command and the command interaction data.
[0343] In some embodiments, in B2, artificial intelligence (AI) large models refer to high-performance AI models built using a large number of training samples and computing resources. They can learn a large amount of language knowledge, image features, and speech patterns, and can reason and generate outputs similar to humans. They have wide applications in natural language processing, image recognition, speech recognition, and other fields. AI large models may include, for example, large language models (LLMs), ChatGPTs (Chat Generative Pre-trained Transformers), multimodal large models, and multimodal cognitive large models.
[0344] As an example, if the input form of the interactive input data is voice input, the main interactive module can send the interactive input data to the large model module after receiving the user's interactive input data. The large model module can input the interactive input data into the artificial intelligence large model, and the artificial intelligence large model recognizes and responds to the interactive input data. Among them, recognizing the interactive input data can generate instruction interaction data. Responding to the interactive input data can generate response data corresponding to the instruction interaction data. If the input form of the interactive input data is non-voice input (including text input, touch input and gesture input), the main interactive module can determine the instruction interaction data corresponding to the interactive input data after receiving the user's interactive input data, and send the instruction interaction data to the large model module. The large model module can input the instruction interaction data into the artificial intelligence large model, and respond to the instruction interaction data through the artificial intelligence large model. Among them, responding to the instruction interaction data can generate response data corresponding to the instruction interaction data.
[0345] As an example, if the large model module is located on the cloud server, and the main interaction module and the dialogue management module are located on the terminal server, then after the terminal server sends the interactive input data to the cloud server through the main interaction module, it can continue to send an asynchronous generation interface to the cloud server so that the large model module returns the instruction interaction data to the terminal server. After the large model module returns the instruction interaction data to the terminal server, it can continue to obtain the response data corresponding to the interactive input data (including the instruction interaction data) and cache the obtained response data. Among them, obtaining the response data corresponding to the interactive input data includes searching for the response data corresponding to the interactive input data in a pre-set question and answer library, searching for the response data corresponding to the interactive input data in the network, and generating the response data corresponding to the interactive input data. In the process of the artificial intelligence large model generating the response data corresponding to the interactive input data, when predicting the next word or character, the previously generated content will be considered. This context-aware generation method enables the artificial intelligence large model to generate text with coherence and a certain logic.
[0346] Furthermore, after receiving the command interaction data, the main interaction module in the terminal server may send the command interaction data to the dialogue management module. After receiving the command interaction data, the dialogue management module may determine the user interaction scenario corresponding to the command interaction data based on an arbitration strategy. User interaction scenarios may include question-and-answer scenarios and interactive control scenarios. Specifically, question-and-answer scenarios may include general question-and-answer scenarios and vehicle question-and-answer scenarios. Interactive control scenarios may include vehicle control scenarios. Furthermore, the arbitration strategy may be a pre-set correspondence between command interaction data and user interaction scenarios.
[0347] If the dialogue management module determines that the user interaction scenario is a target Q&A scenario, it may send the target Q&A scenario to the service management module. Specifically, the target Q&A scenario can be any of a number of Q&A scenarios, such as a general Q&A scenario, a vehicle Q&A scenario, an encyclopedia Q&A scenario, a comparison Q&A scenario, or a plan-making scenario. After receiving the target Q&A scenario, the service management module may determine a target service assistant corresponding to the target Q&A scenario based on the correspondence between the user interaction scenario and the service assistant. The service assistant may correspond to the user interaction scenario and be responsible for executing subsequent actions corresponding to the response data. The target service assistant may be a service assistant corresponding to the target Q&A scenario. For example, the target service assistant may be an AI-based service assistant. An AI-based service assistant may specifically include a dynamic parsing engine and a rich media structuring engine. After determining the target service assistant, the service management module may notify the target service assistant to register the target Q&A scenario and obtain a scenario identifier. The target service assistant may then send the scenario identifier to the service management module. Upon receiving the scenario identifier, the service management module may record the correspondence between the target service assistant and the scenario identifier and send the scenario identifier to the dialogue management module. The dialogue management module can send a streaming data request corresponding to the scene identifier to the large model module. The streaming data request can be used to request the large model module to stream-output multiple response data corresponding to the interactive input data. After receiving the streaming data request, the large model module can respond to the streaming data request and stream-output multiple response data. The latter response data may include the previous response data. For example, if the interactive input data is "Introduce Zhou xx," the multiple response data streamed out may include: the first response data is "Zhou xx," the second response data is "Zhou xx, from xxx," and the third response data is "Zhou xx, from xxx, won x award on x year x month x day..." In addition, the response data may include a scene identifier.
[0348] As a more specific example, if the large model module is located in a cloud server, and the main interaction module and the dialogue management module are located in a terminal server, the dialogue management module can send a streaming data request to the large model module through the main interaction module.
[0349] In addition, in some embodiments, the above B2 may specifically include:
[0350] The interactive input data is input into the artificial intelligence big model to identify the interactive input data through the artificial intelligence big model, obtain the reply type corresponding to the interactive input data, and determine the display template information corresponding to the reply type; and the artificial intelligence big model responds to the interactive input data to obtain and output multiple response data including display template information.
[0351] Here, the presentation template information may be template type information corresponding to the presentation template for the response data. Template types may include description templates, general classification templates, comparison templates, timeline templates, and service expert templates. The service expert template may be a presentation template corresponding to vehicle Q&A. The text description in the service expert template may include operational steps related to vehicle service usage.
[0352] As an example, after receiving interactive input data, the large model module can also determine the response type corresponding to the interactive input data through the artificial intelligence large model, and determine the template type corresponding to the response data of the interactive input data based on the correspondence between the response type and the display template information, obtain the display template information, and stream output multiple response data including the display template information. The display template information included in the multiple response data can be the same.
[0353] In some embodiments, in step B3, the dynamic parsing engine can determine whether to perform rich media structuring on the response data, and which portion of the response data to perform rich media structuring on. Specifically, the dynamic parsing engine can be used to determine target segments for rich media structuring based on the multiple response data streamed out. The target segments can be one or more, and this is not limited here.
[0354] In addition, as described above, among the multiple response data output, the next response data may include the previous response data. Based on this, in some embodiments, the above-mentioned determination of the target paragraph based on the multiple response data output may specifically include:
[0355] For each of the plurality of response data outputted, determining whether the response data includes a preset separator;
[0356] In a case where the response data includes a preset delimiter, determining the response data as a target paragraph;
[0357] In the case that the response data includes a plurality of preset delimiters, the response data between the last two adjacent preset delimiters among the plurality of preset delimiters is determined as the target paragraph.
[0358] Here, the multiple response data outputted may be multiple response data that have already completed streaming output. The dynamic parsing engine may include a delimiter interceptor. The delimiter interceptor may determine whether the response data includes a preset delimiter. If the response data does not include the preset delimiter, the interceptor may continue to determine whether the next response data of the response data includes the preset delimiter. If the response data includes the preset delimiter, the target paragraph for rich media structuring may be determined based on the preset delimiter.
[0359] Specifically, the preset delimiter can be represented by, for example, <|br / >. If the response data is "AAA.BBB.<|br / >," the response data includes one pre-delimiter, and the target paragraph can be "AAA.BBB." If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >," the response data includes two preset delimiters, and the target paragraph can be "CCC.DDD." If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >EEE.FFF.<|br / >," the response data includes three preset delimiters, and the target paragraph can be "EEE.FFF."
[0360] It should be noted that in the embodiments of the present application, the first and second processes can be performed simultaneously without affecting each other. Specifically, the first process can be the process of streaming multiple response data from the artificial intelligence model. The second process can be the process of the dynamic parsing engine determining the target paragraph based on the multiple response data that have been streamed out.
[0361] In addition, as mentioned above, the response data may include a scene identifier. Based on this, in order to ensure the necessity and effectiveness of structured processing, in some embodiments, before determining whether the response data includes a preset delimiter, the following may also be included:
[0362] It is determined whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist, where the structured whitelist includes multiple user interaction scenarios.
[0363] Based on this, the above determination of whether the response data includes the preset delimiter may specifically include:
[0364] When the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, it is determined whether the response data includes a preset separator.
[0365] Here, scenario identifiers can correspond one-to-one with user interaction scenarios. User interaction scenarios can include question-and-answer scenarios and interactive control scenarios. Specifically, question-and-answer scenarios can include general question-and-answer scenarios, vehicle question-and-answer scenarios, encyclopedia question-and-answer scenarios, comparison question-and-answer scenarios, and plan-making scenarios. Specifically, interactive control scenarios can include vehicle control scenarios, smart device control scenarios, and mobile device control scenarios.
[0366] In addition, the structured whitelist can be used to determine whether to perform structured processing on the response data. If the response data is in the structured whitelist, it can be determined that the response data is structured. If the response data is no longer in the structured whitelist, it can be determined that the response data is not structured.
[0367] As an example, a structured whitelist may include multiple user interaction scenarios. The multiple user interaction scenarios in the structured whitelist may include, for example, general Q&A scenarios, vehicle Q&A scenarios, encyclopedia Q&A scenarios, comparison Q&A scenarios, plan-making scenarios, and other Q&A scenarios. If the user interaction scenario corresponding to the scenario identifier in the response data is in the structured whitelist, it can be determined that the response data is in the structured whitelist. If the user interaction scenario corresponding to the scenario identifier in the response data is not in the structured whitelist, it can be determined that the response data is not in the structured whitelist.
[0368] As an example, the dialogue management module can receive multiple response data streamed from the large model module via the main interaction module. After receiving the response data, the dialogue management module can send the response data, including the scenario identifier, to the service management module. The service management module can route the response data to the target service assistant corresponding to the scenario identifier based on the correspondence between the scenario identifier and the target service assistant. After receiving the response data, the target service assistant can perform subsequent structural processing using the dynamic parsing engine within the target service assistant.
[0369] Specifically, each time the dynamic parsing engine receives a response, it determines whether the user interaction scenario corresponding to the scenario identifier in the response is in the structured whitelist. If not, it proceeds to the next response. If it is, it determines that the response meets the preliminary conditions for structured processing, and then proceeds to determine whether the response includes the preset delimiter.
[0370] In this way, by further determining whether the response data includes the preset separator when the response data is in the structured whitelist, the response data can be reasonably arranged to ensure the necessity and effectiveness of structured processing.
[0371] In addition, the response data may also include a word count threshold. The word count threshold may be a pre-set minimum number of words corresponding to the response data, used to determine whether the response data is structured. Based on this, in order to avoid wasting resources and ensure the rational use of resources, in some embodiments, when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, determining whether the response data includes a preset delimiter may specifically include:
[0372] When the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, intercept and record the presentation template information and word count threshold in the response data;
[0373] Get the number of words in the response data;
[0374] Determine the relationship between the number of words and the word count threshold;
[0375] When the word count is greater than the word count threshold, it is determined whether the response data includes a preset delimiter.
[0376] Here, the dynamic parsing module may include a template interceptor. The template interceptor can intercept and record information such as the display template information, word count threshold, and source information of the response data in the response data. Template types may include description templates, overall score templates, comparison templates, timeline templates, and service expert templates. A service expert template can be a display template corresponding to vehicle Q&A. For example, a service expert template may include operational steps related to vehicle usage services.
[0377] In addition, the dynamic parsing module may also include a word count interceptor. The word count interceptor can obtain the number of words included in the response data (the length of the response data) and determine the size relationship between the length and the word count threshold. If the length is less than the word count threshold, no subsequent processing is required. If the length is greater than the word count threshold, it is possible to continue to determine whether the response data includes a preset delimiter. That is, it is necessary to continue to determine whether structured processing is required based on the response data. In actual situations, if the word count is less than the word count threshold, it can be considered that the response data is used for a brief conversation with the user (such as chatting), and the response data may not be structured and displayed.
[0378] In this way, by continuing to determine whether structured processing is required based on the response data when the word count is greater than the word count threshold, resource waste can be avoided and reasonable utilization of resources can be ensured.
[0379] Based on the above embodiments, a specific example is given. For example, the specific process of determining the target paragraph through the dynamic parsing engine can be shown in Figure 4. As shown in Figure 4, the dynamic parsing engine can include an instruction interceptor, a template interceptor, a word count interceptor, a delimiter interceptor, and a structured interceptor.
[0380] The present disclosure also provides a response method, which may include the following steps:
[0381] Step C1: receiving user interaction input data, where the interaction input data includes multimodal feature data and instruction interaction data.
[0382] Step C2: Input the interactive input data into the artificial intelligence model, so that the artificial intelligence model recognizes and responds to the interactive input data, and obtains and outputs multiple response data, wherein the latter response data includes the previous response data.
[0383] Step C3: During the process of the artificial intelligence large model streaming out multiple response data, for each response data, determine whether the response data is in the structured whitelist; if the response data is in the structured whitelist, perform structured processing on the response data to obtain and display the structured data.
[0384] The question-answering method of the embodiment of the present application inputs interactive input data into an artificial intelligence big model, and uses the artificial intelligence big model to identify and respond to the interactive input data, so as to obtain and stream-output multiple response data, thereby achieving the effect of one input and multiple outputs. In this way, by, in the process of streaming output of multiple response data, for each response data, judging whether the response data is in a structured whitelist, and when the response data is in a structured whitelist, performing structured processing on the response data, the process of parsing the response data (including judging whether the response data is in a structured whitelist and performing structured processing on the response data) and outputting the response data do not affect each other, thereby improving the parsing efficiency of multiple response data. By displaying the structured data obtained by the structured processing of the response data, it is possible to obtain a portion of the response data each time, that is, to display a portion of the structured data corresponding to the response data, thereby improving the response speed to user instructions and enhancing the user experience.
[0385] The specific implementation methods of the above steps are introduced below.
[0386] In some embodiments, in C1, the interactive input data may be multimodal instruction interaction data input by the user. That is, the interactive input data may include multimodal feature data and instruction interaction data. Among them, the multimodal feature data may characterize the input form of the interactive input data. The input form of the interactive input data may include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interaction data may characterize the user intention corresponding to the interactive input data. The instruction interaction data may include, for example, control instructions related to vehicle control and vehicle settings, question instructions related to vehicle Q&A and general Q&A, and may also include both control instructions and question instructions, which are not limited here.
[0387] As an example, the question-answering system may include a main interaction module and a large model module. Among them, the main interaction module may have an Automatic Speech Recognition (ASR) function. If the input form of the interactive input data is voice input, the main interaction module can convert the voice instruction into instruction text through the ASR function and send the instruction text to the large model module. After receiving the instruction text, the large model module can perform semantic analysis on the instruction text through the artificial intelligence large model to obtain instruction interaction data (i.e., determine the user's intention), and return the instruction interaction data to the main interaction module.
[0388] In addition, the question-and-answer system may also include a voice user interface (VUI) corresponding to the main interaction module. The voice interaction interface may include a text input box. Users can enter text instructions in the text input box. The main interaction module may determine the instruction interaction data corresponding to the interaction input data based on the correspondence between the text instructions and the instruction interaction data.
[0389] In addition, the voice interaction interface may also include multiple interactive controls, which may include preset commands related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. Users can perform touch commands by clicking on the interactive controls. The main interaction module may determine the command interaction data corresponding to the interactive input data based on the correspondence between the touch commands and the command interaction data.
[0390] In addition, if the voice interaction interface supports gesture air touch, the user can select any preset command by gesture air control interaction control. The main interaction module can determine the command interaction data corresponding to the interaction input data based on the correspondence between the gesture command and the command interaction data.
[0391] In some embodiments, in C2, AI large models refer to high-performance artificial intelligence models built using a large amount of training samples and computing resources. They can learn a large amount of language knowledge, image features, and speech patterns, and can reason and generate outputs similar to humans. They have wide applications in natural language processing, image recognition, speech recognition, and other fields. AI large models may include, for example, large language models (LLM), ChatGPT (Chat Generative Pre-trained Transformer), multimodal large models, and multimodal cognitive large models.
[0392] In addition, the AI model can identify and respond to interactive input data. Specifically, the AI model can first perform semantic recognition on the interactive input data, determine the instruction interaction data corresponding to the interactive input data, and then respond to the instruction interaction data to obtain a complete response text. Specifically, after determining the instruction interaction data, the AI model can, on the one hand, return the instruction interaction data to the main interaction module. On the other hand, it can continue to respond to the instruction interaction data to obtain a complete response text.
[0393] As an example, the question-and-answer system may also include a dialogue management module. After receiving the instruction interaction data, the main interaction module may send the instruction interaction data to the dialogue management module. The dialogue management module may determine the user interaction field corresponding to the instruction interaction data based on a pre-set arbitration strategy. The user interaction field may include a question-and-answer field and a vehicle control field. If the dialogue management module determines that the field corresponding to the instruction interaction data is a question-and-answer field, it may send a data response request to the large model module. The data response request may be used to request the AI large model to respond to the instruction interaction data. At this time, if the AI large model has obtained a complete response text corresponding to the instruction interaction data, it may stream-output multiple response data corresponding to the complete response text. If the AI large model has obtained a partial response text corresponding to the instruction interaction data, the AI large model may, on the one hand, stream-output multiple response data corresponding to the partial response text, and on the other hand, continue to respond to the instruction interaction data, without affecting each other.
[0394] In the multiple response data streamed out by the AI large model, the subsequent response data can include the previous response data. For example, if the instruction interaction data is "Introduce Zhou xx," the multiple response data streamed out may include: the first response data is "Zhou xx," the second response data is "Zhou xx, from xxx," and the third response data is "Zhou xx, from xxx, won x award on x-month-x-year..."
[0395] In some embodiments, at C3, the structured whitelist may include multiple user interaction scenarios. These scenarios may include, for example, general Q&A scenarios, vehicle Q&A scenarios, encyclopedia Q&A scenarios, comparison Q&A scenarios, and plan-making scenarios. Furthermore, the response data may carry a scenario tag. If the scenario tag carried in the response data matches any of the aforementioned user interaction scenarios, the response data may be determined to be in the structured whitelist.
[0396] As an example, the question-answering system may also include a dynamic parsing module. Every time the AI large model outputs a response data, the response data may be parsed by the dynamic parsing module until the multiple response data streamed by the AI large model are parsed. The parsing may include determining whether the response data is in a structured whitelist. If the response data is not in the structured whitelist, it may continue to determine whether the next response data is in the structured whitelist. If the response data is in the structured whitelist, the response data may be structured to obtain and display structured data. Specifically, the structured processing may include rich media structured processing. Rich media structured processing may be to generate card information in rich media format based on the response data. That is, the structured data may include card information in rich media format. The card information in rich media format may include audio, text, links, interactive controls, pictures, etc. corresponding to the response data.
[0397] As a more specific example, the dynamic parsing module can include interceptors and executors. Among them, interceptors are an important component in Java, used to pre-process, post-process and handle exceptions before, after or in abnormal situations before the request reaches the target method. Interceptors are an aspect-oriented programming (AOP) technology that can be used to encapsulate and reuse some common operations, reduce code redundancy, and improve code maintainability. The response data can be parsed through interceptors and executors.
[0398] Based on this, in order to ensure the timeliness of executing user instructions and improve user experience, in some embodiments, before determining whether the response data is in the structured whitelist, the following steps may also be performed:
[0399] Determine whether the response data is in the command whitelist;
[0400] If the response data is in the instruction whitelist, intercepting the control instruction included in the response data;
[0401] When a control instruction is intercepted, the control instruction is executed.
[0402] Based on this, determine whether the response data is in the structured whitelist, which may include:
[0403] determining the response data other than the control instruction in the response data as target response data;
[0404] Determine whether the target response data is in the structured whitelist.
[0405] Here, the command whitelist may include multiple control commands that the vehicle can execute. For example, these multiple control commands may include window control commands, air conditioning control commands, seat control commands, and display control commands. If the control command in the response data matches any control command in the command whitelist, then the response data can be determined to be on the command whitelist. For example, if the control command in the response data is "open window," then it can be determined that the control command matches the window control command, and further, the response data including the control command can be determined to be on the command whitelist.
[0406] As an example, the dynamic parsing module may include an instruction interceptor, which may intercept control instructions included in the response data. If the dynamic parsing module has determined that the response data includes a control instruction, the response data may be sent to the instruction interceptor. The instruction interceptor may intercept the control instruction in the response data. If the instruction interceptor is able to intercept the control instruction, the control instruction may be executed by the instruction executor.
[0407] In this way, by intercepting and executing the control instructions in each response data when the response data is in the instruction whitelist, the timeliness of executing user instructions can be guaranteed and the user experience can be improved.
[0408] Based on this, in order to ensure the timeliness of card information display and improve user experience, in some embodiments, the above S140 may further include:
[0409] If the response data is in the structured whitelist, determining whether the response data includes a preset separator;
[0410] In a case where the response data includes a preset delimiter, determining the target paragraph based on the preset delimiter;
[0411] Use artificial intelligence models to obtain key information and images corresponding to the target paragraph;
[0412] Render key information and images, generate and display card information.
[0413] Here, the dynamic parsing module may also include a delimiter interceptor. The delimiter interceptor may determine whether the response data includes a preset delimiter. If the response data does not include the preset delimiter, the next response data may be determined to determine whether the preset delimiter is included. If the response data includes the preset delimiter, the target paragraph for rich media structuring may be determined based on the preset delimiter. In addition, the card information may be card information in a rich media format. The rich media card information may include audio, text, links, interactive controls, images, etc. corresponding to the target paragraph.
[0414] Furthermore, the dynamic parsing module may also include a structured interceptor. This interceptor can segment the response data based on preset delimiters to obtain target paragraphs, and then use the segmented target paragraphs as structured request parameters for a rich media structured request. In addition to the target paragraph, the structured request parameters may also include the target paragraph's number within the entire response text and a paragraph identifier.
[0415] Based on this, in some embodiments, determining the target paragraph based on a preset delimiter may specifically include:
[0416] In a case where the response data includes a preset delimiter, determining the response data as a target paragraph;
[0417] In the case that the response data includes a plurality of preset delimiters, the response data between the last two adjacent preset delimiters among the plurality of preset delimiters is determined as the target paragraph.
[0418] Here, the preset delimiter can be represented as <|br / >, for example. If the response data is "AAA.BBB.<|br / >," the response data includes one pre-delimiter, and the target paragraph can be "AAA.BBB." If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >," the response data includes two preset delimiters, and the target paragraph can be "CCC.DDD." If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >EEE.FFF.<|br / >," the response data includes three preset delimiters, and the target paragraph can be "EEE.FFF."
[0419] After determining the target paragraph, the structured interceptor can send a rich media structured request to the large model module. This rich media structured request can request the AI large model to obtain the key information and images corresponding to the target paragraph. If the key information and images are obtained, the large model module can return them to the dynamic parsing module.
[0420] As an example, the response system may further include a rich media structuring engine that can render key information and images, and generate and display card information.
[0421] As an example, the AI big model may or may not obtain valid data. Valid data may include at least one of key information and images. Key information may include, for example, a title, subtitle, and summary. The image may be returned in the form of a picture link. In addition, the valid data may also include the target paragraph's number in the entire response text, as well as a paragraph identifier. If the AI big model returns valid data, the rich media structuring engine may render the valid data, generate, and display card information.
[0422] Based on this, in some embodiments, the above-mentioned rendering of key information and images to generate and display card information may specifically include:
[0423] Determine whether valid data is obtained, where valid data is at least one of key information and images;
[0424] When valid data is obtained, the valid data is rendered to generate and display card information.
[0425] If the AI model returns valid data, the rich media structured engine can pull up a Hypertext Markup Language interface (e.g., HTML 5, H5) and transparently pass the valid data to H5 via a JS bridge. The client can load the H5 and valid data through a browser (e.g., a webview container) to generate and display the card information.
[0426] In this way, by determining the target paragraph based on the preset delimiter and generating and displaying the card information corresponding to the target paragraph, the user's reading experience of the displayed response data can be improved. On the other hand, by generating and displaying the card information once for each response data containing the preset delimiter, the timeliness of the card information display can be guaranteed, thereby improving the user experience.
[0427] Based on this, in order to avoid resource waste and ensure reasonable utilization of resources, in some embodiments, when the response data is in the structured whitelist, determining whether the response data includes a preset separator may specifically include:
[0428] When the response data is in the structured whitelist, the template type and word count threshold in the response data are intercepted and recorded;
[0429] Get the number of words in the response data;
[0430] Determine the relationship between the number of words and the word count threshold;
[0431] When the word count is greater than the word count threshold, it is determined whether the response data includes a preset delimiter.
[0432] Here, the dynamic parsing module may include a template interceptor. The template interceptor can intercept and record information such as the template type, word count threshold, and source information of the response data. Template types may include description templates, overall score templates, comparison templates, timeline templates, and service expert templates. A service expert template may be a display template corresponding to vehicle Q&A. For example, a service expert template may include operational steps related to vehicle usage services.
[0433] In addition, the dynamic parsing module may also include a word count interceptor. The word count interceptor can obtain the number of words included in the response data (the length of the response data) and determine the size relationship between the length and the word count threshold. If the length is less than the word count threshold, no subsequent processing is required. If the length is greater than the word count threshold, it is possible to continue to determine whether the response data includes a preset delimiter. That is, it is necessary to continue to determine whether structured processing is required based on the response data. In actual situations, if the word count is less than the word count threshold, it can be considered that the response data is used for a brief conversation with the user (such as chatting), and the response data may not be structured and displayed.
[0434] In this way, by continuing to determine whether structured processing is required based on the response data when the word count is greater than the word count threshold, resource waste can be avoided and reasonable utilization of resources can be ensured.
[0435] Based on this, in order to further improve the display effect of card information and enhance the user's visual experience, in some embodiments, the above rendering of key information and images to generate and display card information may specifically include:
[0436] Determine whether valid data is obtained, where valid data is at least one of key information and images;
[0437] When valid data is obtained, it is rendered according to the template type to generate and display card information.
[0438] Here, if the AI model does not return valid data, no subsequent processing is required. If the AI model returns valid data, the dynamic parsing module can pull up a hypertext markup language interface (such as HTML 5, H5) and pass the valid data to H5 through the JS bridge. The client can use the browser (such as a webview container) to pull up the H5 interface corresponding to the template type, load the H5 template and valid data, and generate and display the card information.
[0439] In this way, by using templates and data to dynamically generate the user interface (i.e., card information), the structure and style of the user interface can be separated from the data, so that different user interfaces can be dynamically generated according to different data, further improving the display effect of the card information and enhancing the user's visual experience.
[0440] In addition, the AI model can detect the network and obtain the network status every time it outputs response data. Therefore, the response data can also include the network status.
[0441] Based on this, in some embodiments, when the word count is greater than the word count threshold, determining whether the response data includes a preset delimiter may specifically include:
[0442] If the word count is greater than a preset threshold, the network status in the response data is intercepted;
[0443] When the network status is normal, it is determined whether the response data includes a preset separator.
[0444] Based on this, in order to further improve the user experience, in some embodiments, when the word count is greater than the word count threshold, and before determining whether the response data includes a preset separator, the following steps may also be included:
[0445] Intercept the network status in the response data;
[0446] When the network status is abnormal, the target control is displayed, and the target control is used to regenerate the card information.
[0447] The dynamic parsing module may also include a network anomaly interceptor. This interceptor can intercept the network status in the response data. If the network status indicates an anomaly, a target control for regenerating card information can be displayed in the voice interaction interface. After seeing the target control, the user can determine that the network anomaly exists and choose whether to regenerate the card information.
[0448] In this way, by providing users with an interactive method to regenerate and display card information when card information display fails, the user experience can be further improved.
[0449] In order to better describe the entire solution, some specific examples are given based on the above embodiments.
[0450] For example, as shown in FIG11 , a schematic diagram of a dynamic parsing process provided by an embodiment of the present disclosure may include the following steps:
[0451] SC1, receives the i-th response data (i≥1) output by the AI large model in streaming format;
[0452] SC2. Determine whether the i-th response data is the last response data. If not, execute S23. If so, execute S24.
[0453] SC3. Determine whether the i-th response data is in the instruction whitelist. If so, execute S24; if not, execute S27.
[0454] SC4. Intercept the control instructions in the response data through the instruction interceptor;
[0455] SC5. Determine whether a control instruction is intercepted. If so, execute S26; if not, execute S27.
[0456] SC6, execute control instructions;
[0457] SC7. Determine whether the response data is in the structured whitelist. If so, execute S28; if not, execute S217.
[0458] SC8, intercept and record key information such as template type, word count threshold, source, etc. in the response data through the template interceptor;
[0459] SC9, obtain the word count in the response data through the word count interceptor;
[0460] SC10, determine whether the word count is greater than the word count threshold through the word count interceptor, if so, execute S211, if not, execute SC17;
[0461] SC11. Determine whether the network is abnormal through the network abnormality interceptor. If so, execute S212; if not, execute SC13;
[0462] SC12. Display the target control for regenerating card information;
[0463] SC13. Determine whether the response data includes a delimiter through a delimiter interceptor. If so, execute S214; if not, execute SC17.
[0464] SC14: Use the structured interceptor to determine the target paragraph based on the delimiter, and initiate a structured request corresponding to the target paragraph to the AI big model. The structured request is used to request the AI big model to obtain key information and images corresponding to the target paragraph.
[0465] SC15: Determine whether valid data is obtained, where valid data includes at least one of key information and images. If so, execute SC16; if not, execute SC17.
[0466] SC16. Pull up the H5 interface corresponding to the template type, load valid data, and generate and display a structured interface (i.e., card information in rich media format).
[0467] SC17: Set i=i+1 and return to execute SC2 until the i-th response data is the last response data.
[0468] Therefore, by simultaneously parsing the previous response data and outputting the next response data simultaneously, the AI model can achieve a more efficient parsing of multiple response data. In addition, by displaying a portion of the structured data corresponding to each response data, the speed of responding to user commands can be increased, thus enhancing the user experience.
[0469] Based on the question-answering method provided in the above embodiment, the present disclosure also provides a specific implementation of a question-answering device. Please refer to the following embodiment.
[0470] As shown in FIG12 , the question-answering device 1000 provided in the embodiment of the present disclosure includes the following modules:
[0471] Receiving module 1010, for receiving interactive input data from a user;
[0472] The first processing module 1020 is configured to process the interactive input data using the artificial intelligence model to obtain a generated content information set, where the generated content information set includes user intention description information, response data, and interaction description information;
[0473] A first determining module 1030 is configured to determine a user interaction scenario corresponding to the user intention description information;
[0474] The second determination module 1040 is configured to determine a target service corresponding to the question-and-answer scenario from multiple services when the user interaction scenario is a question-and-answer scenario. The service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interaction input data.
[0475] The question-answering device 1000 is described in detail below.
[0476] In some embodiments, the second determining module 1040 may specifically include:
[0477] A first acquisition submodule is configured to acquire registration intentions corresponding to multiple services respectively;
[0478] The first determination submodule is configured to determine a target service corresponding to the question-and-answer scenario based on a correspondence between a registration intention and a user interaction scenario.
[0479] In some embodiments, the question-answering device further comprises:
[0480] The second processing module 1050 is configured to perform structured processing on the response data based on the interaction description information through the target service, and obtain and display structured data.
[0481] In some embodiments, the second processing module may specifically include:
[0482] The registration submodule is configured to register the question-answering scenario through the target service and obtain a scenario identifier;
[0483] a routing submodule configured to route the response data corresponding to the scenario identifier to the target service;
[0484] The first processing submodule is configured to perform structured processing on the response data based on the interaction description information through the target service, and obtain and display the structured data.
[0485] In some embodiments, the response data includes multiple sub-response data output in a streaming manner. Based on this, the second processing module may further include:
[0486] A second determining submodule is configured to determine first sub-response data from the plurality of sub-response data, where the first sub-response data is response data for performing structured processing;
[0487] A second acquisition submodule is configured to acquire interaction description information corresponding to the first sub-response data using the artificial intelligence large model;
[0488] The second processing submodule is configured to perform structured processing on the first sub-response data based on the interaction description information, and obtain and display structured data.
[0489] In some embodiments, the first sub-response data includes complete target response data corresponding to the interactive input data, and the interaction description information includes first interaction description information corresponding to the target response data. Based on this, the second processing module may specifically include:
[0490] The third processing submodule is configured to perform tagging processing on the first interaction description information in the target response data to obtain TTS structured data.
[0491] In some embodiments, the response data includes a session identifier; and the determining module includes:
[0492] a first judgment submodule configured to judge, starting from the second sub-response data, whether the first session identifier in the sub-response data is the same as the second session identifier in the previous sub-response data;
[0493] The first determining submodule is configured to determine the response data except the previous sub-response data in the response data as the first target response data when the first session identifier is the same as the second session identifier.
[0494] In some embodiments, the determining module further comprises:
[0495] The second determining submodule is configured to determine the sub-response data as the first sub-response data corresponding to the second session identifier when the first session identifier is different from the second session identifier, and return to perform dialog interaction on the first sub-response data.
[0496] In some embodiments, the second interaction module may include:
[0497] The second judgment submodule is configured to judge whether the response data includes an end marker during the process of performing a dialogue interaction on the first target response data;
[0498] The third determining submodule is configured to determine the response data as the last one of the multiple response data when the response data includes an end marker.
[0499] In some embodiments, the conversational interaction includes data presentation associated with the voice interaction. Based on this, the second interaction module also includes a processing submodule; the processing submodule may specifically include:
[0500] The replacing unit is configured to replace the displayed first response data and at least one first target response data with TTS structured data.
[0501] In some embodiments, the first sub-response data includes a target paragraph, the interaction description information includes second interaction description information corresponding to the target paragraph, and the second interaction description information includes image information. Based on this, the second processing module may specifically include:
[0502] a rendering submodule, configured to render the second interaction description information to generate card information;
[0503] The display submodule is configured to display the card message in the card message display area.
[0504] In some embodiments, there are multiple target paragraphs. Based on this, the display submodule may specifically include:
[0505] an updating unit configured to update the displayed card information based on the multiple target paragraphs in the generation order of the multiple target paragraphs;
[0506] The display unit is configured to display the continuously updated card information in the card information display area until the target paragraph includes an end marker and the complete card information is obtained.
[0507] In some embodiments, among the plurality of sub-response data output in a streamed manner, a subsequent sub-response data includes a previous sub-response data. Based on this, the second determining submodule may specifically include:
[0508] a judging unit configured to judge, for each sub-response data, whether the sub-response data includes a preset separator;
[0509] a first determining unit configured to determine the sub-response data as a target paragraph when the sub-response data includes one of the preset delimiters;
[0510] The second determining unit is configured to determine the sub-response data between the last two adjacent preset delimiters as the target paragraph when the sub-response data includes multiple preset delimiters; wherein the first sub-response data includes the target paragraph.
[0511] In some embodiments, the answer data includes multiple sub-answer data output in a streamed manner, wherein the subsequent sub-answer data includes the previous sub-answer data. Based on this, the question-answering device 1000 may further include:
[0512] A first dialogue module is configured to determine a target service corresponding to the question-answering scenario from among multiple services, and then conduct a dialogue interaction on the first output sub-answer data through the target service, wherein the dialogue interaction includes voice interaction and data presentation associated with the voice interaction;
[0513] a third determining module configured to, starting from the second output sub-response data, determine the sub-response data in the sub-response data except the previous sub-response data as the second sub-response data;
[0514] The second dialogue module is configured to perform dialogue interaction on the second sub-response data until the sub-response data becomes the last one of the multiple sub-response data.
[0515] In some embodiments, the first processing module is configured to use an artificial intelligence big model to process interactive input data to obtain a generated content information set.
[0516] In some embodiments, the second interaction module may further include:
[0517] an input submodule configured to input the second target response data into the artificial intelligence macromodel, so as to extract interaction description information from the second target response data through the artificial intelligence macromodel, wherein the second target response data is the last one of the plurality of response data, and the interaction description information is used to perform structured processing on the second target response data;
[0518] The processing submodule is configured to mark the interaction description information in the second target response data to obtain and display TTS structured data.
[0519] In some embodiments, the second target response data includes a scene identifier. Based on this, the second interaction module may further include:
[0520] The third judgment submodule is configured to determine whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist before inputting the second target response data into the artificial intelligence model, and the structured whitelist includes multiple user interaction scenarios.
[0521] Based on this, the input submodule may specifically include:
[0522] The input unit is configured to input the second target response data into the artificial intelligence model when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist.
[0523] In some embodiments, the question-answering device may further include:
[0524] a judgment module configured to judge, for each response data, whether the response data is in the instruction whitelist during the process of streaming output of the plurality of response data;
[0525] an interception module configured to intercept the control instruction included in the response data if the response data is in the instruction whitelist;
[0526] The execution module is configured to execute the control instruction when the control instruction is intercepted.
[0527] In some embodiments, the interactive input data includes multimodal feature data and instruction interaction data, the multimodal feature data represents the input form of the interactive input data, and the instruction interaction data represents the user intention corresponding to the interactive input data;
[0528] In some embodiments, the question-answering device further includes an input module; the input module may specifically include:
[0529] The first input submodule is configured to input the interactive input data into the artificial intelligence big model, so as to identify the interactive input data through the artificial intelligence big model, obtain the reply type corresponding to the interactive input data, and determine the display template information corresponding to the reply type; and respond to the interactive input data through the artificial intelligence big model, and obtain and output multiple response data including display template information.
[0530] In some embodiments, the target paragraph includes presentation template information. Based on this, the input module may specifically include:
[0531] a second input submodule configured to input the target paragraph into the artificial intelligence big model, determine text interaction description information corresponding to the presentation template information through the artificial intelligence big model, extract the text interaction description information from the target paragraph, and search for multimedia interaction description information corresponding to the text interaction description information;
[0532] The determination submodule is configured to determine the text interaction description information and the multimedia interaction description information as card information.
[0533] In some embodiments, the question-answering device further includes a rendering module; the rendering module may specifically include:
[0534] a filling submodule configured to fill valid information into the presentation template corresponding to the presentation template information to obtain card information corresponding to the target paragraph, wherein the valid information includes at least one of text interaction description information and multimedia interaction description information;
[0535] The display submodule is configured to display card information in the card information display area.
[0536] Based on this, the question-answering device further includes a first judgment module, which may specifically include:
[0537] A first determining submodule is configured to determine the response data other than the control instruction in the response data as target response data;
[0538] The first judgment submodule is configured to judge whether the target response data is in the structured whitelist.
[0539] Based on this, the processing module in the question-answering device includes:
[0540] The second judgment submodule is configured to judge whether the sub-response data includes a preset separator when the sub-response data is in the structured whitelist.
[0541] In some embodiments, the determining unit may further include:
[0542] an interception subunit configured to intercept the network status in the response data when the word count is greater than a word count threshold and before determining whether the response data includes a preset delimiter;
[0543] The display subunit is configured to display a target control when the network status is abnormal, and the target control is used to regenerate card information.
[0544] Based on the above embodiments, the present disclosure further provides another question-answering device, including:
[0545] The question-answering device provided in the embodiment of the present disclosure includes the following modules:
[0546] a receiving module configured to receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing the input form of the interaction input data, and the instruction interaction data representing the user intention corresponding to the interaction input data;
[0547] An input module is configured to input interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and streams multiple response data, wherein a subsequent response data includes a previous response data;
[0548] A first interaction module is configured to perform a dialogue interaction on the first response data during the process of streaming the plurality of response data, wherein the dialogue interaction includes voice interaction;
[0549] a determination module configured to, starting from the second response data, determine the response data other than the previous response data in the response data as the first target response data;
[0550] The second interaction module is configured to perform a dialog interaction on the first target response data until the response data is the last one of the multiple response data.
[0551] The question-and-answer device is described in detail below:
[0552] In some embodiments, the response data includes a session identifier. Based on this, the determination module may specifically include:
[0553] A first judgment submodule is configured to judge, starting from the second response data, whether the first session identifier in the response data is the same as the second session identifier in the previous response data;
[0554] The first determining submodule is configured to determine the response data other than the previous response data in the response data as the first target response data when the first session identifier is the same as the second session identifier.
[0555] In some embodiments, the determining module may further include:
[0556] The second determining submodule is configured to determine the response data as the first response data corresponding to the second session identifier when the first session identifier is different from the second session identifier, and return to perform dialog interaction on the first response data.
[0557] In some embodiments, the second interaction module may specifically include:
[0558] The second judgment submodule is configured to judge whether the response data includes an end marker during the process of performing a dialogue interaction on the first target response data;
[0559] The third determining submodule is configured to determine the response data as the last one of the multiple response data when the response data includes an end marker.
[0560] In some embodiments, the second interaction module may further include:
[0561] an input submodule configured to input the second target response data into the artificial intelligence big model, so as to extract interaction description information from the second target response data through the artificial intelligence big model, wherein the second target response data is the last one of the multiple response data, and the interaction description information is configured to perform structured processing on the second target response data;
[0562] The processing submodule is configured to mark the interaction description information in the second target response data to obtain and display TTS structured data.
[0563] In some embodiments, the conversational interaction includes data presentation associated with the voice interaction. Based on this, the processing submodule may specifically include:
[0564] The replacing unit is configured to replace the displayed first response data and at least one first target response data with TTS structured data.
[0565] In some embodiments, the second target response data includes a scene identifier. Based on this, the second interaction module may further include:
[0566] The third judgment submodule is configured to determine whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist before inputting the second target response data into the artificial intelligence model, and the structured whitelist includes multiple user interaction scenarios.
[0567] Based on this, the input submodule may specifically include:
[0568] The input unit is configured to input the second target response data into the artificial intelligence model when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist.
[0569] In some embodiments, the question-answering device may further include:
[0570] a judgment module configured to judge, for each response data, whether the response data is in the instruction whitelist during the process of streaming output of the plurality of response data;
[0571] an interception module configured to intercept the control instruction included in the response data if the response data is in the instruction whitelist;
[0572] The execution module is configured to execute the control instruction when the control instruction is intercepted.
[0573] The question-answering device of the embodiment of the present disclosure inputs interactive input data into an artificial intelligence big model, uses the artificial intelligence big model to identify and respond to the interactive input data, obtains and streams out multiple response data. On the one hand, it can use the natural language generation technology of the artificial intelligence big model to generate response data corresponding to the interactive input data, thereby ensuring the accuracy of the response data. On the other hand, the multiple response data streamed out can be multiple response data obtained after the artificial intelligence big model extracts key information from the complete response data. Therefore, by conducting a dialogue interaction with the first output response data during the process of the artificial intelligence big model streaming out multiple response data, and starting from the second response data, sequentially determining multiple first target response data from the multiple response data, and conducting dialogue interaction with the multiple first target response data respectively, it is possible to broadcast the relatively important response data corresponding to the interactive input data. Compared with broadcasting the complete response data at one time, it is possible to ensure the validity of the response data, so that the user can promptly determine the information they want from the broadcast response data, thereby improving the user experience.
[0574] Based on the above embodiments, the present disclosure further provides another question-answering device, which may include:
[0575] a receiving module configured to receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing the input form of the interaction input data, and the instruction interaction data representing the user intention corresponding to the interaction input data;
[0576] A first input module is configured to input interactive input data into the artificial intelligence big model, so as to recognize and respond to the interactive input data through the artificial intelligence big model, and obtain and output a plurality of response data;
[0577] a determination module configured to determine a target paragraph based on the outputted plurality of response data when the artificial intelligence large model outputs the plurality of response data;
[0578] A second input module is configured to input the target paragraph into the artificial intelligence big model to obtain card information corresponding to the target paragraph through the artificial intelligence big model;
[0579] The rendering module is configured to render and display card information corresponding to the target paragraph.
[0580] The question-and-answer device is described in detail below:
[0581] In some embodiments, the first input module may specifically include:
[0582] The first input submodule is configured to input the interactive input data into the artificial intelligence big model, so as to identify the interactive input data through the artificial intelligence big model, obtain the reply type corresponding to the interactive input data, and determine the display template information corresponding to the reply type; and respond to the interactive input data through the artificial intelligence big model, and obtain and output multiple response data including display template information.
[0583] In some embodiments, the target paragraph includes presentation template information. Based on this, the second input module may specifically include:
[0584] a second input submodule configured to input the target paragraph into the artificial intelligence big model, determine text interaction description information corresponding to the presentation template information through the artificial intelligence big model, extract the text interaction description information from the target paragraph, and search for multimedia interaction description information corresponding to the text interaction description information;
[0585] The determination submodule is configured to determine the text interaction description information and the multimedia interaction description information as card information.
[0586] In some embodiments, the rendering module may specifically include:
[0587] a filling submodule configured to fill valid information into the presentation template corresponding to the presentation template information to obtain card information corresponding to the target paragraph, wherein the valid information includes at least one of text interaction description information and multimedia interaction description information;
[0588] The display submodule is configured to display card information in the card information display area.
[0589] In some embodiments, there are multiple target paragraphs. Based on this, the display submodule may specifically include:
[0590] an updating unit configured to update the displayed card information based on the multiple target paragraphs in the generation order of the multiple target paragraphs;
[0591] The display unit is configured to display the updated card information in the card information display area until the target paragraph includes an end marker, thereby obtaining complete card information.
[0592] In some embodiments, among the multiple response data outputted, the next response data includes the previous response data. Based on this, the determination module may specifically include:
[0593] The first judgment submodule is configured to judge, for each of the plurality of response data outputted, whether the response data includes a preset separator;
[0594] a first determining submodule, configured to determine the response data as a target paragraph when the response data includes a preset delimiter;
[0595] The second determining submodule is configured to, when the response data includes a plurality of preset delimiters, determine the response data between the last two adjacent preset delimiters among the plurality of preset delimiters as the target paragraph.
[0596] In some embodiments, the response data includes a scene identifier. Based on this, the determination module may further include:
[0597] The second judgment submodule is configured to judge whether the response data includes a preset delimiter and whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist, wherein the structured whitelist includes multiple user interaction scenarios.
[0598] Based on this, the first judgment submodule may specifically include:
[0599] The judgment unit is configured to judge whether the response data includes a preset separator when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, and the structured whitelist includes multiple user interaction scenarios.
[0600] In some embodiments, the response data includes display template information and a word count threshold. Based on this, the judgment unit may specifically include:
[0601] a recording subunit configured to intercept and record presentation template information and a word count threshold in the response data when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist;
[0602] an acquiring subunit, configured to acquire the number of words in the response data;
[0603] A first judging subunit is configured to judge the size relationship between the word count and the word count threshold;
[0604] The second judgment subunit is configured to judge whether the response data includes a preset separator when the word count is greater than a word count threshold.
[0605] In some embodiments, the determining unit may further include:
[0606] an interception subunit configured to intercept the network status in the response data when the word count is greater than a word count threshold and before determining whether the response data includes a preset delimiter;
[0607] The display subunit is configured to display the target control when the network status is abnormal, and the target control is configured to regenerate card information.
[0608] The question-answering device of the disclosed embodiment determines the target paragraph based on the output multiple response data during the process of the artificial intelligence big model outputting multiple response data, and then inputs the target paragraph into the artificial intelligence big model, and uses the artificial intelligence big model to obtain the card information corresponding to the target paragraph, and can determine the key display information corresponding to the target paragraph. In this way, by rendering and displaying the card information corresponding to the target paragraph, the key display information can be displayed. Compared with displaying a large paragraph of plain text and pictures, it can reduce the user's reading burden on the response data, allowing the user to clearly and accurately obtain the desired information, thereby improving the user's reading experience of the response data.
[0609] Based on the above embodiments, the present disclosure further provides another question-answering device, including:
[0610] a receiving module configured to receive user interaction input data, the interaction input data including multimodal feature data and instruction interaction data, the multimodal feature data representing the input form of the interaction input data, and the instruction interaction data representing the user intention corresponding to the interaction input data;
[0611] An input module is configured to input interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and streams multiple response data, wherein a subsequent response data includes a previous response data;
[0612] The first judgment module is configured to judge, for each response data, whether the response data is in the structured whitelist during the process of the artificial intelligence large model streaming output of multiple response data;
[0613] The processing module is configured to perform structured processing on the response data when the response data is in the structured whitelist, and obtain and display the structured data.
[0614] The question-and-answer device is described in detail below:
[0615] In some embodiments, the question-answering device may further include:
[0616] a second determination module configured to determine whether the response data is in the instruction whitelist before determining whether the response data is in the structured whitelist;
[0617] an interception module configured to, if the response data is in the instruction whitelist, intercept the control instruction included in the response data;
[0618] The execution module is configured to execute the control instruction when the control instruction is intercepted.
[0619] Based on this, the first judgment module may specifically include:
[0620] A first determining submodule is configured to determine the response data other than the control instruction in the response data as target response data;
[0621] The first judgment submodule is configured to judge whether the target response data is in the structured whitelist.
[0622] In some embodiments, the processing module may specifically include:
[0623] a second judgment submodule, configured to judge whether the response data includes a preset separator when the response data is in the structured whitelist;
[0624] a second determining submodule, configured to determine a target paragraph based on the preset delimiter when the response data includes the preset delimiter;
[0625] an acquisition submodule, configured to utilize an artificial intelligence large model to acquire key information and images corresponding to a target paragraph;
[0626] The rendering submodule is configured to render key information and images, and generate and display card information.
[0627] In some embodiments, the second determining submodule may specifically include:
[0628] a searching unit configured to search for a second delimiter that precedes the response data and is closest to the response data among the plurality of second response data that have been streamed out, the second delimiter being a delimiter corresponding to the second response data;
[0629] a first determining unit configured to determine the response data as a target paragraph when the response data includes a preset delimiter;
[0630] The second determining unit is configured to, when the response data includes a plurality of preset delimiters, determine the response data between the last two adjacent preset delimiters among the plurality of preset delimiters as the target paragraph.
[0631] In some embodiments, the second determination submodule may specifically include:
[0632] a recording unit configured to intercept and record a template type and a word count threshold in the response data if the response data is in a structured whitelist;
[0633] an acquiring unit configured to acquire the number of words in the response data;
[0634] A first judging unit is configured to judge the relationship between the number of words and a word number threshold;
[0635] The second judgment unit is configured to judge whether the response data includes a preset separator when the word count is greater than a word count threshold.
[0636] In some embodiments, the rendering submodule may specifically include:
[0637] a third determining unit configured to determine whether valid data is obtained, where the valid data is at least one of key information and a picture;
[0638] The rendering unit is configured to render the valid data according to the template type when valid data is obtained, and generate and display card information.
[0639] In some embodiments, the second judgment submodule may further include:
[0640] an interception unit configured to intercept the network status in the response data;
[0641] The display unit is configured to display the target control when the network status is abnormal, and the target control is configured to regenerate card information.
[0642] The question-answering device of the disclosed embodiment inputs interactive input data into an artificial intelligence big model, and uses the artificial intelligence big model to identify and respond to the interactive input data, so as to obtain and stream-output multiple response data, thereby achieving the effect of one input and multiple outputs. In this way, by, in the process of streaming out multiple response data, judging whether the response data is in a structured whitelist for each response data, and performing structured processing on the response data when the response data is in a structured whitelist, the process of parsing the response data (including judging whether the response data is in a structured whitelist and performing structured processing on the response data) and outputting the response data do not affect each other, thereby improving the parsing efficiency of multiple response data. By displaying the structured data obtained by the structured processing of the response data, each part of the response data can be obtained, that is, a part of the structured data corresponding to the response data can be displayed, thereby improving the response speed to user instructions and enhancing the user experience.
[0643] Based on the question-answering method provided in the above embodiment, the present disclosure also provides a specific implementation of an electronic device. FIG13 shows a schematic diagram of an electronic device 1100 provided in an embodiment of the present disclosure.
[0644] The electronic device 1100 may include a processor 1110 and a memory 1120 storing computer program instructions.
[0645] Specifically, the processor 1110 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present disclosure.
[0646] The memory 1120 may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 1120 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 1120 may include removable or non-removable (or fixed) media. Where appropriate, the memory 1120 may be internal or external to the electronic device 1100. In a particular embodiment, the memory 1120 is a non-volatile solid-state memory.
[0647] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the question-and-answer method provided according to the present disclosure.
[0648] The processor 1110 implements any one of the question-answering methods in the above embodiments by reading and executing computer program instructions stored in the memory 1120 .
[0649] In one example, the electronic device 1100 may further include a communication interface 1130 and a bus 1140. As shown in FIG11 , the processor 1110, the memory 1120, and the communication interface 1130 are connected via the bus 1140 and communicate with each other.
[0650] The communication interface 1130 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present disclosure.
[0651] Bus 1140 includes hardware, software or both, couples the parts of electronic equipment to each other.For example, and not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 1140 may include one or more buses. Although the present disclosure describes and shows specific bus, the present disclosure considers any suitable bus or interconnection.
[0652] Illustratively, the electronic device 1100 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0653] The electronic device can execute the question-answering method in the embodiment of the present disclosure, thereby realizing the question-answering method, device, and system described in conjunction with Figures 1 to 10.
[0654] In addition, in conjunction with the question-answering method in the above embodiments, the present disclosure may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the question-answering methods in the above embodiments is implemented.
[0655] In addition, the embodiment of the present disclosure further provides a vehicle, which may include at least one of the following:
[0656] The question-answering device as in any of the previous embodiments;
[0657] The question-answering system as in any of the previous embodiments;
[0658] The electronic device as in any of the preceding embodiments;
[0659] The computer-readable storage medium as in any of the preceding embodiments.
[0660] An embodiment of the present disclosure further provides a computer program, which includes computer-readable code. When the computer-readable code runs in an electronic device, the processor of the electronic device executes the computer program to implement any of the above-described question-and-answer methods.
[0661] An embodiment of the present disclosure also provides a computer program product, which includes a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device implements any of the above-described question-and-answer methods when executing the code.
[0662] It should be understood that the present disclosure is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present disclosure is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present disclosure.
[0663] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card, etc. When implemented in software, the elements of the present disclosure are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0664] It should also be noted that the exemplary embodiments described in this disclosure describe methods or systems based on a series of steps or devices. However, this disclosure is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.
[0665] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0666] The above description is only a specific embodiment of the present disclosure. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present disclosure is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the scope of protection of the present disclosure. Industrial Applicability
[0667] The embodiments of the present disclosure provide a question-and-answer method, apparatus, system, device, medium, vehicle, and program product, including: receiving interactive input data from a user; processing the interactive input data to obtain a generated content information set; the generated content information set includes user intention description information, response data, and interaction description information; determining a user interaction scenario corresponding to the user intention description information; when the user interaction scenario is a question-and-answer scenario, determining a target service corresponding to the question-and-answer scenario from multiple services; the service corresponds to the user interaction scenario, and is used to execute subsequent actions corresponding to the interactive input data.
Claims
1. A question-answering method, comprising: Receive interactive input data from the user; Processing the interactive input data to obtain a generated content information set; The generated content information set includes user intention description information, response data and interaction description information; Determining a user interaction scenario corresponding to the user intention description information; In a case where the user interaction scenario is a question-and-answer scenario, a target service corresponding to the question-and-answer scenario is determined from a plurality of services.
2. The method according to claim 1, wherein: The determining of a target service corresponding to the question-answering scenario from among a plurality of services includes: Obtaining registration intentions corresponding to the multiple services respectively; According to the correspondence between the registration intention and the user interaction scenario, a target service corresponding to the question and answer scenario is determined.
3. The method according to claim 1 or 2, wherein: The method further comprises: The response data is structured through the target service based on the interaction description information to obtain and display structured data.
4. The method according to claim 3, wherein: The step of performing structured processing on the response data based on the interaction description information through the target service to obtain and display structured data includes: Registering the question-answering scenario through the target service to obtain a scenario identifier; Routing the response data corresponding to the scenario identifier to the target service; The response data is structured through the target service based on the interaction description information to obtain and display structured data.
5. The method according to claim 3 or 4, wherein: The response data includes a plurality of sub-response data output in a streaming manner, and the structured processing of the response data based on the interaction description information to obtain and display structured data includes: Determining first sub-response data from the plurality of sub-response data; the first sub-response data is response data for performing structured processing; Using the artificial intelligence big model to obtain the interaction description information corresponding to the first sub-response data; The first sub-response data is structured based on the interaction description information to obtain and display structured data.
6. The method according to claim 5, wherein: The first sub-response data includes complete target response data corresponding to the interactive input data, the interactive description information includes first interactive description information corresponding to the target response data, and the first sub-response data is structured based on the interactive description information to obtain and display structured data, including: Marking the first interaction description information in the target response data to obtain TTS structured data; The TTS structured data is displayed in the TTS display area.
7. The method according to claim 6, wherein: The response data includes a session identifier, and the method further includes: Starting from the second sub-response data, determining whether the first session identifier in the sub-response data is the same as the second session identifier in the previous sub-response data; In the case that the first session identifier is the same as the second session identifier, the sub-response data other than the previous sub-response data in the sub-response data is determined as the first target response data.
8. The method according to claim 7, wherein: After determining whether the first session identifier in the sub-response data is the same as the second session identifier in the previous sub-response data, the method further includes: In the case that the first session identifier is different from the second session identifier, the sub-response data is determined to be the first sub-response data corresponding to the second session identifier, and the dialog interaction with the first sub-response data is returned to be executed.
9. The method according to any one of claims 7 or 8, wherein: The method further comprises: During the dialog interaction with the first target response data, determining whether the response data includes an end identifier; In a case where the response data includes an end marker, the response data is determined to be the last one of the plurality of sub-response data.
10. The method according to any one of claims 6 to 9, wherein: The dialog interaction includes data display associated with the voice interaction, displaying TTS structured data, including: The first sub-response data and at least one of the first target response data that have been displayed are replaced with the TTS structured data.
11. The method according to any one of claims 6 to 10, wherein: The first sub-response data includes a target paragraph, the interaction description information includes second interaction description information corresponding to the target paragraph, the second interaction description information includes picture information, and the first sub-response data is structured based on the interaction description information to obtain and display structured data, further comprising: Rendering the second interaction description information to generate card information; The card information is displayed in the card information display area.
12. The method according to claim 11, wherein: The target paragraphs are multiple, and the card information is displayed in the card information display area, including: According to the generation order of the plurality of target paragraphs, the displayed card information is updated based on the plurality of target paragraphs; The card information is continuously updated and displayed in the card information display area until the target paragraph includes an end mark, thereby obtaining the complete card information.
13. The method according to any one of claims 5 to 11, wherein: Among the plurality of sub-response data output in a streaming manner, a later sub-response data includes a previous sub-response data, and determining the first sub-response data among the plurality of sub-response data comprises: For each of the sub-response data, determining whether the sub-response data includes a preset separator; In a case where the sub-response data includes one of the preset separators, determining the sub-response data as the target paragraph; In the case where the sub-response data includes a plurality of the preset delimiters, the sub-response data between the last two adjacent preset delimiters among the plurality of the preset delimiters is determined as the target paragraph; wherein the first sub-response data includes the target paragraph.
14. The method according to any one of claims 1 to 13, wherein: The answer data includes a plurality of sub-answer data output in a streaming manner, wherein a latter sub-answer data includes a former sub-answer data, and after determining a target service corresponding to the question-and-answer scenario from among the plurality of services, the method further includes: Performing a dialog interaction on the first output sub-response data through the target service, wherein the dialog interaction includes voice interaction and data presentation associated with the voice interaction; Starting from the second output sub-response data, the sub-response data other than the previous sub-response data in the sub-response data are determined as second sub-response data; A dialog interaction is performed on the second sub-response data until the sub-response data is the last one of the multiple sub-response data.
15. The method according to any one of claims 9 to 14, wherein: The step of processing the interactive input data to obtain a generated content information set includes: The interactive input data is processed using an artificial intelligence large model to obtain a generated content information set.
16. The method according to claim 15, wherein: After determining the response data as the last one of the plurality of sub-response data, the method further comprises: Inputting the second target response data into the artificial intelligence big model, so as to extract interaction description information from the second target response data through the artificial intelligence big model, wherein the second target response data is the last one of the plurality of sub-response data, and the interaction description information is used to perform structured processing on the second target response data; The interaction description information in the second target response data is marked to obtain and display TTS structured data.
17. The method according to claim 16, wherein: The second target response data includes a scene identifier. Before inputting the second target response data into the artificial intelligence large model, the method further includes: Determining whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist, wherein the structured whitelist includes a plurality of the user interaction scenarios; The step of inputting the second target response data into the artificial intelligence macro model comprises: In the case where the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, the second target response data is input into the artificial intelligence big model.
18. The method according to any one of claims 12 to 17, wherein: The method further comprises: In the process of streaming output of the plurality of sub-response data, for each of the sub-response data, determining whether the sub-response data is in the instruction whitelist; In the case where the sub-response data is in the instruction whitelist, intercepting a control instruction included in the sub-response data; When the control instruction is intercepted, the control instruction is executed.
19. The method according to any one of claims 1 to 11, wherein: The interactive input data includes multimodal feature data and instruction interaction data, wherein the multimodal feature data represents the input form of the interactive input data; and the instruction interaction data represents the user intention corresponding to the interactive input data.
20. The method according to claim 19, wherein: The method further comprises: The interactive input data is input into the artificial intelligence big model so that the interactive input data is identified by the artificial intelligence big model, a reply type corresponding to the interactive input data is obtained, and display template information corresponding to the reply type is determined; and the interactive input data is responded to by the artificial intelligence big model to obtain and output a plurality of response data including the display template information.
21. The method according to claim 20, wherein: The target paragraph includes display template information, and the inputting the target paragraph into the artificial intelligence big model to obtain card information corresponding to the target paragraph through the artificial intelligence big model includes: Inputting the target paragraph into the artificial intelligence big model, determining text interaction description information corresponding to the display template information through the artificial intelligence big model, extracting the text interaction description information from the target paragraph, and searching for multimedia interaction description information corresponding to the text interaction description information; The text interaction description information and the multimedia interaction description information are determined as the card information.
22. The method according to any one of claims 18 to 21, wherein: The determining whether the sub-response data is in the structured whitelist includes: determining the response data other than the control instruction in the response data as target response data; Determine whether the target response data is in the structured whitelist.
23. The method according to claim 22, wherein: The method further comprises: In the case that the sub-response data is in the structured whitelist, it is determined whether the response data includes a preset separator.
24. The method according to claim 23, wherein: In the case where the number of words in the sub-response data is greater than the word number threshold, and before determining whether the sub-response data includes a preset separator, the method further includes: intercepting the network status in the sub-response data; When the network status is a network abnormality, a target control is displayed, and the target control is used to regenerate card information.
25. A question-answering method, comprising: Receive interactive input data from a user; wherein the interactive input data includes multimodal feature data and instruction interaction data, the multimodal feature data represents the input form of the interactive input data, and the instruction interaction data represents the user intention corresponding to the interactive input data; Inputting the interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and streams multiple response data, wherein the latter response data includes the former response data; In the process of streaming the plurality of response data, performing a dialogue interaction on the first response data, wherein the dialogue interaction includes voice interaction; Starting from the second response data, the response data other than the previous response data in the response data are determined as first target response data; A dialog interaction is performed on the first target response data until the response data is the last one of the multiple response data.
26. A question-answering method, comprising: Receive interactive input data from a user, the interactive input data including multimodal feature data and instruction interaction data, the multimodal feature data representing an input form of the interactive input data, and the instruction interaction data representing a user intention corresponding to the interactive input data; Inputting the interactive input data into the artificial intelligence big model, so that the artificial intelligence big model recognizes and responds to the interactive input data, and obtains and outputs a plurality of response data; In the process of the artificial intelligence big model outputting the plurality of response data, determining a target paragraph based on the plurality of output response data; Inputting the target paragraph into the artificial intelligence big model to obtain card information corresponding to the target paragraph through the artificial intelligence big model; Render and display the card information corresponding to the target paragraph.
27. A question-answering device, comprising: A receiving module, configured to receive interactive input data from a user; A first processing module is configured to process the interactive input data using an artificial intelligence big model to obtain a generated content information set; the generated content information set includes user intention description information, response data and interaction description information; A first determining module is configured to determine a user interaction scenario corresponding to the user intention description information; The second determination module is configured to determine a target service corresponding to the question-and-answer scenario from multiple services when the user interaction scenario is a question-and-answer scenario, wherein the service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interaction input data.
28. A question-answering system, comprising: A main interaction module is configured to receive user interaction input data; A large model module is configured to process the interactive input data through an artificial intelligence large model to obtain a generated content information set, wherein the generated content information set includes user intention description information, response data, and interaction description information; a dialogue management module configured to determine a user interaction scenario corresponding to the user intention description information, and to send the user interaction scenario to the service management module, and further configured to determine a scenario identifier corresponding to the response data, and to send the scenario identifier and the corresponding response data to the service management module; The service management module is further configured to determine a target service assistant corresponding to the user interaction scenario sent by the dialogue management module according to a correspondence between a service assistant and a registration intent, and a correspondence between the registration intent and the user interaction scenario, and is further configured to route the response data corresponding to the scenario identifier to the target service assistant according to a correspondence between the scenario identifier and the target service assistant; The service assistant is configured to register the user interaction scenario of interest to the service management module and obtain the registration intention. It is also configured to perform scene registration on the user interaction scenario and obtain the scene identifier, and process the response data corresponding to the scene identifier. The service assistant includes the target service assistant.
29. An electronic device, comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the question-answering method according to any one of claims 1-24, 25 or 26 is implemented.
30. A computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the question-answering method as described in any one of claims 1-24, 25 or 26.
31. A vehicle comprising at least one of the following: The question-and-answer device as claimed in claim 27; The question-answering system as claimed in claim 28; The electronic device as claimed in claim 29; The computer readable storage medium of claim 30.
32. A computer program, comprising a computer-readable code, wherein when the computer-readable code is run in an electronic device, the processor of the electronic device executes the code to implement the question-answering method as described in any one of claims 1 to 24, 25 or 26.
33. A computer program product, comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device implements the question-and-answer method as described in any one of claims 1 to 24, 25 or 26 when executing the code.
Citation Information
Patent Citations
Search result display method and device, computer equipment and storage medium
CN115544375A
Question and answer data acquisition method, equipment, medium and system for intelligent home linkage
CN115941369A
Document processing and response generation system
CN116157790A
Automatic question answering system based on deep reinforcement learning
CN116910213A
Cited By
Voice interaction method and system based on multi-stage large model
CN121011180A
Information processing method and device, electronic equipment, storage medium and program product
CN121166258A