Question and answer method, device and system and vehicle
Through the artificial intelligence big model, processing interactive input data and determining target services is solved, the problem of difficult identification and poor response data display in multiple user instructions is solved, and the problem of higher response accuracy and user experience is achieved.
Patent Information
- Application Number
- CN202311743968.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
When existing Q&A systems receive multiple user instructions, it is difficult to accurately identify Q&A instructions, resulting in low response accuracy and the response data are displayed in plain text, so that key information cannot be displayed, affecting the user's reading experience.
Through the artificial intelligence big model, a collection of content information is generated, including user intention description information, response data and interaction description information, the user interaction scenario is determined, and the target service is determined in the Q&A scenario, and the response data is structured through the target service.
It improves the accuracy of the response to interactive input data, highlights key information through structured data display, and improves the user's reading experience.
Smart Images

Figure CN120179762A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of intelligent question answering, and particularly relates to a question answering method, device, system and vehicle. Background Art
[0002] With the development of artificial intelligence, the application of question answering systems is becoming more and more widespread. Generally, a question answering system can respond to a user instruction, output and display response data corresponding to the user instruction.
[0003] Currently, a question answering system usually responds to a user instruction after receiving it. However, in actual situations, a question answering system may receive multiple user instructions at once, and the multiple user instructions may include both question answering instructions and control instructions. If a question answering system receives multiple user instructions at once and the multiple user instructions include both question answering instructions and control instructions, the question answering system may not be able to accurately identify the question answering instructions from the multiple user instructions, resulting in a low accuracy in responding to the user instructions. In addition, the currently displayed response data is usually a large block of plain text, and key information cannot be highlighted, resulting in a low reading experience for users of the response data. Summary of the Invention
[0004] Embodiments of this application provide a question answering method, device, equipment, storage medium and vehicle, which can both improve the accuracy of responding to interactive input data and enhance the reading experience of users for the response data.
[0005] In a first aspect, embodiments of this application provide a question answering method, which includes:
[0006] Receiving the interactive input data of the user;
[0007] Using an artificial intelligence large model to process the interactive input data to obtain a generated content information set, the generated content information set including user intention description information, response data and interaction description information;
[0008] Determining a user interaction scenario corresponding to the user intention description information;
[0009] In the case where the user interaction scenario is a question answering scenario, determining a target service corresponding to the question answering scenario among multiple services, the service corresponding to the user interaction scenario and being used to perform subsequent actions corresponding to the interactive input data;
[0010] Through the target service, performing structured processing on the response data based on the interaction description information, and obtaining and displaying structured data.
[0011] In a possible implementation, determining the target service corresponding to the Q&A scenario among multiple services includes:
[0012] Obtain the registration intents respectively corresponding to the multiple services;
[0013] Determine the target service corresponding to the Q&A scenario according to the correspondence between the registration intent and the user interaction scenario.
[0014] In a possible implementation, structuring the response data based on the interaction description information through the target service, and obtaining and presenting structured data includes:
[0015] Perform scenario registration for the Q&A scenario through the target service to obtain a scenario identifier;
[0016] Route the response data corresponding to the scenario identifier to the target service;
[0017] Through the target service, structure the response data based on the interaction description information, and obtain and present structured data.
[0018] In a possible implementation, the response data includes multiple sub-response data output in a streaming manner. Structuring the response data based on the interaction description information, and obtaining and presenting structured data includes:
[0019] Determine the first sub-response data among the multiple sub-response data, where the first sub-response data is the response data for structuring;
[0020] Use an artificial intelligence large model to obtain the interaction description information corresponding to the first sub-response data;
[0021] Structure the first sub-response data based on the interaction description information, and obtain and present structured data.
[0022] In a possible implementation, the first sub-response data includes the complete target response data corresponding to the interaction input data, and the interaction description information includes the first interaction description information corresponding to the target response data. Structuring the first sub-response data based on the interaction description information, and obtaining and presenting structured data includes:
[0023] Perform marking processing on the first interaction description information in the target response data to obtain TTS structured data;
[0024] Display the TTS structured data in the TTS display area.
[0025] In a possible implementation, the first sub-response data includes a target paragraph, the interaction description information includes second interaction description information corresponding to the target paragraph, the second interaction description information includes picture information, and the step of performing structured processing on the first sub-response data based on the interaction description information to obtain and display structured data further includes:
[0026] Render the second interaction description information to generate card information;
[0027] Display the card message in the card message display area.
[0028] In a possible implementation, there are multiple target paragraphs, and the step of displaying the card information in the card information display area includes:
[0029] Update the displayed card information based on the multiple target paragraphs according to the generation order of the multiple target paragraphs;
[0030] Display the continuously updated card information in the card information display area until an end identifier is included in the target paragraphs to obtain the complete card information.
[0031] In a possible implementation, among the multiple sub-response data output in a streaming manner, the latter sub-response data includes the former sub-response data. The step of determining the first sub-response data among the multiple sub-response data includes:
[0032] For each sub-response data, determine whether a preset delimiter is included in the sub-response data;
[0033] In the case where one preset delimiter is included in the sub-response data, determine the sub-response data as the target paragraph;
[0034] In the case where multiple preset delimiters are included in the sub-response data, determine the sub-response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
[0035] In a possible implementation, the response data includes multiple sub-response data output in a streaming manner, where the latter sub-response data includes the former sub-response data. After determining the target service corresponding to the Q&A scenario among multiple services, the method further includes:
[0036] Perform a dialogue interaction on the first output sub-response data through the target service, where the dialogue interaction includes voice interaction and data display associated with the voice interaction;
[0037] Starting from the sub-response data of the second output, determine the sub-response data in the sub-response data except the previous sub-response data as the second sub-response data;
[0038] Perform a dialogue interaction on the second sub-response data until the sub-response data is the last one among the multiple sub-response data.
[0039] In a possible implementation manner, the interaction input data includes multimodal feature data, and the multimodal feature data characterizes the input form of the interaction input data. The input form of the interaction input data includes at least one of voice input, text input, touch input, and gesture input.
[0040] In a second aspect, an embodiment of the present application provides a question-and-answer device, and the device includes:
[0041] A receiving module, configured to receive the interaction input data of the user;
[0042] A first processing module, configured to process the interaction input data by using an artificial intelligence large model to obtain a generated content information set, and the generated content information set includes user intention description information, response data, and interaction description information;
[0043] A first determination module, configured to determine a user interaction scenario corresponding to the user intention description information;
[0044] A second determination module, configured to, when the user interaction scenario is a question-and-answer scenario, determine a target service corresponding to the question-and-answer scenario among multiple services, where the service corresponds to the user interaction scenario and is used to perform subsequent actions corresponding to the interaction input data;
[0045] A second processing module, configured to, through the target service, perform structured processing on the response data based on the interaction description information to obtain and display structured data.
[0046] In a third aspect, an embodiment of the present application provides a question-and-answer system, and the system includes:
[0047] A main interaction module, configured to receive the interaction input data of the user;
[0048] A large model module, configured to process the interaction input data by using an artificial intelligence large model to obtain a generated content information set, and the generated content information set includes user intention description information, response data, and interaction description information;
[0049] The dialogue management module is used to determine the user interaction scenario corresponding to the user intention description information, and send the user interaction scenario to the service management module. It is also used to determine the scenario identifier corresponding to the response data, and send the scenario identifier and its corresponding response data to the service management module;
[0050] The service management module is used to determine the target service assistant corresponding to the user interaction scenario sent by the dialogue management module according to the correspondence between the service assistant and the registered intention, and the correspondence between the registered intention and the user interaction scenario. It is also used to route the response data corresponding to the scenario identifier to the target service assistant according to the correspondence between the scenario identifier and the target service assistant;
[0051] The service assistant is used to register the user interaction scenarios it is concerned about with the service management module to obtain the registered intention. It is also used to perform scenario registration on the user interaction scenario to obtain a scenario identifier, and process the response data corresponding to the scenario identifier. The service assistant includes the target service assistant.
[0052] In a possible implementation manner, the service assistant includes a task-based service assistant and an AI-based service assistant;
[0053] The task-based service assistant is used to process tasks corresponding to the interaction control scenario;
[0054] The AI-based service assistant is used to process tasks corresponding to the Q&A scenario.
[0055] In a possible implementation manner, the response data includes multiple sub-response data output in a streaming manner. The AI-based service assistant includes a TTS playback engine, a TTS structuring engine, a rich media structuring engine, and a dynamic parsing engine;
[0056] The TTS playback engine is used to perform dialogue interaction on the multiple sub-response data output in a streaming manner;
[0057] The TTS structuring engine is used to perform TTS structuring processing on the complete target response data corresponding to the interaction input data to obtain TTS structured data;
[0058] The rich media structuring engine is used to perform rich media structuring processing on the target paragraph corresponding to the interaction input data to obtain card information;
[0059] The dynamic parsing engine is used to dynamically parse the multiple sub-response data, the target response data, and the multiple target paragraphs during the process of streaming output of the multiple sub-response data, conducting the conversation interaction, performing the TTS structured processing, and performing the rich media structured processing.
[0060] In a possible implementation manner, the system further includes a generative user interface service;
[0061] The generative user interface service is used to dynamically display the TTS structured data and the card information.
[0062] In a fourth aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;
[0063] When the processor executes the computer program instructions, the method in any possible implementation method in the first aspect above is implemented.
[0064] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method in any possible implementation method in the first aspect above is implemented.
[0065] In a sixth aspect, an embodiment of the present application provides a vehicle, which includes at least one of the following:
[0066] A question and answer device as in any embodiment of the second aspect;
[0067] A question and answer system as in any embodiment of the third aspect;
[0068] An electronic device as in any embodiment of the fourth aspect;
[0069] A computer-readable storage medium as in any embodiment of the fifth aspect.
[0070] In the question-and-answer method, apparatus, system, device, storage medium, and vehicle according to the embodiments of the present application, the artificial intelligence large model can process the interactive input data to obtain a generated content information set including user intention description information, response data, and interactive description information. Based on this, since the service corresponds to the user interaction scenario, by determining the user interaction scenario corresponding to the user intention description information, and then performing subsequent actions corresponding to the interactive input data through the service corresponding to the user interaction scenario, the professionalism of performing subsequent actions can be ensured, and the accuracy of performing subsequent actions can be improved. When the user interaction scenario is a question-and-answer scenario, by determining the target service corresponding to the question-and-answer scenario among multiple services, and performing subsequent actions corresponding to the interactive input data through the target service, the accuracy of answering the interactive input data can be improved. In addition, through the target service, the response data is structured based on the interactive description information, and the structured data is obtained and displayed. Compared with displaying a large paragraph of plain text, the key information can be highlighted, and the reading experience of the user for the response data can be improved. Thus, through the embodiments of the present application, the accuracy of answering the interactive input data can be improved, and the reading experience of the user for the response data can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0072] Figure 1 is a schematic structural diagram of a question-and-answer system provided by an embodiment of the present application;
[0073] Figure 2 is a schematic flowchart of a question-and-answer method provided by an embodiment of the present application;
[0074] Figure 3 is a schematic comparison diagram between plain text and TTS structured data provided by an embodiment of the present application;
[0075] Figure 4 is a schematic flowchart of determining a target paragraph through a dynamic parsing engine provided by an embodiment of the present application;
[0076] Figure 5 is a schematic diagram of card information provided by an embodiment of the present application;
[0077] Figure 6 is a schematic diagram of a voice interaction interface including a TTS display area and a card information display area provided by an embodiment of the present application;
[0078] Figure 7It is a schematic diagram of the correspondence between a template type and a card type provided by an embodiment of the present application;
[0079] Figure 8 It is a schematic diagram of the process of performing dialogue interaction on sub-response data by a TTS structured engine provided by an embodiment of the present application;
[0080] Figure 9 It is an interaction schematic diagram of a question-and-answer method provided by an embodiment of the present application;
[0081] Figure 10 It is a schematic diagram of the structure of a question-and-answer device provided by an embodiment of the present application;
[0082] Figure 11 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0083] In order to more clearly understand the above objects, features and advantages of the present application, the solution of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.
[0084] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application, but the present application may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all the embodiments.
[0085] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0086] As described in the background art section, to solve the problems of the prior art, embodiments of the present application provide a question-and-answer method, apparatus, system, device, storage medium, and vehicle. Among them, the question-and-answer system can be set in a terminal server, can be set in a cloud server, or can be set in a distributed server, which is not limited herein. That is to say, the execution subject of the question-and-answer method can be a server. The server can be a terminal server, can be a cloud server, or can be a distributed server, which is not limited herein. Among them, the terminal server can include, for example, servers corresponding to mobile devices such as mobile phones, computers, and tablets, servers corresponding to smart wearable devices such as smart glasses and smart watches, and vehicle servers. In addition, the question-and-answer method can be specifically applied to the voice question-and-answer scenario in the in-vehicle environment. That is to say, the terminal server can be a vehicle server. In addition, the execution subject of the question-and-answer method can also include a terminal server and a cloud server at the same time. That is to say, the question-and-answer system can also set some of its functional modules in the terminal server and the other part in the cloud server. The terminal server and the cloud server can communicate through a network.
[0087] The question-and-answer system provided by the embodiments of the present application will be introduced below.
[0088] Figure 1 Fig. shows a schematic structural diagram of a question-and-answer system 100 provided by an embodiment of the present application. As Figure 1 shown, the question-and-answer system 100 provided by the embodiment of the present application can include a main interaction module 11, a dialogue management module 12, a large model module 13, a service management module 14, a service assistant 15, and a generative user interface service 16. Among them, the service assistant 15 can include a task-based service assistant 151 and an AI-based service assistant 152. The AI-based service assistant can be the abbreviation of the Artificial Intelligence-based service assistant. The AI-based service assistant 152 can include a text-to-speech (TTS) broadcast engine 1521, a TTS structured engine 1522, a rich media structured engine 1523, and a dynamic parsing engine 1524.
[0089] First, the large model module 13 may include an artificial intelligence large model. An artificial intelligence (AI) large model refers to a high-performance artificial intelligence model constructed through a large number of training samples and computing resources. They can learn a large amount of language knowledge, image features, and speech patterns, and can reason and generate outputs similar to humans, with wide applications in the fields of natural language processing, image recognition, speech recognition, etc. The artificial intelligence large model may include, for example, a large language model (LLM), ChatGPT (Chat Generative Pre-trained Transformer), a multi-modal large model, and a multi-modal cognitive large model, etc. Specifically, the artificial intelligence large model can process the interactive input data to obtain a set of generated content information. The set of generated content information may include user intention description information, response data, and interactive description information. The response data may include multiple sub-response data output in a streaming manner, where the latter sub-response data may include the former sub-response data. The interactive description information can be used to perform structured processing on the response data.
[0090] In addition, the large model module 13 may also include TaskFormer. TaskFormer can be regarded as a model that improves and expands on the basis of GPT (Generative Pre-trained Transformer) for task-based natural language processing problems. It can better handle diverse tasks and has high flexibility and versatility. In the embodiments of this application, TaskFormer can be equivalent to the central control of the artificial intelligence large model and can perform pre-processing and post-processing on the input and output content of the artificial intelligence large model.
[0091] In addition, the large model module 13 can be set in a terminal server, can be set in a cloud server, or can be set in a distributed server, which is not limited here.
[0092] The main interaction module 11 can be used to receive the interactive input data of the user, determine the user intention description information corresponding to the interactive input data, and send the user intention description information to the dialogue management module 12. Among them, the interactive input data can be multimodal instruction interactive data input by the user. That is to say, the interactive input data can include multimodal feature data and instruction interactive data. Among them, the multimodal feature data can characterize the input form of the interactive input data. The input form of the interactive input data can include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interactive data can characterize the user intention corresponding to the interactive input data. The instruction interactive data can include, for example, control instructions related to vehicle control and vehicle settings, question instructions related to vehicle Q&A and general Q&A, and can also include both control instructions and question instructions, which are not limited here.
[0093] As an example, the main interaction module 11 can have the function of Automatic Speech Recognition (ASR). If the input form of the interactive input data is voice input, the main interaction module can convert the voice instruction into an instruction text through the ASR function and send the instruction text to the large model module 13. After receiving the instruction text, the large model module 13 can perform semantic parsing on the instruction text through the artificial intelligence large model to obtain the user intention description information and return the user intention description information to the main interaction module 11.
[0094] In addition, the question and answer system 100 can also include a voice user interface (VUI) corresponding to the main interaction module 11. The voice user interface can include a text input box. The user can enter a text instruction in the text input box. The main interaction module 11 can determine the user intention description information corresponding to the interactive input data based on the correspondence between the text instruction and the user intention.
[0095] In addition, the voice user interface can also include multiple interactive controls, and the interactive controls can include preset instructions related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. The user can perform a touch instruction by clicking on the interactive control. The main interaction module 11 can determine the user intention description information corresponding to the interactive input data based on the correspondence between the touch instruction and the user intention.
[0096] In addition, if the voice user interface supports gesture air touch, the user can use gesture air control to select any preset instruction for the interactive control. The main interaction module 11 can determine the user intention description information corresponding to the interactive input data based on the correspondence between the gesture instruction and the user intention.
[0097] In addition, the main interaction module 11 is also responsible for building the basic voice capabilities, including the software development kit (SDK) access for dialogue types, dialogue execution, interaction modes, user modes, large models, and free conversations, as well as the maintenance of voice images. Among them, the SDK can be used as a general capability of the in-vehicle intelligent cockpit and is widely applied to scenarios such as large models, input methods, voice assistants, and gesture recognition.
[0098] The dialogue management module 12 can receive the user intention description information sent by the main interaction module 11 through the SDK. After receiving the user intention description information, the dialogue management module 12 can determine the user interaction scenario corresponding to the user intention description information according to the arbitration strategy and send the user interaction scenario to the service management module 14. Among them, the user interaction scenario can include a question-and-answer scenario and an interaction control scenario. Among them, the arbitration strategy can be the pre-defined corresponding relationship between the user intention description information and the user interaction scenario.
[0099] In the service management module 14, the corresponding relationship between multiple service assistants 15 and registered intentions can be pre-stored. Among them, the registered intentions and the user interaction scenarios can be in one-to-one correspondence. The registered intention can represent the user interaction scenario that the service assistant can handle. After receiving the user interaction scenario, the service management module 14 can first determine the target registered intention corresponding to the received user interaction scenario according to the corresponding relationship between the user interaction scenario and the registered intention, and then determine the service assistant corresponding to the target registered intention according to the corresponding relationship between the registered intention and the service assistant.
[0100] It should be noted that in order for the service management module 14 to determine the corresponding relationship between the service assistant and the registered intention, each service assistant 15 can pre-register the user interaction scenarios it is concerned about with the service management module 14 to obtain the registered intention.
[0101] In addition, after determining the target service assistant, the service management module 14 can notify the target service assistant to register its corresponding user interaction scenario to obtain a scenario identifier. The service management module 14 can record the correspondence between the target service assistant and the scenario identifier, and send the scenario identifier to the dialogue management module 12. After receiving the scenario identifier, the dialogue management module 12 can initiate a streaming data request corresponding to the scenario identifier to the large model module 13. The streaming data request can be used to request the artificial intelligence large model to stream out multiple sub-response data. After receiving the streaming data request, the large model module 13 can stream out multiple sub-response data (i.e., response data) to the dialogue management module 12 through the main interaction module 11. After receiving the response data, the dialogue management module 12 can determine the scenario identifier corresponding to the response data, and send the scenario identifier and its corresponding response data to the service management module 14. After receiving the scenario identifier and its corresponding response data, the service management module 14 can route the response data corresponding to the scenario identifier to the target service assistant. After receiving the response data, the target service assistant can process the response data (the response data corresponding to the scenario identifier registered by the target service assistant).
[0102] In addition, the dialogue management module 12 is also responsible for skill operation, service operation, exception handling operation, and exception handling, including skill whitelist configuration, script configuration, expression configuration, exception monitoring, etc.
[0103] Specifically, the service assistant 15 can include a task-based service assistant 151 and an AI-based service assistant 152. Among them, the task-based service assistant 151 can be used to process tasks corresponding to the interaction control scenario. The specific process of the task-based service assistant 151 processing tasks can include: determining the target application corresponding to the user instruction, and sending the user instruction to the target application so that the target application executes the user instruction. In addition, the AI-based service assistant 152 can be used to process tasks corresponding to the Q&A scenario. The AI-based service assistant can be an auxiliary tool integrated in the voice system, which can help users complete various Q&As. The specific process of the AI-based service assistant 152 processing tasks can include: sending a structured request to the large model module 13. Among them, the structured request can be used to request the artificial intelligence large model to obtain the interaction description information corresponding to the response data. After receiving the interaction description information, the AI-based service assistant 152 can perform structured processing on the response data based on the interaction description information to obtain and display the structured data.
[0104] The AI type service assistant 152 may specifically include a TTS broadcast engine 1521, a TTS structuring engine 1522, a rich media structuring engine 1523, and a dynamic parsing engine 1524. Among them, the TTS broadcast engine 1521 can be used for dialogue interaction with multiple sub-response data of streaming output. The TTS broadcast engine 1521 can generate voice frame-by-frame according to the text based on the TTS interruption broadcast strategy and the append broadcast strategy, making the speech synthesis a streaming process, capable of real-time output of sound and playing simultaneously with the generation of speech. This way can make the synthesized speech more fluent and reduce latency. The TTS structuring engine 1522 can be used for TTS structuring processing of the complete target response data corresponding to the interaction input data to obtain TTS structured data. The rich media structuring engine 1523 can be used for rich media structuring processing of the target paragraph corresponding to the interaction input data to obtain card information. The dynamic parsing engine 1524 can be used for dynamically parsing multiple sub-response data, target response data, and multiple target paragraphs during the process of streaming output of multiple sub-response data, dialogue interaction, TTS structuring processing, and rich media structuring processing.
[0105] The generative user interface service 16 can dynamically display the generated TTS structured data and card information. The generative user interface service 16 can include a card service, a page service, a widget, and a webview container. Among them, the webview container can be a commonly used component in mobile application development and can be used to display web (Web) content. The webview container provides a browser view that can be embedded in the application, enabling developers to load and display Web pages, HTML files, JavaScript applications, etc. through the webview. In addition, the generative user interface (User Interface, UI) technology can be composed of the concept combination of a container + a template. After the large model outputs the display template corresponding to the response data, interface rendering can be achieved by loading the container. The generative UI technology can be specifically implemented based on fusion technologies such as Native, ReactNative, HTML5 (H5), etc. Among them, H5 is a standard technology for building and presenting Web content, which consists of HTML (HyperText Markup Language), CSS (Cascading Style Sheets), and JavaScript, etc. The emergence and development of H5 technology have brought a richer and more interactive Internet experience. The advantage of generative UI lies in flexibility and scalability. By separating the UI from the data, different UIs can be dynamically generated according to the changes in the data at runtime, which is very suitable for processing dynamic content and variable structure interfaces.
[0106] In addition, the question-and-answer system may further include a control in a Graphical User Interface (GUI). The GUI control may be responsible for overall GUI scheduling, including card management, card creation and destruction, data update of widgets, creation and destruction of webview containers, etc.
[0107] Based on this, in order to improve the response efficiency of the question-and-answer system 100. In some embodiments, the main interaction module 11, the dialogue management module 12, the service management module 14, the service assistant 15, and the generative user interface service 16 may be set in a terminal server. The terminal server may include, for example, a server corresponding to mobile devices such as mobile phones, computers, and tablets, a server corresponding to smart wearable devices such as smart glasses and smart watches, and a vehicle server, etc. The large model module 13 may be set in a cloud server. Additionally, the large model module 13 may also be deployed in a cloud server, and an offline large model module may be set in the terminal server.
[0108] If the large model module 13 is set in a cloud server, the question-and-answer system 100 may further include a cloud control / dialogue. The cloud control / dialogue may be responsible for dialogue management in the cloud. The cloud control / dialogue may be a cloud computing-based service that can receive and process instructions input by users through voice, and implement voice control and dialogue functions.
[0109] Specifically, the cloud control may refer to sending instructions input by users through voice to a remote cloud server for processing and parsing. The cloud server may use speech recognition technology to convert speech into text, and then through technologies such as semantic parsing and intent recognition, understand the instructions and intentions of users. Voice dialogue may refer to the function of communicating and having a dialogue through voice. The cloud control may, through natural language processing technology, convert the voice instructions of users into a dialogue, providing an interactive experience similar to human-machine dialogue. By converting voice into text, and using technologies such as semantic understanding and language generation, the cloud control can understand the questions or needs of users and provide corresponding answers or responses.
[0110] The advantage of the cloud control / dialogue is that it can utilize the powerful computing power of cloud computing and rich speech processing technologies to provide a high-quality and intelligent voice interaction experience. Users can control devices, obtain information through voice input, and can also communicate and interact with the intelligent system through voice dialogue.
[0111] If the large model module 13 is set in a cloud server, the main interaction module 11 may also send the instruction data of users to the cloud control through the SDK.
[0112] Next, the question-and-answer method provided by the embodiments of the present application will be introduced.
[0113] Figure 2 shows a schematic flowchart of a question-and-answer method provided by an embodiment of the present application. The question-and-answer method can be executed by a processor in a vehicle. As Figure 2 shown, the question-and-answer method provided by an embodiment of the present application includes the following steps:
[0114] S210. Receive the interactive input data of the user;
[0115] S220. Use an artificial intelligence large model to process the interactive input data to obtain a generated content information set, and the generated content information set includes user intention description information, response data, and interactive description information;
[0116] S230. Determine the user interaction scenario corresponding to the user intention description information;
[0117] S240. When the user interaction scenario is a question-and-answer scenario, determine a target service corresponding to the question-and-answer scenario among multiple services. The service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interactive input data;
[0118] S250. Through the target service, perform structured processing on the response data based on the interactive description information to obtain and display structured data.
[0119] In the question-and-answer method of the embodiment of the present application, the artificial intelligence large model can process the interactive input data to obtain a generated content information set including user intention description information, response data, and interactive description information. Based on this, since the service corresponds to the user interaction scenario, therefore, by determining the user interaction scenario corresponding to the user intention description information, and then executing subsequent actions corresponding to the interactive input data through the service corresponding to the user interaction scenario, the professionalism of executing subsequent actions can be ensured, and further the accuracy of executing subsequent actions can be improved. By determining a target service corresponding to the question-and-answer scenario among multiple services when the user interaction scenario is a question-and-answer scenario, and executing subsequent actions corresponding to the interactive input data through the target service, the accuracy of answering the interactive input data can be improved. In addition, through the target service, performing structured processing on the response data based on the interactive description information to obtain and display structured data, compared with displaying a large paragraph of plain text, can highlight key information, and further can enhance the user's reading experience of the response data. In this way, through the embodiment of the present application, the accuracy of answering the interactive input data can be improved, and the user's reading experience of the response data can be enhanced.
[0120] The following introduces the specific implementation manners of the above steps.
[0121] In some embodiments, in S210, the interactive input data may be multimodal instruction interactive data input by the user. That is, the interactive input data may include multimodal feature data and instruction interactive data. Among them, the multimodal feature data may characterize the input form of the interactive input data. The input form of the interactive input data may include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interactive data may characterize the user intention corresponding to the interactive input data. The instruction interactive data may include, for example, control instructions related to vehicle control and vehicle settings, may include question instructions related to vehicle Q&A and general Q&A, and may also include both control instructions and question instructions at the same time, which is not limited here.
[0122] As an example, the Q&A system may include a main interaction module and a large model module. Among them, the main interaction module may have the function of Automatic Speech Recognition (ASR). If the input form of the interactive input data is voice input, the main interaction module may convert the voice instruction into an instruction text through the ASR function and send the instruction text to the large model module. After receiving the instruction text, the large model module may input the interactive input data into the artificial intelligence large model, perform semantic parsing on the instruction text through the artificial intelligence large model, obtain the user intention description information (i.e., the instruction interactive data), and return the user intention description information to the main interaction module.
[0123] In addition, the Q&A system may also include a voice user interface (VUI) corresponding to the main interaction module. The voice user interface may include a text input box. The user can input a text instruction in the text input box. The main interaction module may determine the instruction interactive data corresponding to the interactive input data based on the correspondence between the text instruction and the instruction interactive data.
[0124] In addition, the voice user interface may also include multiple interactive controls, and the interactive controls may include preset instructions related to vehicle control, vehicle settings, vehicle Q&A, general Q&A, etc. The user can perform touch instructions by clicking on the interactive controls. The main interaction module may determine the instruction interactive data corresponding to the interactive input data based on the correspondence between the touch instruction and the instruction interactive data.
[0125] In addition, if the voice user interface supports gesture air touch, the user can use gesture air control to select any preset instruction for the interactive control. The main interaction module may determine the instruction interactive data corresponding to the interactive input data based on the correspondence between the gesture instruction and the instruction interactive data.
[0126] In some embodiments, in S220, if the input form of the interactive input data is voice input, after receiving the user's interactive input data, the main interaction module may send the interactive input data to the large model module. The large model module may input the interactive input data into the artificial intelligence large model, and the artificial intelligence large model may identify and respond to the interactive input data. Among them, identifying the interactive input data may generate user intention description information. Responding to the interactive input data may generate response data corresponding to the interactive input data. If the input form of the interactive input data is non-voice input (including text input, touch input, and gesture input), after receiving the user's interactive input data, the main interaction module may determine the instruction interaction data corresponding to the interactive input data by itself and send the instruction interaction data to the large model module. The large model module may input the instruction interaction data into the artificial intelligence large model, and the artificial intelligence large model may identify and respond to the instruction interaction data. Among them, identifying the instruction interaction data may generate user intention description information. Responding to the instruction interaction data may generate response data corresponding to the instruction interaction data.
[0127] As an example, if the large model module is located in the cloud server and the main interaction module and the dialogue management module are located in the terminal server, after the terminal server sends the interactive input data to the cloud server through the main interaction module, it may continue to send an asynchronous generation interface to the cloud server, so that the large model module returns the user intention description information to the terminal server. After receiving the user intention description information, the main interaction module in the terminal server may send the user intention description information to the dialogue management module.
[0128] In addition, after the large model module returns the user intention description information to the terminal server, it may continue to obtain the response data corresponding to the interactive input data (including the instruction interaction data) and cache the obtained response data. Among them, obtaining the response data corresponding to the interactive input data includes searching for the response data corresponding to the interactive input data in the pre-set Q&A library, searching for the response data corresponding to the interactive input data in the network, and generating the response data corresponding to the interactive input data. In the process of the artificial intelligence large model generating the response data corresponding to the interactive input data, when predicting the next word or character, the content already generated before will be considered. This context-aware generation method enables the artificial intelligence large model to generate text with coherence and certain logic.
[0129] In some embodiments, in S230, after receiving the user intention description information, the dialogue management module may determine the user interaction scenario corresponding to the user intention description information according to the arbitration policy. Among them, the user interaction scenario may include a question-and-answer scenario and an interaction control scenario. Specifically, the question-and-answer scenario may include a general question-and-answer scenario and a vehicle question-and-answer scenario. The interaction control scenario may include a vehicle control scenario. In addition, the arbitration policy may be a pre-set correspondence between the user intention description information and the user interaction scenario.
[0130] As an example, when the dialogue management module determines that the user interaction scenario is a question-and-answer scenario, it may send a streaming data request to the large model module. Among them, the streaming data request may be used to request the large model module to stream out multiple sub-response data corresponding to the interaction input data. That is to say, the response data may include multiple sub-response data output in a stream. After receiving the streaming data request, the large model module may, in response to the streaming data request, stream out multiple sub-response data. Among them, the latter sub-response data may include the previous sub-response data. For example, if the interaction input data is "Introduce Zhou xx", the multiple sub-response data output in a stream may include: the first sub-response data is "Zhou xx,", the second sub-response data is "Zhou xx, from xxx,", and the third sub-response data is "Zhou xx, from xxx, won the x award on x month x day..."
[0131] As a more specific example, if the large model module is located in the cloud server and the main interaction module and the dialogue management module are located in the terminal server, the dialogue management module may send a streaming data request to the large model module through the main interaction module.
[0132] As another example, after determining the user interaction scenario, the dialogue management module may also send the user interaction scenario to the service management module, so that the service management module selects the service corresponding to the user interaction scenario from multiple services to perform subsequent actions corresponding to the response data.
[0133] In some embodiments, in S240, the service may correspond to the user interaction scenario and be used to perform subsequent actions corresponding to the interaction input data. For example, if the user interaction scenario is a question-and-answer scenario, the service corresponding to the question-and-answer scenario may be used to perform dialogue interaction on the multiple sub-response data output in a stream and to perform structured data on the response data. Among them, the dialogue interaction may include voice interaction and data display corresponding to the voice interaction. Another example, if the user interaction scenario is an interaction control scenario, the service corresponding to the interaction control scenario may be used to send the user intention description information to the corresponding application, so that the corresponding application executes the control instruction corresponding to the user intention description information. In addition, the target service may be the service corresponding to the question-and-answer scenario.
[0134] As an example, different services can be executed by different service assistants. That is, there can be a one-to-one correspondence between service assistants and services. For example, the target service corresponding to the Q&A scenario can be executed by the AI-type service assistant corresponding to the Q&A scenario.
[0135] Based on this, in some embodiments, the above S240 may specifically include:
[0136] Obtain registration intents respectively corresponding to multiple services;
[0137] Determine the target service corresponding to the Q&A scenario according to the correspondence between the registration intent and the user interaction scenario.
[0138] Here, the service management module may pre-store the correspondence between multiple service assistants (i.e., services) and registration intents. Among them, there can be a one-to-one correspondence between the registration intent and the user interaction scenario. The registration intent can represent the user interaction scenario that the service assistant can handle. After receiving the user interaction scenario (which is the Q&A scenario), the service management module can first determine the target registration intent corresponding to the Q&A scenario according to the correspondence between the user interaction scenario and the registration intent, and then determine the service assistant (i.e., the AI-type service assistant) corresponding to the target registration intent according to the correspondence between the registration intent and the service assistant. After determining the AI-type service assistant, the target service corresponding to the AI-type service assistant can be determined.
[0139] In some embodiments, in S250, after determining the target service (i.e., the AI-type service assistant), a structured request can be sent to the large model module through the AI-type service assistant. The structured request can be used to request the artificial intelligence large model to obtain the interaction description information corresponding to the response data. After receiving the structured request, the large model module can input the response data into the artificial intelligence large model, and generate the interaction description information corresponding to the response data through the artificial intelligence large model. Then, the large model module can return the interaction description information to the AI-type service assistant. After receiving the interaction description information, the AI-type service assistant can, through the target service, perform structured processing on the response data based on the interaction description information, and obtain and display the structured data.
[0140] Based on this, in order to implement structured processing of the response data based on the interaction description information through the target service, in some embodiments, the above S250 may specifically include:
[0141] Perform scenario registration on the Q&A scenario through the target service to obtain a scenario identifier;
[0142] Route the response data corresponding to the scenario identifier to the target service;
[0143] Based on the target service, the response data is structurally processed according to the interaction description information, and the structured data is obtained and displayed.
[0144] Here, after determining the target service (i.e., the AI-based service assistant), the service management module can notify the target service to register the Q&A scenario to obtain a scenario identifier. The service management module can record the correspondence between the target service and the scenario identifier, and send the scenario identifier to the dialogue management module. After receiving the scenario identifier, the dialogue management module can initiate a streaming data request corresponding to the scenario identifier to the large model module. After receiving the streaming data request, the large model module can stream out multiple sub-response data (i.e., response data) to the dialogue management module through the main interaction module. After receiving the response data, the dialogue management module can determine the scenario identifier corresponding to the response data, and send the scenario identifier and its corresponding response data to the service management module. After receiving the scenario identifier and its corresponding response data, the service management module can route the response data corresponding to the scenario identifier to the target service. After receiving the response data, the target service can structurally process the response data based on the interaction description information, and obtain and display the structured data.
[0145] In this way, by routing the response data corresponding to the scenario identifier to the target service, it is possible to structurally process the response data based on the interaction description information through the target service.
[0146] Based on this, in order to improve the display efficiency of the structured data and enhance the user experience, in some embodiments, the above-mentioned structurally processing the response data according to the interaction description information to obtain and display the structured data may specifically include:
[0147] Determine the first sub-response data among the multiple sub-response data, where the first sub-response data is the response data for structural processing;
[0148] Input the first sub-response data into the artificial intelligence large model to obtain the interaction description information corresponding to the first sub-response data through the artificial intelligence large model;
[0149] Structurally process the first sub-response data based on the interaction description information, and obtain and display the structured data.
[0150] Here, the structured processing may include TTS structured processing and rich media structured processing. If the first sub-response data is used for TTS structured processing, the first sub-response data may be the complete target response data corresponding to the interactive input data. The interactive description information corresponding to the target response data may be the first interactive description information. Among them, the first interactive description information may include a title, a subtitle, preset keywords, etc. If the first sub-response data is used for rich media structured processing, the first sub-response data may be the target paragraph corresponding to the interactive input data. The interactive description information corresponding to the target paragraph may be the second interactive description information. Among them, there may be multiple target paragraphs. If the complete target response data includes three major paragraphs in total, the target paragraph may be any one of the three major paragraphs. In addition, the second interactive description information may include a title, a subtitle, an abstract, preset keywords, etc. In addition, the second interactive description information may further include picture information. In addition, if the first sub-response data is the target paragraph related to comparison, the second interactive description information may further include comparison objects, comparison items, comparison contents, etc.
[0151] Based on this, in some embodiments, determining the first sub-response data among multiple sub-response data may specifically include:
[0152] When the sub-response data includes an end flag, determining the sub-response data as the complete target response data.
[0153] Here, the sub-response data may include an end flag. If the sub-response data includes an end flag, the sub-response data may be the last sub-response data among the multiple sub-response data. Since the subsequent sub-response data includes the previous sub-response data, when the sub-response data includes an end flag, the sub-response data may be the complete target response data corresponding to the interactive input data.
[0154] Based on this, in order to improve the user's reading experience of the target response data, in some embodiments, the above-mentioned structured processing of the first sub-response data based on the interactive description information to obtain and display the structured data may specifically include:
[0155] Performing a marking process on the first interactive description information in the target response data to obtain TTS structured data;
[0156] Displaying the TTS structured data in the TTS display area.
[0157] Here, the markup processing may include highlighting, bolding, italicizing, underlining, increasing the font size, etc. Additionally, the TTS structured data may include interactive input data and response data after performing TTS structuring processing on the target response data of the interactive input data. Additionally, the TTS display area may be a preset display area of the voice interaction interface. The TTS display area may be used to display the TTS structured data.
[0158] As an example, the TTS structuring engine may perform markup processing on the first interactive description information based on the Markdown syntax to convert the target response data into TTS structured data. After obtaining the TTS structured data, the TTS structured data can be displayed in the TTS display area.
[0159] As an example, a comparison schematic diagram between plain text and TTS structured data provided by the embodiments of the present application may be as Figure 3 shown.
[0160] In this way, by displaying the TTS structured data instead of a large block of plain text, the effect of highlighting key information can be achieved, improving the user's reading experience of the target response data.
[0161] Additionally, in some embodiments, the above-mentioned structuring processing of the first sub-response data based on the interactive description information, obtaining and displaying the structured data may specifically further include:
[0162] For each sub-response data, determine whether the sub-response data includes a preset delimiter;
[0163] In the case where the sub-response data includes a preset delimiter, determine the sub-response data as the target paragraph;
[0164] In the case where the sub-response data includes multiple preset delimiters, determine the sub-response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
[0165] Here, the delimiter interceptor may be included in the dynamic parsing engine. The delimiter interceptor may determine whether the sub-response data includes a preset delimiter. If the sub-response data does not include a preset delimiter, it may continue to determine whether the next sub-response data of the sub-response data includes a preset delimiter. If the sub-response data includes a preset delimiter, the target paragraph for rich media structuring may be determined based on the preset delimiter.
[0166] Specifically, the preset delimiter can be denoted as <|br / >, for example. If the sub-response data is "AAA. BBB. <|br / >", then there is one pre-delimiter in the response data, and the target paragraph can be "AAA. BBB.". If the sub-response data is "AAA. BBB. <|br / >CCC. DDD. <|br / >", then there are two preset delimiters in the sub-response data, and the target paragraph can be "CCC. DDD.". If the sub-response data is "AAA. BBB. <|br / >CCC. DDD. <|br / >EEE. FFF. <|br / >", then there are three preset delimiters in the sub-response data, and the target paragraph can be "EEE. FFF.".
[0167] It should be noted that in the embodiments of the present application, the first process and the second process can be carried out synchronously without affecting each other. Among them, the first process can be a process of determining the target response data from multiple sub-response data and performing TTS structuring based on the target response data. The second process can be a process of determining the target paragraph from multiple sub-response data and performing rich media structuring based on the target paragraph.
[0168] Based on this, in order to reasonably arrange the sub-response data and ensure the necessity and effectiveness of the structuring process, in some embodiments, before determining the target paragraph from multiple sub-response data, it is also possible to determine whether it is necessary to determine the target paragraph through a dynamic parsing engine. Among them, the dynamic parsing engine can include an instruction interceptor, a template interceptor, a word count interceptor, a delimiter interceptor, and a structuring interceptor.
[0169] As an example, as Figure 4 shown, a flowchart of determining the target paragraph through a dynamic parsing engine provided by the embodiments of the present application may include the following steps:
[0170] S41. Determine whether the i-th sub-response data is in the instruction whitelist. If so, execute S42; if not, execute S45.
[0171] S42. Intercept the control instructions in the sub-response data through the instruction interceptor.
[0172] S43. Determine whether a control instruction is intercepted. If so, execute S44; if not, execute S45.
[0173] S44. Execute the control instruction.
[0174] S45. Determine whether the sub-response data is in the structuring whitelist. If so, execute S46; if not, set i = i + 1 and return to execute S41.
[0175] S46. Intercept and record key information such as template type, word count threshold, source, etc. in the sub-response data through a template interceptor;
[0176] S47. Obtain the word count in the sub-response data through a word count interceptor;
[0177] S48. Determine whether the word count is greater than the word count threshold through the word count interceptor. If so, execute S49; if not, set i = i + 1 and return to execute S41;
[0178] S49. Determine whether the sub-response data includes a preset delimiter through a delimiter interceptor. If so, execute S410; if not, set i = i + 1 and return to execute S41;
[0179] S410. Determine the target paragraph based on the preset delimiter through a structuring interceptor, and initiate a rich media structuring request corresponding to the target paragraph to the large model module. The rich media structuring request is used to request the large model module to obtain the second interaction description information corresponding to the target paragraph.
[0180] In Figure 4 , on the one hand, by determining the target paragraph only when the sub-response data meets the above multiple conditions, the sub-response data can be reasonably arranged to ensure the necessity and effectiveness of the structuring process. On the other hand, for each sub-response data, when the sub-response data is in the instruction whitelist, intercept and execute the control instructions in the sub-response data, which can ensure the timeliness of executing user instructions and improve the user experience.
[0181] In addition, in some embodiments, the above-mentioned inputting the first sub-response data into the artificial intelligence large model to obtain the interaction description information corresponding to the first sub-response data through the artificial intelligence large model may specifically include:
[0182] Input the target paragraph into the artificial intelligence large model, first obtain text interaction description information such as the title, subtitle, abstract, preset keywords, etc. corresponding to the target paragraph through the artificial intelligence large model, and then search for pictures corresponding to the text interaction description information. Among them, pictures may or may not be found, which is not limited here. In addition, the picture information may include, for example, a picture link.
[0183] Based on this, in order to further improve the user's visual experience, in some embodiments, the above-mentioned structuring process of the first sub-response data based on the interaction description information to obtain and display the structured data may specifically further include:
[0184] Render the second interaction description information to generate card information;
[0185] Display the card information in the card information display area.
[0186] Here, the second interaction description information may further include which paragraph the target paragraph is in the complete target response data, as well as a paragraph identifier. After the large artificial intelligence model returns the second interaction description information, the rich media structuring engine may render the second interaction description information to generate and display card information. Among them, the card information may be card information in rich media form. The card information in rich media form may include audio, text, links, interactive controls, pictures, etc. corresponding to the target paragraph. In addition, the card information may also include interactive input data and the response data after structured processing. For example, a schematic diagram of a kind of card information provided by an embodiment of the present application may be as Figure 5 shown.
[0187] As an example, after the large artificial intelligence model returns the second interaction description information, the rich media structuring engine may pull up a hypertext markup language interface (such as HTML 5, H5) interface, and transmit the second interaction description information to H5 in the way of JS bridge. The client may load H5 and the second interaction description information through a browser (such as a webview container) to generate and display card information.
[0188] In addition, the card information display area may be a preset display area of the voice interaction interface. The card information display area may be adjacent to the TTS display area and located on the right side of the TTS display area. Of course, the card information display area may also be located on the left side of the TTS display area, which is not limited here. A schematic diagram of a voice interaction interface including a TTS display area and a card information display area provided by an embodiment of the present application may be as Figure 6 shown. In Figure 6 it, the left side may be the TTS display area. The right side may be the card information display area.
[0189] In this way, the visual experience of the user can be further improved by displaying the card information.
[0190] In addition, as described above, there may be multiple target paragraphs. Based on this, in order to ensure the timeliness of card information display and further improve the visual experience of the user, in some embodiments, the above-mentioned displaying card information in the card information display area may specifically include:
[0191] Updating the displayed card information based on multiple target paragraphs according to the generation order of the multiple target paragraphs;
[0192] Displaying the continuously updated card information in the card information display area until the target paragraph includes an end identifier to obtain complete card information.
[0193] Here, for each determined target paragraph, the corresponding card information can be generated and displayed once. During the process of displaying the card information corresponding to the first target paragraph, the dynamic parsing engine can synchronously determine a second target paragraph from multiple sub-response data streamed out after the first target paragraph. After determining the second target paragraph, the card information being displayed can be updated. For example, as Figure 6 shown, the card information display area can update the chart. Additionally, the TTS display area can display a prompt message "Generating chart..." to prompt the user that the card information has not been fully displayed.
[0194] If the sub-response data includes an end flag, it can be determined that the sub-response data is the complete target response data. Therefore, the target paragraph determined based on this sub-response data can also include an end flag. If the target paragraph includes an end flag, it can be determined that the target paragraph is the last target paragraph. Furthermore, after generating and displaying the card information corresponding to the last target paragraph, the update of the card information can be completed to obtain the complete card information.
[0195] In this way, by displaying a part of the card information for each obtained target paragraph, the timeliness of card information display can be ensured, further enhancing the user's visual experience.
[0196] Additionally, the sub-response data output by the large model module can also include display template information. Among them, the display template information included in multiple sub-response data can be the same. Additionally, the display template information can be the template type information corresponding to the display template of the sub-response data. Among them, the template type can include descriptive templates, general classification templates, comparison templates, timeline templates, and service expert templates. Additionally, different template types can correspond to different card types. Among them, the card types can include descriptive cards, general classification cards, and service expert cards. Among the descriptive cards, general classification cards, and service expert cards, the cards can further include cards with pictures and cards without pictures. The correspondence between the template type and the card type can be as Figure 7 shown.
[0197] As an example, after receiving the interactive input data, the large model module can also determine the reply type corresponding to the interactive input data, and determine the display template type of the response data corresponding to the interactive input data based on the reply type.
[0198] Based on this, in order to further improve the display effect of the card information and enhance the user's visual experience, in some embodiments, the above-mentioned rendering of the second interactive description information to generate card information can specifically include:
[0199] Render the second interactive description information according to the display template information to generate card information.
[0200] Here, the dynamic parsing engine may include a template interceptor, which can intercept and record the display template information in the sub-response data. If the dynamic parsing engine can intercept the display template information, the rich media structuring engine can render the second interaction description information according to the display template information to generate card information. Specifically, the rich media structuring engine can launch a HyperText Markup Language interface (such as HTML 5, H5), and transmit the second interaction description information to H5 through the JS bridge. The client can launch an H5 interface corresponding to the template type through a browser (such as a webview container), load the H5 template and the second interaction description information, and generate card information.
[0201] In this way, by using templates and data to dynamically generate the user interface (i.e., card information), the structure and style of the user interface can be separated from the data, so that different user interfaces can be dynamically generated according to different data, further improving the display effect of card information and enhancing the user's visual experience.
[0202] In addition, the sub-response data may further include the source information of the sub-response data. The source information can characterize the way the artificial intelligence large model obtains the response data. For example, if the response data is the response data corresponding to the interaction input data obtained by the artificial intelligence large model by searching in a pre-set question and answer library corresponding to the vehicle, the source information of the response data (including multiple sub-response data output in a streaming manner) can be the question and answer library corresponding to the vehicle.
[0203] As an example, when the artificial intelligence large model obtains the response data, it can synchronously determine the source information corresponding to the response data and mark the source information in the response data (including multiple sub-response data output in a streaming manner).
[0204] Based on this, in order to further enhance the user experience of using the question and answer system, in some embodiments, the above-mentioned rendering of the second interaction description information according to the display template information to generate card information may specifically include:
[0205] In the case where the source information of the response data is the question and answer library corresponding to the vehicle, modify the display template information to a service expert template, and the service expert template is a display template corresponding to the vehicle question and answer;
[0206] Render the second interaction description information according to the service expert template to generate card information.
[0207] Here, the service expert template can be a display template corresponding to the vehicle Q&A. The text description in the service expert template can be the operation steps related to the vehicle usage service. The card information generated based on the service expert template can include the following content, for example: "The usage steps of the child safety lock are as follows: 1. Click 'Vehicle' in the central control screen settings, select 'Door Locks', select the option below the 'Child Lock Button' to set the child lock button; 2. After selecting 'Left' or 'Right', when the child safety lock is turned on, the corresponding rear door cannot be opened from the inside and the window cannot be controlled."
[0208] In this way, when the source information of the response data is the Q&A library corresponding to the vehicle, by rendering the second interaction description information according to the service expert template corresponding to the vehicle Q&A to generate card information, the professionalism of the card information can be improved, and the user experience of using the Q&A model can be further enhanced.
[0209] In addition, in order to ensure the timeliness of the response data broadcast and display and improve the user experience, in some embodiments, after determining the target service corresponding to the Q&A scenario among multiple services, it may further include:
[0210] Performing a dialogue interaction on the first output sub-response data through the target service, where the dialogue interaction includes voice interaction and data display associated with the voice interaction;
[0211] Starting from the second output sub-response data, determine the sub-response data other than the previous sub-response data in the sub-response data as the second sub-response data;
[0212] Performing a dialogue interaction on the second sub-response data until the sub-response data is the last one among the multiple sub-response data.
[0213] Here, the voice interaction can include voice broadcasting of the sub-response data. The data display associated with the voice interaction can be the on-screen display of the text information during the voice interaction. Among them, the data can be displayed in the TTS display area. The displayed data can include text data, image data, video data, etc.
[0214] As an example, if the Q&A system is an in-vehicle Q&A system, and the large model module in the in-vehicle Q&A system is located in the cloud server, and the other modules except the large model module are located in the terminal server, then every time the large model module outputs a sub-response data, the sub-response data can be sent to the terminal server through the SDK. The terminal server can receive the sub-response data through the main interaction module. After the main interaction module receives the sub-response data, it can send the sub-response data to the dialogue management module and send the sub-response data to the TTS structured engine through the dialogue management module. Every time the TTS structured engine receives a sub-response data, it can process the sub-response data based on the append broadcast strategy to obtain the second sub-response data for voice broadcast and on-screen display. After obtaining the second sub-response data, on the one hand, the TTS structured engine can send the second sub-response data to the audio module through the TTS service, and the audio module can perform voice broadcast on the second sub-response data. On the other hand, the TTS structured engine can send the second sub-response data to the generative user interface service through the TTS service, and the generative user interface service can perform streaming on-screen display on the second sub-response data.
[0215] In addition, the TTS structured engine can also determine whether to perform voice broadcast and on-screen display on the second sub-response data based on the interruption broadcast strategy, and manage the TTS broadcast status. Among them, the TTS broadcast status can include which sub-response data is being broadcast. In addition, the interruption broadcast strategy can be a broadcast strategy executed when the session identifier of the latter sub-response data is different from that of the previous sub-response data. The append broadcast strategy can be a broadcast strategy executed when the session identifiers of the latter sub-response data and the previous sub-response data are the same.
[0216] As an example, the sub-response data can include a session identifier. If the session identifier of the latter sub-response data is the same as that of the previous sub-response data, the two sub-response data can be the sub-response data corresponding to the same interaction input data, and then the voice broadcast and on-screen display of the second sub-response data corresponding to the session identifier can be continued. If the session identifier of the latter sub-response data is different from that of the previous sub-response data, the two sub-response data are not the sub-response data corresponding to the same interaction input data. That is, during the process of answering the current interaction input data, if another interaction input data is received, the answer to the current interaction input data can be stopped and the answer to the other interaction input data can be switched to.
[0217] As a more specific example, as Figure 8 shown, the schematic diagram of the process of dialogue interaction on the sub-response data by the TTS structured engine can include the following steps:
[0218] S81. Receive the first sub-response data streamed out by the large model module, and perform voice broadcast and on-screen display on the first sub-response data;
[0219] S82. Receive the i-th (i≥2) sub-response data streamed out by the large model module;
[0220] S83. Determine whether the session identifier of the i-th sub-response data is the same as that of the (i - 1)-th sub-response data. If so, execute S84; if not, execute S86;
[0221] S84. Record the index value of the i-th sub-response data;
[0222] S85. Based on the append broadcast strategy, determine the second sub-response data as the part of the i-th sub-response data except the (i - 1)-th sub-response data, and perform voice broadcast and on-screen display on the second sub-response data;
[0223] S86. Reset the index value and session identifier of the i-th sub-response data;
[0224] S87. Based on the interrupt broadcast strategy, perform voice broadcast and on-screen display on the i-th sub-response data;
[0225] S88. Let i = i + 1, and return to execute S82 until the sub-response data includes an end identifier.
[0226] In this way, by generating voice frame-by-frame according to the text, voice synthesis becomes a streaming process, which can output sound in real time and play simultaneously with the generation of voice. This method can make the synthesized voice smoother and reduce latency. Thus, by performing streaming broadcast and streaming on-screen display on multiple sub-response data streamed out by the large model module, it is possible to broadcast and display a part of the sub-response data every time a part of the sub-response data is generated, ensuring the timeliness of the broadcast and display of the response data and improving the user experience.
[0227] Based on this, since in the process of the large model module streaming out multiple sub-response data, the first sub-response data and at least one second sub-response data have been streamed on-screen and are displayed in the TTS display area. Therefore, in order to ensure both the timeliness of the display of the response data and the user's reading experience of the response data, and further improve the user experience, in some embodiments, the above-mentioned display of TTS structured data in the TTS display area may specifically include:
[0228] Replace the displayed first sub-response data and at least one second sub-response data with TTS structured data.
[0229] Here, the process of streaming multiple sub-response data, the process of streaming the sub-response data onto the screen, and the process of TTS structuring can be carried out synchronously without affecting each other. That is, after the large model module outputs the last sub-response data (i.e., the complete target response data), the streaming display of the sub-response data on the screen may not have ended yet. On the one hand, the sub-response data can continue to be streamed onto the screen. On the other hand, the TTS structuring engine can structure the target response data to obtain TTS structured data. After obtaining the TTS structured data, the streaming display of the sub-response data on the screen may have ended or may not have ended yet, which is not limited here. If the streaming display has ended, a large block of plain text that has been displayed on the screen can be replaced with the TTS structured data. If the streaming display has not ended, the first sub-response data and at least one second sub-response data that have been displayed on the screen can be replaced with the TTS structured data.
[0230] In this way, by generating and displaying a part of the sub-response data each time, the timeliness of the display of the response data can be ensured. By replacing the streamed data that has been displayed with the TTS structured data after generating the TTS structured data, the reading experience of the user for the response data can be improved. In this way, it is possible to ensure both the timeliness of the display of the response data and the reading experience of the user for the response data, further enhancing the user experience.
[0231] To better describe the entire solution, based on the above embodiments, some specific examples are given.
[0232] For example, the interaction schematic diagram of a question-and-answer method provided by an embodiment of the present application can be as Figure 9 shown.
[0233] In Figure 9Among them, the question-and-answer method can be applied to an in-vehicle question-and-answer system. The large model module can be set in the cloud server. The main interaction module, the dialogue management module, the service management module, and the AI-type service assistant can be set in the terminal server. Specifically, the main interaction module of the terminal server can receive the interaction input data input by the user. If the input form of the interaction input data is non-voice input, the main interaction module can determine the user intention description information corresponding to the interaction input data by itself. If the input form of the interaction input data is voice input, the main interaction module can send the interaction input data to the large model module. The large model module can input the interaction input data into the artificial intelligence large model to identify the interaction input data through the large voice model and generate the user intention description information. After determining the user intention description information, the large model module can return the user intention description information to the main interaction module. After receiving the user intention description information, the main interaction module can send the user intention description information to the dialogue management module. In addition, after determining the user intention description information, the large model module can continue to answer the interaction input data based on the user intention description information, generate the answer data and cache it.
[0234] After receiving the user intention description information, the dialogue management module can determine the user interaction scenario corresponding to the user intention description information according to the arbitration strategy, and send the user interaction scenario to the service management module. After receiving the user interaction scenario, the service management module can determine the target service assistant corresponding to the user interaction scenario according to the correspondence between the user interaction scenario and the registered intention, and the correspondence between the registered intention and the service assistant. When the user interaction scenario is a question-and-answer scenario, the service management module can determine that the target service assistant is the AI-type service assistant, and notify the AI-type service assistant to perform scenario registration to obtain the scenario identifier. The service management module can record the correspondence between the scenario identifier and the AI-type service assistant, and send the scenario identifier to the dialogue management module.
[0235] After receiving the scenario identifier, the dialogue management module can send a streaming data request corresponding to the scenario identifier to the large model module through the main interaction module. After receiving the streaming data request, the large model module can stream output the answer data corresponding to the scenario identifier. That is, the large model module can stream send multiple sub-answer data (including the scenario identifier) to the main interaction module. Each time the main interaction module receives a sub-answer data, it can send a sub-answer data to the dialogue management module. Each time the dialogue management module receives a sub-answer data, it can send the sub-answer data to the service management module once. Each time the service management module receives a sub-answer data, it can route the sub-answer data to the AI-type service assistant corresponding to the scenario identifier according to the correspondence between the scenario identifier and the service assistant.
[0236] After receiving the sub-response data, on the first hand, the AI type service assistant can perform TTS voice broadcast and streaming screen display based on multiple sub-response data through the TTS broadcast engine. On the second hand, the AI type service assistant can determine the target paragraph based on multiple sub-response data through the rich media structuring engine, send a rich media structuring request to the large model module, and generate and display card information based on the second interaction description information returned by the large model module. On the third hand, the AI type service assistant can determine the complete target response data based on multiple sub-response data through the TTS structuring engine, send a TTS structuring request to the large model module, and perform marking processing based on the first interaction description information returned by the large model module to obtain and display TTS structured data. Among them, displaying the TTS structured data can be replacing the first sub-response data that has been displayed on the screen and multiple second sub-response data with the TTS structured data.
[0237] In addition, the large model module can also be set in the terminal server, which is not limited here.
[0238] Thus, through the embodiments of the present application, on the first hand, by generating a part of the sub-response data each time and immediately broadcasting and displaying it, the timeliness of the response data broadcast and display can be ensured, enhancing the user experience. On the second hand, by performing TTS structuring on the complete target response data, generating and displaying the TTS structured data, compared with displaying a large paragraph of plain text, the reading experience of the user for the target response data can be improved. On the third hand, by performing rich media structuring on the target paragraph, generating card information, and displaying the card information based on the generative UI technology, the visual experience of the user can be further enhanced. On the fourth hand, by synchronously performing multiple processes such as TTS voice broadcast, TTS streaming screen display, TTS structuring, and rich media structuring based on the dynamic parsing engine during the process of streaming output of multiple sub-response data, the response speed to the user instruction (i.e., the interactive input data) can be increased, further enhancing the user experience.
[0239] Based on the question-answering method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of the question-answering device. Please refer to the following embodiments.
[0240] As Figure 10 shown, the question-answering device 1000 provided by the embodiments of the present application includes the following modules:
[0241] A receiving module 1010, configured to receive the interactive input data of the user;
[0242] A first processing module 1020, configured to process the interactive input data by using an artificial intelligence large model to obtain a generated content information set, where the generated content information set includes user intention description information, response data, and interaction description information;
[0243] The first determination module 1030 is configured to determine a user interaction scenario corresponding to the user intention description information;
[0244] The second determination module 1040 is configured to, when the user interaction scenario is a question-and-answer scenario, determine a target service corresponding to the question-and-answer scenario from multiple services, where the service corresponds to the user interaction scenario and is used to perform subsequent actions corresponding to the interaction input data;
[0245] The second processing module 1050 is configured to perform structured processing on the answer data based on the interaction description information through the target service, and obtain and display the structured data.
[0246] The above question-and-answer device 1000 will be described in detail below, as follows:
[0247] In some embodiments, the second determination module 1040 may specifically include:
[0248] The first acquisition sub-module is configured to acquire registration intents respectively corresponding to multiple services;
[0249] The first determination sub-module is configured to determine a target service corresponding to the question-and-answer scenario according to the corresponding relationship between the registration intent and the user interaction scenario.
[0250] In some embodiments, the second processing module 1050 may specifically include:
[0251] The registration sub-module is configured to perform scenario registration on the question-and-answer scenario through the target service to obtain a scenario identifier;
[0252] The routing sub-module is configured to route the answer data corresponding to the scenario identifier to the target service;
[0253] The first processing sub-module is configured to perform structured processing on the answer data based on the interaction description information through the target service, and obtain and display the structured data.
[0254] In some embodiments, the answer data includes a plurality of sub-answer data output in a streaming manner. Based on this, the second processing module 1050 may specifically further include:
[0255] The second determination sub-module is configured to determine first sub-answer data from the plurality of sub-answer data, where the first sub-answer data is the answer data for performing structured processing;
[0256] The second acquisition sub-module is configured to use an artificial intelligence large model to acquire interaction description information corresponding to the first sub-answer data;
[0257] The second processing sub-module is configured to perform structured processing on the first sub-answer data based on the interaction description information, and obtain and display the structured data.
[0258] In some of these embodiments, the first sub-response data includes complete target response data corresponding to the interactive input data, and the interactive description information includes first interactive description information corresponding to the target response data. Based on this, the second processing module 1050 may specifically include:
[0259] A third processing sub-module, configured to perform marking processing on the first interactive description information in the target response data to obtain TTS structured data;
[0260] A first display sub-module, configured to display the TTS structured data in the TTS display area.
[0261] In some of these embodiments, the first sub-response data includes a target paragraph, and the interactive description information includes second interactive description information corresponding to the target paragraph, and the second interactive description information includes picture information. Based on this, the second processing module 1050 may specifically include:
[0262] A rendering sub-module, configured to render the second interactive description information to generate card information;
[0263] A second display sub-module, configured to display the card message in the card message display area.
[0264] In some of these embodiments, there are multiple target paragraphs. Based on this, the second display sub-module may specifically include:
[0265] An update unit, configured to update the displayed card information based on the multiple target paragraphs in the generation order of the multiple target paragraphs;
[0266] A display unit, configured to display the continuously updated card information in the card information display area until an end identifier is included in the target paragraph to obtain complete card information.
[0267] In some of these embodiments, among the multiple sub-response data output in a streaming manner, the latter sub-response data includes the former sub-response data. Based on this, the second determination sub-module may specifically include:
[0268] A judgment unit, configured to judge whether a preset delimiter is included in each sub-response data;
[0269] A first determination unit, configured to determine the sub-response data as a target paragraph when the sub-response data includes one of the preset delimiters;
[0270] A second determination unit, configured to determine the sub-response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph when the sub-response data includes multiple preset delimiters.
[0271] In some of these embodiments, the response data includes a plurality of sub-response data output in a streaming manner, where the subsequent sub-response data includes the previous sub-response data. Based on this, the Q&A device 1000 may further include:
[0272] A first dialogue module, configured to, after determining a target service corresponding to the Q&A scenario among a plurality of services, perform a dialogue interaction on the first output sub-response data through the target service, where the dialogue interaction includes voice interaction and data display associated with the voice interaction;
[0273] A third determination module, configured to start from the second output sub-response data, determine the sub-response data other than the previous sub-response data in the sub-response data as the second sub-response data;
[0274] A second dialogue module, configured to perform a dialogue interaction on the second sub-response data until the sub-response data is the last one among the plurality of sub-response data.
[0275] In some of these embodiments, the interaction input data includes multi-modal feature data, and the multi-modal feature data characterizes the input form of the interaction input data, where the input form of the interaction input data includes at least one of voice input, text input, touch input, and gesture input.
[0276] In the Q&A device according to the embodiments of the present application, the artificial intelligence large model can process the interaction input data to obtain a generated content information set including user intention description information, response data, and interaction description information. Based on this, since the service corresponds to the user interaction scenario, by determining the user interaction scenario corresponding to the user intention description information, and then performing subsequent actions corresponding to the interaction input data through the service corresponding to the user interaction scenario, the professionalism of performing subsequent actions can be ensured, and thus the accuracy of performing subsequent actions can be improved. By determining a target service corresponding to the Q&A scenario among a plurality of services when the user interaction scenario is the Q&A scenario, and performing subsequent actions corresponding to the interaction input data through the target service, the accuracy of answering the interaction input data can be improved. In addition, through the target service, the response data is structurally processed based on the interaction description information to obtain and display structured data. Compared with displaying a large paragraph of pure text, the key information can be highlighted, and thus the reading experience of the user for the response data can be improved. In this way, through the embodiments of the present application, both the accuracy of answering the interaction input data and the reading experience of the user for the response data can be improved.
[0277] Based on the Q&A method provided in the above embodiments, the embodiments of the present application also provide a specific implementation manner of an electronic device. Figure 11 FIG. 1100 shows a schematic diagram of an electronic device 1100 provided by the embodiments of the present application.
[0278] The electronic device 1100 may include a processor 1110 and a memory 1120 storing computer program instructions.
[0279] Specifically, the processor 1110 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0280] The memory 1120 may include a mass storage for data or instructions. By way of example and not limitation, the memory 1120 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In a suitable case, the memory 1120 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 1120 may be internal or external to the electronic device 1100. In a particular embodiment, the memory 1120 is a non-volatile solid state memory.
[0281] The memory may include a read only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer readable storage media (e.g., memory devices) encoded with software including computer executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0282] The processor 1110 reads and executes the computer program instructions stored in the memory 1120 to implement any one of the question and answer methods in the above embodiments.
[0283] In one example, the electronic device 1100 may further include a communication interface 1130 and a bus 1140. Among them, as Figure 11 shown, the processor 1110, the memory 1120, and the communication interface 1130 are connected through the bus 1140 and complete communication with each other.
[0284] The communication interface 1130 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application.
[0285] The bus 1140 includes hardware, software, or both, and couples the components of the electronic device together. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 1140 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0286] Exemplarily, the electronic device 1100 may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.
[0287] The electronic device may execute the Q&A method in the embodiments of the present application, thereby implementing the combination Figures 1 to 10 The described Q&A method, apparatus, and system.
[0288] In addition, in combination with the Q&A method in the above embodiments, the embodiments of the present application may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the Q&A methods in the above embodiments is implemented.
[0289] In addition, the embodiments of the present application further provide a vehicle, which may include at least one of the following:
[0290] The Q&A apparatus in any one of the embodiments of the second aspect;
[0291] The Q&A system in any one of the embodiments of the third aspect;
[0292] The electronic device in any one of the embodiments of the fourth aspect;
[0293] The computer-readable storage medium in any one of the embodiments of the fifth aspect. Details are not described herein again.
[0294] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0295] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0296] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0297] The various aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to generate a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0298] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. A question-answering method, characterized in that, Including: Receiving the interactive input data of the user; Processing the interactive input data by using an artificial intelligence large model to obtain a set of generated content information, where the set of generated content information includes user intention description information, response data, and interactive description information; Determining the user interaction scenario corresponding to the user intention description information; In the case where the user interaction scenario is a question-and-answer scenario, determining a target service corresponding to the question-and-answer scenario among multiple services, where the service corresponds to the user interaction scenario and is used to execute subsequent actions corresponding to the interactive input data; Through the target service, based on the interactive description information, performing structured processing on the response data to obtain and display structured data.
2. The method according to claim 1, characterized in that, The determining the target service corresponding to the question-and-answer scenario among multiple services includes: Obtaining the registered intentions corresponding to the multiple services respectively; Determining the target service corresponding to the question-and-answer scenario according to the corresponding relationship between the registered intention and the user interaction scenario.
3. The method according to claim 1, characterized in that, The performing structured processing on the response data based on the interactive description information through the target service to obtain and display structured data includes: Performing scenario registration on the question-and-answer scenario through the target service to obtain a scenario identifier; Routing the response data corresponding to the scenario identifier to the target service; Through the target service, based on the interactive description information, performing structured processing on the response data to obtain and display structured data.
4. The method according to claim 1 or 3, characterized in that, The response data includes a plurality of sub-response data output in a streaming manner, and the performing structured processing on the response data based on the interactive description information to obtain and display structured data includes: Determining a first sub-response data among the plurality of sub-response data, where the first sub-response data is the response data for performing structured processing; Using an artificial intelligence large model to obtain the interactive description information corresponding to the first sub-response data; Based on the interactive description information, performing structured processing on the first sub-response data to obtain and display structured data.
5. The method according to claim 4, characterized in that, The first sub-response data includes the complete target response data corresponding to the interactive input data, and the interactive description information includes the first interactive description information corresponding to the target response data. The performing structured processing on the first sub-response data based on the interactive description information to obtain and display structured data includes: Performing marking processing on the first interactive description information in the target response data to obtain TTS structured data; Displaying the TTS structured data in the TTS display area.
6. The method according to claim 4, characterized in that, The first sub-response data includes a target paragraph, and the interactive description information includes the second interactive description information corresponding to the target paragraph. The second interactive description information includes picture information. The performing structured processing on the first sub-response data based on the interactive description information to obtain and display structured data further includes: Rendering the second interactive description information to generate card information; Displaying the card message in the card message display area.
7. The method according to claim 6, characterized in that, The target paragraphs are multiple, and the displaying the card information in the card information display area includes: Update the displayed card information based on multiple target paragraphs in the generation order of the multiple target paragraphs; Display the continuously updated card information in the card information display area until an end identifier is included in the target paragraph to obtain the complete card information.
8. The method according to claim 6, characterized in that, Among the multiple sub-response data output in a streaming manner, each subsequent sub-response data includes the previous sub-response data. Determining the first sub-response data among the multiple sub-response data includes: For each sub-response data, determine whether a preset delimiter is included in the sub-response data; When one preset delimiter is included in the sub-response data, determine the sub-response data as the target paragraph; When multiple preset delimiters are included in the sub-response data, determine the sub-response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
9. The method according to claim 1, characterized in that, The response data includes multiple sub-response data output in a streaming manner, where each subsequent sub-response data includes the previous sub-response data. After determining the target service corresponding to the Q&A scenario among multiple services, the method further includes: Perform a dialogue interaction on the first output sub-response data through the target service, where the dialogue interaction includes voice interaction and data display associated with the voice interaction; Starting from the second output sub-response data, determine the sub-response data other than the previous sub-response data in the sub-response data as the second sub-response data; Perform a dialogue interaction on the second sub-response data until the sub-response data is the last one among the multiple sub-response data.
10. According to the method described in claim 1, characterized in that, The interaction input data includes multi-modal feature data, and the multi-modal feature data characterizes the input form of the interaction input data. The input form of the interaction input data includes at least one of voice input, text input, touch input, and gesture input.
11. A question-answering device, characterized in that, The device includes: A receiving module, configured to receive the interaction input data of the user; A first processing module, configured to process the interaction input data by using an artificial intelligence large model to obtain a generated content information set, where the generated content information set includes user intention description information, response data, and interaction description information; A first determination module, configured to determine the user interaction scenario corresponding to the user intention description information; A second determination module, configured to, when the user interaction scenario is a Q&A scenario, determine the target service corresponding to the Q&A scenario among multiple services, where the service corresponds to the user interaction scenario and is used to perform subsequent actions corresponding to the interaction input data; A second processing module, configured to perform structured processing on the response data based on the interaction description information through the target service to obtain and display structured data.
12. A question-answering system, characterized in that, Includes: A main interaction module, configured to receive the interaction input data of the user; A large model module, configured to process the interaction input data by using an artificial intelligence large model to obtain a generated content information set, where the generated content information set includes user intention description information, response data, and interaction description information; The dialogue management module is used to determine the user interaction scenario corresponding to the user intention description information, and send the user interaction scenario to the service management module. It is also used to determine the scenario identifier corresponding to the response data, and send the scenario identifier and the corresponding response data to the service management module; The service management module is used to determine the target service assistant corresponding to the user interaction scenario sent by the dialogue management module according to the correspondence between the service assistant and the registered intention, and the correspondence between the registered intention and the user interaction scenario. It is also used to route the response data corresponding to the scenario identifier to the target service assistant according to the correspondence between the scenario identifier and the target service assistant; The service assistant is used to register the user interaction scenario it is concerned about with the service management module to obtain the registered intention. It is also used to perform scenario registration on the user interaction scenario to obtain a scenario identifier, and process the response data corresponding to the scenario identifier. The service assistant includes the target service assistant.
13. According to the system described in claim 12, characterized in that, The service assistant includes a task-based service assistant and an AI-based service assistant; The task-based service assistant is used to process tasks corresponding to the interaction control scenario; The AI-based service assistant is used to process tasks corresponding to the question-and-answer scenario.
14. According to the system described in claim 13, characterized in that, The response data includes a plurality of sub-response data output in a streaming manner. The AI-based service assistant includes a TTS playback engine, a TTS structuring engine, a rich media structuring engine, and a dynamic parsing engine; The TTS playback engine is used to perform dialogue interaction on the plurality of sub-response data output in a streaming manner; The TTS structuring engine is used to perform TTS structuring processing on the complete target response data corresponding to the interaction input data to obtain TTS structured data; The rich media structuring engine is used to perform rich media structuring processing on the target paragraph corresponding to the interaction input data to obtain card information; The dynamic parsing engine is used to perform dynamic parsing on the plurality of sub-response data, the target response data, and the plurality of target paragraphs during the process of outputting the plurality of sub-response data in a streaming manner, performing the dialogue interaction, performing the TTS structuring processing, and performing the rich media structuring processing.
15. According to the system described in claim 14, characterized in that, The system further includes a generative user interface service; The generative user interface service is used to perform dynamic display on the TTS structured data and the card information.
16. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the question-and-answer method according to any one of claims 1-10.
17. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium. When the computer program instructions are executed by the processor, they implement the question-and-answer method according to any one of claims 1-10.
18. A vehicle, characterized in that, Including at least one of the following: The question-answering device according to claim 11; The question-and-answer system according to any one of claims 12-15; The electronic device according to claim 16; The computer-readable storage medium according to claim 17.
Citation Information
Cited By
Interaction method and interaction device based on large model and vehicle
CN120705296A
Voice interaction system for vehicle
CN120895036A
Interaction method and device, electronic equipment and computer readable storage medium
CN121116146A
Interaction method and device, electronic equipment and computer readable storage medium
CN121116146B