Question and answer method, device and equipment, storage medium and vehicle
By using artificial intelligence models in the Q&A system to identify and respond to user input, determine the target paragraph and obtain card information, the problem of low reading experience of answer data in the existing Q&A system is solved, and more efficient information display and user experience are achieved.
Patent Information
- Application Number
- CN202311746295.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-20
AI Technical Summary
When existing question and answer systems display response data, text and pictures are usually structured up and down, resulting in a lower user reading experience.
By receiving user's interactive input data, inputting artificial intelligence models to identify and respond, determining the target paragraph, and obtaining card information corresponding to the target paragraph, rendering and displaying card information.
It improves the user's reading experience with response data, reduces the user's reading burden, and allows users to obtain the desired information clearly and accurately.
Smart Images

Figure CN120181218A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of intelligent question answering, and particularly relates to a question answering method, device, equipment, storage medium and vehicle. Background Art
[0002] With the development of artificial intelligence, question answering systems are increasingly widely used. Generally, a question answering system can respond to a user instruction, output and display response data corresponding to the user instruction.
[0003] Currently, a question answering system usually first obtains the user's voice data, and then determines the graphic information (i.e., response data) corresponding to the voice data according to pre-set question-answer pairs. Among them, the question-answer pairs are usually fixed, and the response data corresponding to the voice data usually includes a large paragraph of plain text and pictures. That is, the currently displayed response data can include a large paragraph of plain text and pictures corresponding to the voice data pre-set. However, the text and pictures are usually in an upper-lower structure. That is, a large paragraph of plain text is displayed in the upper half of the display interface, and pictures are displayed in the lower half. Or pictures are displayed in the upper half, and a large paragraph of plain text is displayed in the lower half. The above all result in a low reading experience of the response data for users. Summary of the Invention
[0004] Embodiments of this application provide a question answering method, device, equipment, storage medium and vehicle, which can improve the reading experience of users for response data.
[0005] In a first aspect, embodiments of this application provide a question answering method, which includes:
[0006] Receiving the user's interactive input data, where the interactive input data includes multi-modal feature data and instruction interaction data, the multi-modal feature data characterizes the input form of the interactive input data, and the instruction interaction data characterizes the user intention corresponding to the interactive input data;
[0007] Inputting the interactive input data into an artificial intelligence large model to identify and respond to the interactive input data through the artificial intelligence large model, and obtaining and outputting multiple response data;
[0008] During the process of the artificial intelligence large model outputting the multiple response data, determining a target paragraph based on the output multiple response data;
[0009] Inputting the target paragraph into the artificial intelligence large model to obtain card information corresponding to the target paragraph through the artificial intelligence large model;
[0010] Rendering and displaying the card information corresponding to the target paragraph.
[0011] In a possible implementation, inputting the interactive input data into the large artificial intelligence model to identify and respond to the interactive input data through the large artificial intelligence model, and obtaining and outputting multiple response data, including:
[0012] Input the interactive input data into the large artificial intelligence model to identify the interactive input data through the large artificial intelligence model, obtain a response type corresponding to the interactive input data, and determine display template information corresponding to the response type; and respond to the interactive input data through the large artificial intelligence model, and obtain and output multiple response data including the display template information.
[0013] In a possible implementation, the target paragraph includes display template information. Inputting the target paragraph into the large artificial intelligence model to obtain card information corresponding to the target paragraph through the large artificial intelligence model, including:
[0014] Input the target paragraph into the large artificial intelligence model to determine text interaction description information corresponding to the display template information through the large artificial intelligence model, extract the text interaction description information from the target paragraph, and search for multimedia interaction description information corresponding to the text interaction description information;
[0015] Determine the text interaction description information and the multimedia interaction description information as the card information.
[0016] In a possible implementation, rendering and displaying the card information corresponding to the target paragraph includes:
[0017] Fill valid information into the display template corresponding to the display template information to obtain card information corresponding to the target paragraph, where the valid information includes at least one of the text interaction description information and the multimedia interaction description information;
[0018] Display the card information in the card information display area.
[0019] In a possible implementation, there are multiple target paragraphs. Displaying the card information in the card information display area includes:
[0020] Update the displayed card information based on the multiple target paragraphs according to the generation order of the multiple target paragraphs;
[0021] Display the updated card information in the card information display area until an end identifier is included in the target paragraph to obtain the complete card information.
[0022] In a possible implementation, among the multiple response data output, the subsequent response data includes the previous response data. Determining the target paragraph based on the multiple response data output in a streaming manner includes:
[0023] Among the multiple response data output in a streaming manner, for each response data, determine whether the response data includes a preset delimiter;
[0024] In the case where the response data includes one preset delimiter, determine the response data as the target paragraph;
[0025] In the case where the response data includes multiple preset delimiters, determine the response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
[0026] In a possible implementation, the response data includes a scene identifier. Before determining whether the response data includes a preset delimiter, the method further includes:
[0027] Determine whether the user interaction scene corresponding to the scene identifier is in a structured whitelist, where the structured whitelist includes multiple user interaction scenes;
[0028] Determining whether the response data includes a preset delimiter includes:
[0029] In the case where the user interaction scene corresponding to the scene identifier is in the structured whitelist, determine whether the response data includes a preset delimiter.
[0030] In a possible implementation, the response data includes display template information and a word count threshold. In the case where the user interaction scene corresponding to the scene identifier is in the structured whitelist, determining whether the response data includes a preset delimiter includes:
[0031] In the case where the user interaction scene corresponding to the scene identifier is in the structured whitelist, intercept and record the display template information and the word count threshold in the response data;
[0032] Obtain the word count in the response data;
[0033] Determine the magnitude relationship between the word count and the word count threshold;
[0034] In the case where the word count is greater than the word count threshold, determine whether the response data includes a preset delimiter.
[0035] In a possible implementation, when the number of words is greater than the word count threshold and before determining whether the response data includes a preset delimiter, the method further includes:
[0036] Intercept the network status in the response data;
[0037] When the network status is a network anomaly, display a target control for regenerating the card information.
[0038] In a second aspect, an embodiment of the present application provides a question-and-answer device, which includes:
[0039] A receiving module, configured to receive user interaction input data, where the interaction input data includes multi-modal feature data and instruction interaction data, the multi-modal feature data characterizes the input form of the interaction input data, and the instruction interaction data characterizes the user intention corresponding to the interaction input data;
[0040] A first input module, configured to input the interaction input data into an artificial intelligence large model to identify and respond to the interaction input data through the artificial intelligence large model, and obtain and output multiple response data;
[0041] A determination module, configured to determine a target paragraph based on the multiple output response data during the process of the artificial intelligence large model outputting the multiple response data;
[0042] A second input module, configured to input the target paragraph into the artificial intelligence large model to obtain card information corresponding to the target paragraph through the artificial intelligence large model;
[0043] A rendering module, configured to render and display the card information corresponding to the target paragraph.
[0044] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;
[0045] When the processor executes the computer program instructions, the method in any possible implementation method in the first aspect above is implemented.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method in any possible implementation method in the first aspect above is implemented.
[0047] In a fifth aspect, an embodiment of the present application provides a vehicle, which includes at least one of the following:
[0048] The Q&A device in any of the embodiments of the second aspect;
[0049] The electronic device in any of the embodiments of the third aspect;
[0050] The computer-readable storage medium in any of the embodiments of the fourth aspect.
[0051] The Q&A method, device, equipment, storage medium and vehicle according to the embodiments of the present application, by determining a target paragraph based on a plurality of response data output during the process of the artificial intelligence large model outputting a plurality of response data, and then inputting the target paragraph into the artificial intelligence large model, and using the artificial intelligence large model to obtain card information corresponding to the target paragraph, can determine key display information corresponding to the target paragraph. In this way, by rendering and displaying the card information corresponding to the target paragraph, the key display information can be displayed. Compared with displaying a large paragraph of pure text and pictures, it can reduce the reading burden of the user on the response data, enable the user to clearly and accurately obtain the information they want, and thus improve the user's reading experience of the response data. Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 It is a schematic flow chart of a Q&A method provided by the embodiments of the present application;
[0054] Figure 2 It is a schematic flow chart of determining a target paragraph through a dynamic parsing engine provided by the embodiments of the present application;
[0055] Figure 3 It is a schematic diagram of card information provided by the embodiments of the present application;
[0056] Figure 4 It is a schematic diagram of the corresponding relationship between a template type and a card type provided by the embodiments of the present application;
[0057] Figure 5 It is a schematic diagram of a voice interaction interface provided by the embodiments of the present application;
[0058] Figure 6 It is a schematic structural diagram of a Q&A device provided by the embodiments of the present application;
[0059] Figure 7 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed Embodiments
[0060] In order to more clearly understand the above-mentioned objects, features, and advantages of the present application, the solutions of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0061] In the following description, many specific details are set forth in order to fully understand the present application, but the present application can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all the embodiments.
[0062] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.
[0063] As described in the background art section, to solve the problems of the prior art, embodiments of the present application provide a question-and-answer method, apparatus, device, storage medium, and vehicle. Among them, the execution subject of the question-and-answer method can be a server. The server can be a terminal server, a cloud server, or a distributed server, which is not limited herein. Among them, the terminal server can include, for example, servers corresponding to mobile devices such as mobile phones, computers, and tablets, servers corresponding to smart wearable devices such as smart glasses and smart watches, and vehicle servers. In addition, the question-and-answer method can be specifically applied to the question-and-answer scenario in a vehicle-mounted environment. Based on this, the terminal server can be a vehicle server. In addition, the execution subject of the question-and-answer method can also include a terminal server and a cloud server at the same time. That is, if the question-and-answer method is applied to a question-and-answer system, some of the functional modules of the question-and-answer system can be set in the terminal server, and some other functional modules can be set in the cloud server. Among them, the terminal server and the cloud server can communicate through a network.
[0064] First, the question-and-answer method provided by the embodiments of the present application will be introduced below.
[0065] Figure 1 The flowchart of a question-and-answer method provided by an embodiment of the present application is shown. As Figure 1As shown in the figure, the Q&A method provided by the embodiment of the present application includes the following steps:
[0066] S110. Receive the interactive input data of the user. The interactive input data includes multi-modal feature data and instruction interaction data. The multi-modal feature data represents the input form of the interactive input data, and the instruction interaction data represents the user intention corresponding to the interactive input data;
[0067] S120. Input the interactive input data into the artificial intelligence large model to identify and answer the interactive input data through the artificial intelligence large model, and obtain and output multiple response data;
[0068] S130. During the process of the artificial intelligence large model outputting multiple response data, determine the target paragraph based on the output multiple response data;
[0069] S140. Input the target paragraph into the artificial intelligence large model to obtain the card information corresponding to the target paragraph through the artificial intelligence large model;
[0070] S150. Render and display the card information corresponding to the target paragraph.
[0071] The Q&A method of the embodiment of the present application determines the target paragraph based on the output multiple response data during the process of the artificial intelligence large model outputting multiple response data, and then inputs the target paragraph into the artificial intelligence large model. By using the artificial intelligence large model to obtain the card information corresponding to the target paragraph, the key display information corresponding to the target paragraph can be determined. In this way, by rendering and displaying the card information corresponding to the target paragraph, the key display information can be displayed. Compared with displaying a large paragraph of pure text and pictures, it can reduce the reading burden of the user on the response data, enable the user to clearly and accurately obtain the information they want, and thus improve the reading experience of the user on the response data.
[0072] The specific implementation manners of the above steps are introduced below.
[0073] In some embodiments, in S110, the interactive input data may be multi-modal instruction interaction data input by the user. That is, the interactive input data may include multi-modal feature data and instruction interaction data. Among them, the multi-modal feature data may represent the input form of the interactive input data. The input form of the interactive input data may include, for example, voice input, text input, touch input, and gesture input. In addition, the instruction interaction data may represent the user intention corresponding to the interactive input data. The instruction interaction data may include, for example, control instructions related to vehicle control and vehicle settings, question instructions related to vehicle Q&A and general Q&A, and may also include both control instructions and question instructions, which are not limited herein.
[0074] As an example, a question-and-answer system may include a main interaction module and a large model module. Among them, the main interaction module may have the function of Automatic Speech Recognition (ASR). If the input form of the interaction input data is voice input, the main interaction module may convert the voice command into command text through the ASR function and send the command text to the large model module. After receiving the command text, the large model module may perform semantic parsing on the command text through an artificial intelligence large model to obtain command interaction data (i.e., determine the user's intention), and return the command interaction data to the main interaction module.
[0075] In addition, the question-and-answer system may further include a voice user interface (VUI) corresponding to the main interaction module. The voice user interface may include a text input box. The user may enter a text command in the text input box. The main interaction module may determine the command interaction data corresponding to the interaction input data based on the correspondence between the text command and the command interaction data.
[0076] In addition, the voice user interface may further include a plurality of interaction controls, and the interaction controls may include preset commands related to vehicle control, vehicle settings, vehicle question-and-answer, general question-and-answer, etc. The user may perform a touch command by clicking on the interaction control. The main interaction module may determine the command interaction data corresponding to the interaction input data based on the correspondence between the touch command and the command interaction data.
[0077] In addition, if the voice user interface supports gesture air touch, the user may select any one of the preset commands by gesture air controlling the interaction control. The main interaction module may determine the command interaction data corresponding to the interaction input data based on the correspondence between the gesture command and the command interaction data.
[0078] In some embodiments, in S120, the artificial intelligence (AI) large model refers to a high-performance artificial intelligence model constructed through a large number of training samples and computing resources. They can learn a large amount of language knowledge, image features, and speech patterns, and can reason and generate outputs similar to humans, and have a wide range of applications in the fields of natural language processing, image recognition, speech recognition, etc. The artificial intelligence large model may include, for example, a large language model (LLM), ChatGPT (Chat Generative Pre-trained Transformer), a multi-modal large model, and a multi-modal cognitive large model, etc.
[0079] As an example, if the input form of the interactive input data is voice input, after receiving the user's interactive input data, the main interaction module can send the interactive input data to the large model module. The large model module can input the interactive input data into the artificial intelligence large model, and the artificial intelligence large model can identify and respond to the interactive input data. Among them, identifying the interactive input data can generate instruction interactive data. Responding to the interactive input data can generate response data corresponding to the instruction interactive data. If the input form of the interactive input data is non-voice input (including text input, touch input, and gesture input), after receiving the user's interactive input data, the main interaction module can determine the instruction interactive data corresponding to the interactive input data by itself and send the instruction interactive data to the large model module. The large model module can input the instruction interactive data into the artificial intelligence large model, and the artificial intelligence large model can respond to the instruction interactive data. Among them, responding to the instruction interactive data can generate response data corresponding to the instruction interactive data.
[0080] As an example, if the large model module is located in the cloud server and the main interaction module and the dialogue management module are located in the terminal server, after the terminal server sends the interactive input data to the cloud server through the main interaction module, it can continue to send an asynchronous generation interface to the cloud server, so that the large model module returns the instruction interactive data to the terminal server. After the large model module returns the instruction interactive data to the terminal server, it can continue to obtain the response data corresponding to the interactive input data (including the instruction interactive data) and cache the obtained response data. Among them, obtaining the response data corresponding to the interactive input data includes searching for the response data corresponding to the interactive input data in a pre-set Q&A library, searching for the response data corresponding to the interactive input data in the network, and generating the response data corresponding to the interactive input data. During the process of the artificial intelligence large model generating the response data corresponding to the interactive input data, when predicting the next word or character, the previously generated content will be considered. This context-aware generation method enables the artificial intelligence large model to generate text with coherence and certain logic.
[0081] In addition, after receiving the instruction interactive data, the main interaction module in the terminal server can send the instruction interactive data to the dialogue management module. After receiving the instruction interactive data, the dialogue management module can determine the user interaction scenario corresponding to the instruction interactive data according to the arbitration strategy. Among them, the user interaction scenario can include a Q&A scenario and an interaction control scenario. Specifically, the Q&A scenario can include a general Q&A scenario and a vehicle Q&A scenario. The interaction control scenario can include a vehicle control scenario. In addition, the arbitration strategy can be the corresponding relationship between the pre-set instruction interactive data and the user interaction scenario.
[0082] When the dialogue management module determines that the user interaction scenario is the target Q&A scenario, it can send the target Q&A scenario to the service management module. Among them, the target Q&A scenario can specifically be any one of multiple Q&A scenarios such as a general Q&A scenario, a vehicle Q&A scenario, an encyclopedia Q&A scenario, a comparison Q&A scenario, a plan-making scenario, etc. After receiving the target Q&A scenario, the service management module can determine the target service assistant corresponding to the target Q&A scenario according to the correspondence between the user interaction scenario and the service assistant. Among them, the service assistant can correspond to the user interaction scenario and is used to execute subsequent actions corresponding to the response data. The target service assistant can be the service assistant corresponding to the target Q&A scenario. The target service assistant can be, for example, an AI-type service assistant. The AI-type service assistant can specifically include a dynamic parsing engine and a rich media structuring engine. After the service management module determines the target service assistant, it can notify the target service assistant to register the target Q&A scenario to obtain a scenario identifier. After that, the target service assistant can send the scenario identifier to the service management module. After receiving the scenario identifier, on the one hand, the service management module can record the correspondence between the target service assistant and the scenario identifier, and on the other hand, it can send the scenario identifier to the dialogue management module. The dialogue management module can send a streaming data request corresponding to the scenario identifier to the large model module. Among them, the streaming data request can be used to request the large model module to stream output multiple response data corresponding to the interaction input data. After receiving the streaming data request, the large model module can, in response to the streaming data request, stream output multiple response data. Among them, the latter response data can include the former response data. For example, if the interaction input data is "Introduce Zhou xx", the multiple response data streamed output can include: the first response data is "Zhou xx,", the second response data is "Zhou xx, from xxx,", and the third response data is "Zhou xx, from xxx, won the x award on x month x day...". In addition, the response data can include the scenario identifier.
[0083] As a more specific example, if the large model module is located in the cloud server and the main interaction module and the dialogue management module are located in the terminal server, the dialogue management module can send a streaming data request to the large model module through the main interaction module.
[0084] In addition, in some embodiments, the above S120 can specifically include:
[0085] Input the interaction input data into the artificial intelligence large model to identify the interaction input data through the artificial intelligence large model, obtain the reply type corresponding to the interaction input data, and determine the display template information corresponding to the reply type; and use the artificial intelligence large model to answer the interaction input data to obtain and output multiple response data including the display template information.
[0086] Here, the display template information may be template type information corresponding to the display template of the response data. Among them, the template type may include a description template, a general classification template, a comparison template, a timeline template, and a service expert template. The service expert template may be a display template corresponding to vehicle Q&A. The text description in the service expert template may be operation steps related to the usage service of the vehicle.
[0087] As an example, after receiving the interactive input data, the large model module can also determine the response type corresponding to the interactive input data through the artificial intelligence large model, and determine the template type corresponding to the response data of the interactive input data according to the correspondence between the response type and the display template information, obtain the display template information, and stream out multiple response data including the display template information. Among them, the display template information included in the multiple response data may be the same.
[0088] In some embodiments, in S130, the dynamic parsing engine can determine whether to perform rich media structuring processing on the response data, and which part of the response data to perform rich media structuring processing on. The dynamic parsing engine can specifically be used to determine the target paragraph for rich media structuring based on the multiple response data that have been streamed out. The target paragraph may be one or multiple, and is not limited herein.
[0089] In addition, as described above, among the multiple response data output, the latter response data may include the former response data. Based on this, in some embodiments, determining the target paragraph based on the multiple response data output specifically may include:
[0090] Among the multiple response data output, for each response data, determine whether the response data includes a preset delimiter;
[0091] In the case where the response data includes a preset delimiter, determine the response data as the target paragraph;
[0092] In the case where the response data includes multiple preset delimiters, determine the response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
[0093] Here, the multiple response data output may be multiple response data that have completed streaming output. The dynamic parsing engine may include a delimiter interceptor. The delimiter interceptor can determine whether the response data includes a preset delimiter. If the response data does not include a preset delimiter, it can continue to determine whether the next response data of the response data includes a preset delimiter. If the response data includes a preset delimiter, it can determine the target paragraph for rich media structuring based on the preset delimiter.
[0094] Specifically, the preset delimiter can be denoted as <|br / >, for example. If the response data is "AAA.BBB.<|br / >", then there is one preset delimiter in the response data, and the target paragraph can be "AAA.BBB.". If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >", then there are two preset delimiters in the response data, and the target paragraph can be "CCC.DDD.". If the response data is "AAA.BBB.<|br / >CCC.DDD.<|br / >EEE.FFF.<|br / >", then there are three preset delimiters in the response data, and the target paragraph can be "EEE.FFF.".
[0095] It should be noted that in the embodiments of the present application, the first process and the second process can be carried out synchronously without affecting each other. Among them, the first process can be a process of streaming out multiple response data by the artificial intelligence large model. The second process can be a process in which the dynamic parsing engine determines the target paragraph based on the multiple response data that have been streamed out.
[0096] In addition, as described above, the response data can include a scene identifier. Based on this, in order to ensure the necessity and effectiveness of structured processing, in some embodiments, before determining whether the response data includes a preset delimiter, it may further include:
[0097] Determine whether the user interaction scene corresponding to the scene identifier is in the structured whitelist, and the structured whitelist includes multiple user interaction scenes.
[0098] Based on this, determining whether the response data includes a preset delimiter can specifically include:
[0099] When the user interaction scene corresponding to the scene identifier is in the structured whitelist, determine whether the response data includes a preset delimiter.
[0100] Here, the scene identifier and the user interaction scene can be in one-to-one correspondence. The user interaction scene can include a question-and-answer scene and an interaction control scene. The question-and-answer scene can specifically include a general question-and-answer scene, a vehicle question-and-answer scene, an encyclopedia question-and-answer scene, a comparison question-and-answer scene, a plan-making scene, etc. The interaction control scene can specifically include a vehicle control scene, an intelligent device control scene, a mobile device control scene, etc.
[0101] In addition, the structured whitelist can be used to determine whether to perform structured processing on the response data. If the response data is in the structured whitelist, it can be determined that structured processing is performed on the response data. If the response data is not in the structured whitelist, it can be determined that structured processing is not performed on the response data.
[0102] As an example, the structured whitelist may include multiple user interaction scenarios. The multiple user interaction scenarios in the structured whitelist may include, for example, multiple Q&A scenarios such as general Q&A scenarios, vehicle Q&A scenarios, encyclopedia Q&A scenarios, comparison Q&A scenarios, plan-making scenarios, etc. If the user interaction scenario corresponding to the scenario identifier in the response data is in the structured whitelist, it can be determined that the response data is in the structured whitelist. If the user interaction scenario corresponding to the scenario identifier in the response data is not in the structured whitelist, it can be determined that the response data is not in the structured whitelist.
[0103] As an example, the dialogue management module can receive multiple response data streamed out by the large model module through the main interaction module. After receiving the response data, the dialogue management module can send the response data including the scenario identifier to the service management module. The service management module can route the response data to the target service assistant corresponding to the scenario identifier according to the correspondence between the scenario identifier and the target service assistant. After receiving the response data, the target service assistant can perform subsequent structured processing through the dynamic parsing engine in the target service assistant.
[0104] Specifically, every time the dynamic parsing engine receives a response data, it can determine whether the user interaction scenario corresponding to the scenario identifier in the response data is in the structured whitelist. If it is not in the structured whitelist, it can continue to judge the next response data. If it is in the structured whitelist, it can be determined that the response data meets the preliminary conditions for structured processing, and then it can continue to judge whether the response data includes a preset delimiter.
[0105] In this way, by continuing to judge whether the response data includes a preset delimiter when the response data is in the structured whitelist, the response data can be reasonably arranged to ensure the necessity and effectiveness of structured processing.
[0106] In addition, the response data may also include a word count threshold. The word count threshold may be a pre-set minimum number of words corresponding to the response data, used to judge whether to perform structured processing on the response data. Based on this, in order to avoid resource waste and ensure the reasonable use of resources, in some embodiments, the above-mentioned judgment of whether the response data includes a preset delimiter when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist may specifically include:
[0107] When the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, intercept and record the display template information and word count threshold in the response data;
[0108] Obtain the word count in the response data;
[0109] Judge the size relationship between the word count and the word count threshold;
[0110] When the number of words is greater than the word count threshold, it is determined whether the response data includes a preset delimiter.
[0111] Here, the dynamic parsing module may include a template interceptor. The template interceptor can intercept and record the display template information, word count threshold, source information of the response data, etc. in the response data. Among them, the template types may include a description template, a total score template, a comparison template, a timeline template, a service expert template, etc. The service expert template may be a display template corresponding to vehicle Q&A. For example, the service expert template may include operation steps related to the use service of the vehicle.
[0112] In addition, the dynamic parsing module may also include a word count interceptor. The word count interceptor can obtain the number of words (the length of the response data) included in the response data and determine the size relationship between the length and the word count threshold. If the length is less than the word count threshold, no subsequent processing may be performed. If the length is greater than the word count threshold, it is determined whether the response data includes a preset delimiter. That is, it is continued to determine whether structured processing needs to be performed based on the response data. In actual situations, if the number of words is less than the word count threshold, it can be considered that the response data is used for a short conversation with the user (such as chatting), and thus no structured processing and display of the response data may be performed.
[0113] In this way, by continuing to determine whether structured processing needs to be performed based on the response data when the number of words is greater than the word count threshold, resource waste can be avoided and the reasonable utilization of resources can be ensured.
[0114] Based on the above embodiments, a specific example is given. For example, the specific process of determining the target paragraph by the dynamic parsing engine may be as Figure 2 shown. As Figure 2 shown, the dynamic parsing engine may include an instruction interceptor, a template interceptor, a word count interceptor, a delimiter interceptor, and a structured interceptor.
[0115] As Figure 2 shown, a flow diagram of determining the target paragraph by the dynamic parsing engine provided by the embodiment of the present application may include the following steps:
[0116] S21. Determine whether the i-th response data is in the instruction whitelist. If so, execute S22; if not, execute S25;
[0117] S22. Intercept the control instruction in the response data through the instruction interceptor;
[0118] S23. Determine whether a control instruction is intercepted. If so, execute S24; if not, execute S25;
[0119] S24. Execute the control instruction;
[0120] S25. Determine whether the response data is in the structured whitelist. If so, execute S26; if not, set i = i + 1 and return to execute S21.
[0121] S26. Intercept and record key information such as the template type, word count threshold, source, etc. in the response data through the template interceptor.
[0122] S27. Obtain the word count in the response data through the word count interceptor.
[0123] S28. Determine whether the word count is greater than the word count threshold through the word count interceptor. If so, execute S29; if not, set i = i + 1 and return to execute S21.
[0124] S29. Determine whether the response data includes a preset delimiter through the delimiter interceptor. If so, execute S210; if not, set i = i + 1 and return to execute S21.
[0125] S210. Determine the target paragraph based on the preset delimiter through the structured interceptor, and initiate a rich media structured request corresponding to the target paragraph to the large model module. The rich media structured request is used to request the large model module to obtain the interactive description information corresponding to the target paragraph.
[0126] In Figure 2 On the one hand, by determining the target paragraph only when the response data meets the above multiple conditions, the response data can be reasonably arranged to ensure the necessity and effectiveness of the structured processing. On the other hand, for each response data, by intercepting and executing the control instruction in the response data when the response data is in the instruction whitelist, the timeliness of executing the user instruction can be ensured, improving the user experience.
[0127] In some embodiments, in S140, the card information may be card information in rich media form. The card information in rich media form may include audio, text, links, interactive controls, pictures, etc. corresponding to the target paragraph. Additionally, the artificial intelligence large model obtaining the card information corresponding to the target paragraph may specifically be obtaining the interactive description information corresponding to the target paragraph. The interactive description information may be used to generate the card information corresponding to the target paragraph. Among them, the interactive description information may include text interactive description information and multimedia interactive description information. The text interactive description information may, for example, include key information such as titles, subtitles, abstracts, preset keywords, etc. If the target paragraph is a target paragraph related to comparison, the interactive description information may further include comparison objects, comparison items, comparison contents, etc. Additionally, the multimedia interactive description information may, for example, include multimedia information such as picture information, animation information, short video information, etc. Among them, the presentation form of the multimedia information may be a link.
[0128] As an example, as described above, the structured interceptor in the dynamic parsing engine can determine the target paragraph based on a preset delimiter and initiate a rich media structured request corresponding to the target paragraph to the large model module. After receiving the rich media structured request, the large model module can input the target paragraph into the artificial intelligence large model. The artificial intelligence large model can obtain the interactive description information corresponding to the target paragraph and send the obtained interactive description information to the dialogue management module through the main interactive module. Among them, the interactive description information can also include a scene identifier, a paragraph identifier, and which paragraph the target paragraph is in the entire response text. After receiving the interactive description information, the dialogue management module can send the interactive description information including the scene identifier to the service management module. After receiving the scene identifier, the service management module can route the interactive description information to the target service assistant corresponding to the scene identifier based on the correspondence between the scene identifier and the target service assistant. After receiving the interactive description information, the target service assistant can perform subsequent structured processing through the rich media structured engine in the target service assistant.
[0129] Based on this, in some embodiments, the above S140 may specifically include:
[0130] Input the target paragraph into the artificial intelligence large model to determine the text interactive description information corresponding to the display template information through the artificial intelligence large model, extract the text interactive description information from the target paragraph, and search for the multimedia interactive description information corresponding to the text interactive description information;
[0131] Determine the interactive description information and the multimedia description information as card information.
[0132] Here, since the response data may include display template information, therefore, the target paragraph determined based on the response data may also include display template information. After the large model module inputs the target paragraph into the artificial intelligence large model, the artificial intelligence large model can first identify the display template information and determine the text interactive description information corresponding to the display template information. For example, if the display template information represents that the template type is a description class template, the text interactive description information may include information such as a title, a subtitle, and preset keywords. If the display template information represents that the template type is a comparison class template, the text interactive description information may include information such as comparison objects, comparison items, and comparison contents.
[0133] After the artificial intelligence large model determines the text interactive description information, it can extract the text interactive description information from the target paragraph and then search for multimedia interactive description information such as picture information, animation information, and short video information corresponding to the text interactive description information. Among them, the multimedia interactive description information may or may not be searched, which is not limited here.
[0134] In some embodiments, in S150, the card information may be card information in rich media form. The card information in rich media form may include audio, text, links, interactive controls, pictures, etc. corresponding to the target paragraph. In addition, the card information may further include interactive input data and response data after structured processing. For example, a schematic diagram of a kind of card information provided by an embodiment of the present application may be as Figure 3 shown.
[0135] As an example, after the artificial intelligence large model returns the interactive description information, the rich media structuring engine may pull up a hypertext markup language interface (such as HTML 5, H5), and transmit the interactive description information to H5 through the way of JS bridge. The client can load H5 and the interactive description information through a browser (such as a webview container), and generate and display the card information.
[0136] In addition, the response data may further include network status. The network status may include normal network and abnormal network. Based on this, in some embodiments, when the above-mentioned number of words is greater than the word count threshold, determining whether the response data includes a preset delimiter may specifically include:
[0137] When the number of words is greater than the preset threshold, intercept the network status in the response data;
[0138] When the network status is normal network, determine whether the response data includes a preset delimiter.
[0139] Based on this, in order to further improve the user experience, in some embodiments, when the number of words is greater than the word count threshold and before determining whether the response data includes a preset delimiter, it may further include:
[0140] Intercept the network status in the response data;
[0141] When the network status is abnormal network, display a target control, and the target control is used to regenerate the card information.
[0142] Here, the dynamic parsing module may further include a network exception interceptor. The network exception interceptor can intercept the network status in the response data. If the network status is abnormal network, the target control for regenerating the card information may be displayed on the voice interaction interface. After the user sees the target control, it can be determined that the network is abnormal, and the user can choose whether to regenerate the card information.
[0143] In this way, by providing an interactive method for the user to regenerate and display the card information when the card information display fails, the user experience can be further improved.
[0144] In addition, as described above, the display template information may be template type information corresponding to the display template of the response data. Among them, the template type may include a description template, a general classification template, a comparison template, a timeline template, and a service expert template. In addition, different template types may correspond to different card types. Among them, the card type may include a description card, a general classification card, and a service expert card. Among the description card, the general classification card, and the service expert card, the card may further include a card with a picture and a card without a picture. The corresponding relationship between the template type and the card type may be as Figure 4 shown. It should be noted that the relationship between the template type and the card type may also be a one-to-one correspondence, which is not limited herein.
[0145] Based on this, in order to further improve the display effect of the card information and enhance the user's visual experience, in some embodiments, the above S150 may specifically include:
[0146] Filling the valid information into the display template corresponding to the display template information to obtain card information corresponding to the target paragraph, where the valid information includes at least one of text interaction description information and multimedia interaction description information;
[0147] Displaying the card information in the card information display area.
[0148] Here, the card information display area may be a preset display area of the voice interaction interface. In the case where the voice interaction interface includes multiple display areas, a schematic diagram of a voice interaction interface provided by an embodiment of the present application may be as Figure 5 shown. In Figure 5 it, the right side of the voice interaction interface may be the card information display area.
[0149] In addition, the valid information may include at least one of text interaction description information and multimedia interaction description information. For example, the valid information may only include a title, may only include a subtitle, or may include both a title and a picture link. That is, as long as the large model module can return any valid information, the card information corresponding to the valid information can be generated.
[0150] In addition, the dynamic parsing engine may include a template interceptor, which can intercept and record the display template information in the response data. If the dynamic parsing engine can intercept the display template information, the rich media structuring engine can render the interactive description information according to the display template information to generate card information corresponding to the target paragraph. Specifically, the rich media structuring engine can launch a HyperText Markup Language interface (such as HTML 5, H5), and transmit the interactive description information to H5 through the JSbridge method. The client can launch an H5 interface corresponding to the template type through a browser (such as a webview container), load the H5 template and the interactive description information, fill the valid information into the corresponding positions of the H5 template, and generate card information corresponding to the target paragraph.
[0151] In this way, by using templates and data to dynamically generate the user interface (i.e., card information), the structure and style of the user interface can be separated from the data, so that different user interfaces can be dynamically generated according to different data, further improving the display effect of the card information and enhancing the user's visual experience.
[0152] In addition, the response data may further include the source information of the response data. The source information can characterize the way the artificial intelligence large model obtains the response data. As described above, obtaining the response data corresponding to the interactive input data includes searching for the response data corresponding to the interactive input data in a pre-set Q&A library, searching for the response data corresponding to the interactive input data in the network, and generating the response data corresponding to the interactive input data. Based on this, for example, if the response data is the response data corresponding to the interactive input data obtained by the artificial intelligence large model by searching in a pre-set Q&A library corresponding to the vehicle, the source information of the response data can be the Q&A library corresponding to the vehicle.
[0153] As an example, when the artificial intelligence large model obtains the response data, it can synchronously determine the source information corresponding to the response data and mark the source information in the response data.
[0154] Based on this, in order to further enhance the user experience of using the question and answer system, in some embodiments, rendering the interactive description information according to the display template information to generate card information corresponding to the target paragraph may specifically include:
[0155] In the case where the source information of the response data is the Q&A library corresponding to the vehicle, modify the display template information to the service expert template, and the service expert template is the display template corresponding to the vehicle Q&A;
[0156] Render the interactive description information according to the service expert template to generate card information corresponding to the target paragraph.
[0157] Here, the service expert template can be a display template corresponding to vehicle Q&A. The service expert template can include two upper and lower display areas. Among them, the upper display area can display multimedia interaction description information, and the lower display area can display text interaction description information. The text interaction description information in the service expert template can be operation steps related to the vehicle's usage service. The card information generated based on the service expert template can include the following content, for example: "The usage steps of the child safety lock are as follows: 1. Click 'Vehicle' in the central control screen settings, select 'Vehicle Lock', select the option below the 'Child Lock Button' to set the child lock button; 2. After selecting 'Left' or 'Right', when the child safety lock is turned on, the corresponding rear door on that side cannot be opened from inside the vehicle and the window cannot be controlled."
[0158] In this way, when the source information of the response data is the Q&A library corresponding to the vehicle, by rendering the interaction description information according to the service expert template corresponding to the vehicle Q&A to generate card information, the professionalism of the card information can be improved, and further enhance the user experience of using the Q&A model.
[0159] In addition, as described above, there can be multiple target paragraphs. Based on this, in order to ensure the timeliness of card information display and further enhance the user's visual experience, in some embodiments, the above-mentioned display of card information in the card information display area can specifically include:
[0160] Update the displayed card information based on multiple target paragraphs according to the generation order of the multiple target paragraphs;
[0161] Display the updated card information in the card information display area until the end identifier is included in the target paragraph to obtain the complete card information.
[0162] Here, each time a target paragraph is determined, the card information corresponding to the target paragraph can be generated and displayed once. During the process of displaying the card information corresponding to the first target paragraph, the dynamic parsing engine can synchronously determine the second target paragraph from multiple response data streamed out after the first target paragraph. After determining the second target paragraph, the card information being displayed can be updated. For example, as Figure 5 shown, the card information display area can update the chart. In addition, a prompt message "Generating chart..." can be displayed on the left side of the card information display area to prompt the user that the card information has not been fully displayed.
[0163] If the response data includes an end flag, it can be determined that the response data is a complete target response data. Therefore, the target paragraph determined based on the response data may also include an end flag. If the target paragraph includes an end flag, it can be determined that the target paragraph is the last target paragraph. Furthermore, after generating and displaying the card information corresponding to the last target paragraph, the update of the card information can be completed, and the complete card information can be obtained.
[0164] In this way, by displaying a part of the card information every time a target paragraph is obtained, the timeliness of the card information display can be ensured, and the user's visual experience can be further improved.
[0165] Based on the Q&A method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of the Q&A device. Please refer to the following embodiments.
[0166] As Figure 6 shown, the Q&A device 600 provided in the embodiments of the present application includes the following modules:
[0167] A receiving module 610, configured to receive the user's interaction input data, where the interaction input data includes multimodal feature data and instruction interaction data, the multimodal feature data represents the input form of the interaction input data, and the instruction interaction data represents the user intention corresponding to the interaction input data;
[0168] A first input module 620, configured to input the interaction input data into an artificial intelligence large model, so as to identify and respond to the interaction input data through the artificial intelligence large model, and obtain and output multiple response data;
[0169] A determining module 630, configured to determine a target paragraph based on the multiple output response data during the process of the artificial intelligence large model outputting multiple response data;
[0170] A second input module 640, configured to input the target paragraph into the artificial intelligence large model, so as to obtain card information corresponding to the target paragraph through the artificial intelligence large model;
[0171] A rendering module 650, configured to render and display the card information corresponding to the target paragraph.
[0172] The above Q&A device 600 will be described in detail below, as follows:
[0173] In some of the embodiments, the first input module 620 may specifically include:
[0174] The first input sub-module is used to input the interactive input data into the artificial intelligence large model, so as to identify the interactive input data through the artificial intelligence large model, obtain the reply type corresponding to the interactive input data, and determine the display template information corresponding to the reply type; and reply to the interactive input data through the artificial intelligence large model, and obtain and output a plurality of reply data including display template information.
[0175] In some embodiments, the target paragraph includes display template information. Based on this, the second input module 640 may specifically include:
[0176] The second input sub-module is used to input the target paragraph into the artificial intelligence large model, so as to determine the text interaction description information corresponding to the display template information through the artificial intelligence large model, extract the text interaction description information from the target paragraph, and search for the multimedia interaction description information corresponding to the text interaction description information;
[0177] The determination sub-module is used to determine the text interaction description information and the multimedia interaction description information as card information.
[0178] In some embodiments, the rendering module 650 may specifically include:
[0179] The filling sub-module is used to fill the valid information into the display template corresponding to the display template information to obtain the card information corresponding to the target paragraph, and the valid information includes at least one of the text interaction description information and the multimedia interaction description information;
[0180] The display sub-module is used to display the card information in the card information display area.
[0181] In some embodiments, there are multiple target paragraphs. Based on this, the display sub-module may specifically include:
[0182] The update unit is used to update the displayed card information based on the multiple target paragraphs according to the generation order of the multiple target paragraphs;
[0183] The display unit is used to display the updated card information in the card information display area until the end identifier is included in the target paragraph to obtain the complete card information.
[0184] In some embodiments, among the multiple reply data output, the latter reply data includes the former reply data. Based on this, the determination module 630 may specifically include:
[0185] The first judgment sub-module is used to judge whether a preset delimiter is included in each reply data among the multiple reply data output;
[0186] The first determination sub-module is configured to determine the response data as the target paragraph when the response data includes a preset delimiter.
[0187] The second determination sub-module is configured to determine the response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph when the response data includes multiple preset delimiters.
[0188] In some embodiments, the response data includes a scenario identifier. Based on this, the determination module 630 may specifically further include:
[0189] The second judgment sub-module is configured to judge whether the user interaction scenario corresponding to the scenario identifier is in a structured whitelist including multiple user interaction scenarios before judging whether the response data includes a preset delimiter.
[0190] Based on this, the first judgment sub-module may specifically include:
[0191] The judgment unit is configured to judge whether the response data includes a preset delimiter when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist, and the structured whitelist includes multiple user interaction scenarios.
[0192] In some embodiments, the response data includes display template information and a word count threshold. Based on this, the judgment unit may specifically include:
[0193] The recording sub-unit is configured to intercept and record the display template information and the word count threshold in the response data when the user interaction scenario corresponding to the scenario identifier is in the structured whitelist.
[0194] The obtaining sub-unit is configured to obtain the word count in the response data.
[0195] The first judgment sub-unit is configured to judge the magnitude relationship between the word count and the word count threshold.
[0196] The second judgment sub-unit is configured to judge whether the response data includes a preset delimiter when the word count is greater than the word count threshold.
[0197] In some embodiments, the judgment unit may specifically further include:
[0198] The interception sub-unit is configured to intercept the network status in the response data when the word count is greater than the word count threshold and before judging whether the response data includes a preset delimiter.
[0199] The display sub-unit is configured to display a target control for regenerating card information when the network status is network exception.
[0200] In the process of the Q&A device according to the embodiment of the present application outputting multiple response data, the target paragraph is determined based on the multiple output response data, and then the target paragraph is input into the artificial intelligence large model, and the artificial intelligence large model is used to obtain the card information corresponding to the target paragraph, so as to determine the key display information corresponding to the target paragraph. In this way, by rendering and displaying the card information corresponding to the target paragraph, the key display information can be displayed. Compared with displaying a large paragraph of pure text and pictures, it can reduce the reading burden of the user on the response data, enable the user to clearly and accurately obtain the desired information, and thus improve the user's reading experience of the response data.
[0201] Based on the Q&A method provided in the above embodiment, the embodiment of the present application also provides a specific implementation manner of an electronic device. Figure 7 The schematic diagram of the electronic device 700 provided by the embodiment of the present application is shown.
[0202] The electronic device 700 may include a processor 710 and a memory 720 storing computer program instructions.
[0203] Specifically, the above-mentioned processor 710 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0204] The memory 720 may include a mass storage for data or instructions. By way of example and not limitation, the memory 720 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 720 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 720 may be inside or outside the electronic device 700. In a specific embodiment, the memory 720 is a non-volatile solid state memory.
[0205] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0206] The processor 710 implements any one of the Q&A methods in the above embodiments by reading and executing computer program instructions stored in the memory 720.
[0207] In one example, the electronic device 700 may further include a communication interface 730 and a bus 740. Among them, as Figure 7 shown, the processor 710, the memory 720, and the communication interface 730 are connected through the bus 740 and complete communication with each other.
[0208] The communication interface 730 is mainly used to implement communication between each module, device, unit, and / or device in the embodiments of the present application.
[0209] The bus 740 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 740 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0210] Exemplarily, the electronic device 700 may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.
[0211] The electronic device can execute the Q&A method in the embodiments of the present application, thereby implementing the Q&A method and device combined with Figure 1 to and Figure 6 described Q&A methods and devices.
[0212] In addition, in combination with the Q&A method in the above embodiments, the embodiments of the present application may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the Q&A methods in the above embodiments is implemented.
[0213] In addition, the embodiments of the present application further provide a vehicle, which may include at least one of the following:
[0214] a question-and-answer device as in any of the embodiments of the second aspect;
[0215] an electronic device as in any of the embodiments of the third aspect;
[0216] a computer-readable storage medium as in any of the embodiments of the fourth aspect. Details are not described herein again.
[0217] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0218] The functional blocks shown in the above structure block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0219] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0220] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the present application. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine, such that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0221] As described above, the above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present application.
Claims
1. A question-and-answer method, characterized in that, including: Receiving interactive input data from a user, where the interactive input data includes multimodal feature data and instruction interaction data. The multimodal feature data characterizes the input form of the interactive input data, and the instruction interaction data characterizes the user intention corresponding to the interactive input data; Inputting the interactive input data into an artificial intelligence large model to identify and respond to the interactive input data through the artificial intelligence large model, and obtaining and outputting multiple response data; During the process of the artificial intelligence large model outputting the multiple response data, determining a target paragraph based on the multiple output response data; Inputting the target paragraph into the artificial intelligence large model to obtain card information corresponding to the target paragraph through the artificial intelligence large model; Rendering and displaying the card information corresponding to the target paragraph.
2. The method according to claim 1, characterized in that, The step of inputting the interactive input data into an artificial intelligence large model to identify and respond to the interactive input data through the artificial intelligence large model, and obtaining and outputting multiple response data includes: Inputting the interactive input data into the artificial intelligence large model to identify the interactive input data through the artificial intelligence large model, obtaining a response type corresponding to the interactive input data, and determining display template information corresponding to the response type; and responding to the interactive input data through the artificial intelligence large model to obtain and output multiple response data including the display template information.
3. The method according to claim 1, characterized in that, The target paragraph includes display template information. The step of inputting the target paragraph into the artificial intelligence large model to obtain card information corresponding to the target paragraph through the artificial intelligence large model includes: Inputting the target paragraph into the artificial intelligence large model to determine text interaction description information corresponding to the display template information through the artificial intelligence large model, extracting the text interaction description information from the target paragraph, and searching for multimedia interaction description information corresponding to the text interaction description information; Determining the text interaction description information and the multimedia interaction description information as the card information.
4. The method according to claim 3, characterized in that, The step of rendering and displaying the card information corresponding to the target paragraph includes: Filling valid information into a display template corresponding to the display template information to obtain card information corresponding to the target paragraph, where the valid information includes at least one of the text interaction description information and the multimedia interaction description information; Displaying the card information in a card information display area.
5. The method according to claim 4, characterized in that, When there are multiple target paragraphs, the step of displaying the card information in the card information display area includes: Updating the displayed card information based on the multiple target paragraphs in the generation order of the multiple target paragraphs; Displaying the updated card information in the card information display area until an end identifier is included in the target paragraph to obtain the complete card information.
6. The method according to claim 1, characterized in that, Among the multiple output response data, the latter response data includes the former response data. The step of determining a target paragraph based on the multiple output response data includes: Among the multiple pieces of the response data output, for each piece of the response data, determine whether the response data includes a preset delimiter; When the response data includes one of the preset delimiters, determine the response data as the target paragraph; When the response data includes multiple preset delimiters, determine the response data between the last two adjacent preset delimiters among the multiple preset delimiters as the target paragraph.
7. The method according to claim 6, characterized in that, The response data includes a scene identifier. Before determining whether the response data includes a preset delimiter, the method further includes: Determine whether the user interaction scene corresponding to the scene identifier is in a structured whitelist, where the structured whitelist includes multiple user interaction scenes; The determination of whether the response data includes a preset delimiter includes: When the user interaction scene corresponding to the scene identifier is in the structured whitelist, determine whether the response data includes a preset delimiter.
8. The method according to claim 7, characterized in that, The response data includes display template information and a word count threshold. When the user interaction scene corresponding to the scene identifier is in the structured whitelist, the determination of whether the response data includes a preset delimiter includes: When the user interaction scene corresponding to the scene identifier is in the structured whitelist, intercept and record the display template information and the word count threshold in the response data; Obtain the word count in the response data; Determine the size relationship between the word count and the word count threshold; When the word count is greater than the word count threshold, determine whether the response data includes a preset delimiter.
9. The method according to claim 8, characterized in that, When the word count is greater than the word count threshold and before determining whether the response data includes a preset delimiter, the method further includes: Intercept the network status in the response data; When the network status is a network anomaly, display a target control for regenerating the card information.
10. A question-and-answer device, characterized in that, The device includes: A receiving module, configured to receive the user's interactive input data, where the interactive input data includes multimodal feature data and instruction interaction data, the multimodal feature data represents the input form of the interactive input data, and the instruction interaction data represents the user intention corresponding to the interactive input data; A first input module, configured to input the interactive input data into an artificial intelligence large model to identify and respond to the interactive input data through the artificial intelligence large model, and obtain and output multiple pieces of response data; A determination module, configured to determine a target paragraph based on the multiple pieces of response data output during the process of the artificial intelligence large model outputting the multiple pieces of response data; A second input module, configured to input the target paragraph into the artificial intelligence large model to obtain card information corresponding to the target paragraph through the artificial intelligence large model; A rendering module, configured to render and display the card information corresponding to the target paragraph.
11. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the question-answering method described in any one of claims 1-9 is implemented.
12. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by the processor, the question-answering method described in any one of claims 1-9 is implemented.
13. A vehicle, characterized in that, Including at least one of the following: The question-and-answer device according to claim 10; The electronic device according to claim 11; The computer-readable storage medium according to claim 12.
Citation Information
Cited By
Interaction method and device, electronic equipment and computer readable storage medium
CN121116146A
Message card processing method, device and equipment based on artificial intelligence model
CN121833785A