Interaction method, interaction device, electronic equipment and computer readable storage medium
By introducing guided information and user interaction into the AI question-telling system, simulating the interaction of human teachers, and adjusting the explanation strategy in real time, the existing AI question-telling problem is solved, and a personalized and efficient learning experience is achieved.
Patent Information
- Application Number
- CN202510534077.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
The existing AI question-teaching methods are usually an overall explanation from scratch, which makes students unable to directly advance to problematic areas and is inefficient.
It provides an interactive method, by receiving questions input by users and displaying guidance information on the interactive interface, guides users to determine the method of teaching questions, generates and outputs content, interacts based on user replies, simulates the emotional interaction and heuristic teaching of human teachers, and adjusts explanation strategies in real time.
It improves the user's learning experience, stimulates in-depth thinking through personalized explanation strategies, and improves the efficiency of the topic and the user's understanding effect.
Smart Images

Figure CN120412355A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to an interaction method, an interaction device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the development of artificial intelligence (AI) technology and large language models, many software applications that utilize AI technology for intelligent interaction have emerged. Common applications of AI interaction in the field of education include AI problem-solving lectures. Compared with the situation of live lectures by human teachers, general AI problem-solving lecture products mainly adopt an overall lecture method starting from the beginning, with the AI as the "lecture teacher" asking questions and the user as the "problem-solving student" answering. This lecture method may cause students to be unable to directly jump to the problematic part and quickly obtain answers, resulting in a reduction in the overall lecture efficiency. Summary of the Invention
[0003] According to some embodiments of the present disclosure, there is provided an interaction method, including: receiving a question input by a user in an interaction interface of an agent; displaying guiding information on the interaction interface, where the guiding information is used to guide the user to determine a lecture method; generating and outputting first content on the interaction interface based on the question and the lecture method; receiving a first reply of the user to the first content; and generating and outputting second content on the interaction interface based on the question and the first reply.
[0004] According to some other embodiments of the present disclosure, there is provided an interaction device, including: a receiving module configured to receive a question input by a user in an interaction interface of an agent; a guiding module configured to display guiding information on the interaction interface, where the guiding information is used to guide the user to determine a lecture method; a first output module configured to generate and output first content on the interaction interface based on the question and the lecture method; a first receiving module configured to receive a first reply of the user to the first content; and a second output module configured to generate and output second content on the interaction interface based on the question and the first reply.
[0005] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, where the processor is configured to execute the method according to any one of the embodiments of the present disclosure based on instructions stored in the memory.
[0006] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it executes the method according to any one of the embodiments of the present disclosure.
[0007] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Description of the Drawings
[0008] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation on the present disclosure. In the drawings:
[0009] Figure 1 A flowchart of an interaction method for an agent to explain a problem is shown according to some embodiments of the present disclosure;
[0010] Figures 2A to 2F A schematic diagram of an interaction interface between a user and an agent is shown according to some embodiments of the present disclosure;
[0011] Figures 3A to 3C A schematic diagram of an agent generating problem-explanation content is shown according to some embodiments of the present disclosure;
[0012] Figures 4A to 4D A schematic diagram of a method for training an agent is shown according to some embodiments of the present disclosure;
[0013] Figure 5 A schematic block diagram of an interaction device for an agent to explain a problem is shown according to some embodiments of the present disclosure;
[0014] Figure 6 A block diagram of an electronic device is shown according to some embodiments of the present disclosure;
[0015] Figure 7 A block diagram of an electronic device is shown according to other embodiments of the present disclosure.
[0016] It should be understood that, for the sake of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to represent the same or similar components. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. Detailed Embodiments
[0017] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0018] It should be understood that the various steps described in the method embodiments of the present disclosure may be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and not limiting the scope of the present disclosure.
[0019] The term "comprising" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features, but do not exclude other elements / features, that is, "including but not limited to". The term "based on" means "at least partially based on".
[0020] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to define the order of functions executed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts such as "first", "second", etc. are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0024] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined by those of ordinary skill in the art in any suitable manner that will be apparent from the present disclosure.
[0025] It should be understood that the present disclosure also places no restrictions on how to obtain the image to be applied / processed. In some embodiments of the present disclosure, it can be obtained from a storage device, such as an internal memory or an external storage device. In other embodiments of the present disclosure, the photography component can be mobilized to take pictures. It should be noted that the obtained image can be a single captured image or a frame of the captured video, and is not particularly limited thereto.
[0026] In the context of the present disclosure, the image can refer to any one of a variety of images, such as a color image, a grayscale image, etc. It should be pointed out that in the context of this specification, the type of the image is not specifically limited. In addition, the image can be any appropriate image, such as the original image obtained by a camera device, or the image that has been subjected to specific processing on the original image, such as preliminary filtering, anti-aliasing, color adjustment, contrast adjustment, normalization, etc. It should be pointed out that the preprocessing operation can also include other types of preprocessing operations known in the art, which will not be described in detail here.
[0027] In the related art, an intelligent agent in the field of education can provide a function of explaining questions. With the help of technologies such as natural language processing and knowledge graph construction, it can identify, understand, and generate explanatory content for test questions. With the significant development of AI question-explaining software in recent years, it can cover a large number of question banks in multiple disciplines and all school stages, and can also provide personalized learning path planning according to students' answering situations. When the intelligent agent combines the Vision Language Model (VLM) and / or the Large Language Model (LLM) technology, it can provide an interactive question-explaining experience for users and improve the intelligent agent's knowledge understanding and integration ability in explaining complex test questions. For this reason, the present disclosure proposes an interaction method, in which the emotional interaction and heuristic teaching of a human teacher can be simulated, and the explanation strategy can be adjusted in real time according to the user's question-explaining needs, knowledge level, understanding degree, etc., so as to stimulate the user's in-depth thinking and improve the user's usage experience.
[0028] Specifically, Figure 1 shows a flowchart of an interaction method for an intelligent agent to explain questions according to some embodiments of the present disclosure. As Figure 1 shown, in step S101, the question input by the user in the interaction interface of the intelligent agent is received; in step S102, guiding information is displayed on the interaction interface, and the guiding information is used to guide the user to determine the question-explaining method; in step S103, based on the question and the question-explaining method, the first content is generated and output on the interaction interface; in step S104, the first reply of the user to the first content is received; in step S105, based on the question and the first reply, the second content is generated and output on the interaction interface.
[0029] The recommendation method of this embodiment can be executed on the client side or partially on the server side.
[0030] Generally, the questions in this disclosure are input by the user on the interaction interface, and their forms can be text, multimedia resources, or a combination of both. The user can manually input the question text, or use the voice input function of the interaction interface to dictate the question content, or take a picture of the question and upload the image or video of the question to the interaction interface for subsequent image recognition by the intelligent agent. The ways of explaining questions mainly include overall explanation, key point explanation, and thinking guidance summarized from the teaching process of teachers to students in reality, and the process planning of the explanation is designed around the actual needs of the user's questions. It can be understood that in the process of the intelligent agent imitating a human teacher to explain, the first content, the second content, etc. displayed on the interaction interface, as well as other content output during the interaction, can have rich forms of expression. In addition to conventional text and voice, it can also include but is not limited to static pictures, dynamic images, short videos, slides, virtual reality / augmented reality content, social media content, animated films, interactive elements, etc., and can be set to one or a combination of the above forms according to the questions input by the user and the detailed information of their answers.
[0031] The interaction method provided by this disclosure is generally applied to intelligent agent interaction software with the function of explaining questions. Specifically, reference can be made to Figures 2A to 2F which shows a schematic diagram of the interaction interface between the user and the intelligent agent according to some embodiments of this disclosure. Figure 2A It shows the interaction interface of the question-explaining software in the related art, which generally includes the question 1 input by the user and the answer content 2 output by the intelligent agent. Generally, the user inputs the question 1 to be explained into the question-explaining software, and the question-explaining software recognizes and understands the content of the question 1, thereby generating the answer content 2 corresponding to the question 1 and presenting it completely on the software interface. In a non-limiting embodiment, if the question input by the user contains three questions, the answer content 2 can respectively display the detailed analysis and guidance of "the first step", "the second step", and "the third step" in the order of the questions.
[0032] Correspondingly, Figure 2BAn example of an agent interface for an interaction method according to the present disclosure is shown. In response to receiving problem 1, the agent first displays guiding information 3 in the interaction interface, and the guiding information 3 is used to guide the user to determine the problem-solving method. In a non-limiting embodiment, the user may have different levels of mastery of the input problem 1. For example, the user may have no idea at all and need a complete explanation from the beginning; or the user may only have some parts that are not understood and need to explain these parts specifically; or the user is not confident enough about the problem-solving idea and only needs some appropriate hints to solve the problem by themselves. That is to say, the problem-solving methods required by the user may include, but are not limited to, complete explanation, key-point explanation, and idea hint, etc.
[0033] In some embodiments, if the user determines the problem-solving method according to the guiding information, the agent generates and outputs the first content based on the specific content of problem 1 and the determined problem-solving method. As Figure 2C shown, in response to the user determining that the problem-solving method is key-point explanation, the first machine learning model is used to perform semantic analysis on problem 1 to obtain the key content of problem 1 as the first content 4. Further, when the question type of problem 1 is different, the obtained key content has different presentation forms accordingly. In a non-limiting embodiment, if the problem 1 input by the user is a multiple-choice question, the semantic analysis of problem 1 should include the identification and knowledge point summary of the stem content, and the analysis and knowledge point summary of each option. For example, when problem 1 includes four options, the key content includes at least five parts. Alternatively, in another embodiment, if the problem 1 input by the user is a solution question with multiple sub-questions / steps, the semantic analysis of problem 1 includes the analysis and knowledge point summary of each sub-question / step. For example, when problem 1 includes three sub-questions, the key content includes at least three parts.
[0034] In some embodiments, in response to the user selecting the problem-solving method as key-point explanation, the problem points selected by the user from the key content are received as the first reply 41. Figure 2C The situation where three main problem-solving steps are obtained after semantic analysis of problem 1 is shown. Then the first content 4 is the three-step key content generated and output. At this time, the first reply 41 of the user to the first content 4 is to select the problem point "the second step" of the key content. It can be understood that the first reply 41 of the user may include input operations such as selecting the interaction options given by the first content 4 in the interface, or input contents such as voice or text in response to the first content 4. Further, as Figure 2DAs shown, the agent generates and outputs a second content 5 on the interaction interface based on the problem 1 and the first response 41. Specifically, since the user's first response 41 selects the problem point "the second step", the agent accordingly explains the second problem-solving step of the problem 1. In a non-limiting embodiment, the agent can simulate a real teaching process in the second content 5 to conduct a problem-guided explanation, that is, prompt the user to think about the corresponding part of the problem 1. For example, in the second content 5, it asks: "The second step involves the variable a. What is the corresponding formula A?" to teach the user to autonomously analyze the "second step" from the knowledge points associated with the formula A.
[0035] In some embodiments, in response to the second content including a question, receive a second response from the user to the question; and based on the second content and the second response, generate and output a third content on the interaction interface. As mentioned above, Figure 2D the second content 5 in is in the form of a question. The user can answer the question by means of voice or text input. For example, answer the specific content of the formula A as the second response. At this time, the agent can judge whether the user understands the knowledge points expected to be explained in the second content 5 based on the user's answer, so as to generate the third content for the next explanation. Refer to Figure 2E , based on the user answering the formula A, the agent can continue to ask "What is the relationship between the variable a and the variable b according to the formula A?" to guide the user to gradually understand the problem point of the "second step" of the problem 1.
[0036] Specifically, in some embodiments, in response to the time from outputting the second content to receiving the second reply exceeding a specified time interval, or the second reply containing specified negative semantics, determine the teaching bottleneck in the second content; and generate a reference hint for the teaching bottleneck as the third content. When the user doesn't know how to answer the question about the second content output by the agent, it indicates that the current question of the agent is a difficult point for the user's knowledge reserve, that is, there is a bottleneck in the teaching process that is difficult to advance. In a non-limiting embodiment, the user takes a long time to input the second reply after receiving the second content of the agent. For example, it takes a long time to recall or think to answer the question, indicating unfamiliarity or uncertainty about the knowledge point involved in the question. Thus, a specified time interval for the user's thinking time is set, and the agent calculates the time from outputting the second content to receiving the user's second reply. If this time exceeds the specified time interval, it can be determined that the second content encounters a teaching bottleneck at this question. In another non-limiting embodiment, the user may directly input "I don't know this formula" or "I don't understand your question" as the second reply. The agent's model can set specified negative semantics during training, such as the aforementioned "don't know", "don't understand", etc., to identify the reply indicating that the user encounters a teaching bottleneck. Alternatively, the specified negative semantics may also include that the second reply is a wrong answer to the question in the second content, etc.
[0037] Further, please refer to Figure 2F , which shows the reference hint 52 generated by the agent. In a non-limiting embodiment, when the agent determines the teaching bottleneck in the second content, the icon of the reference hint 52 can be displayed in the interaction interface. The icon of the reference hint 52 can be located at any position in the interaction interface, including near the user input box, or specific positions such as the top bar, bottom bar, and side bar of the interface. The icon of the reference hint 52 can not be displayed during the normal interaction process and be displayed in the interface after the agent determines the teaching bottleneck; or it can be displayed as one appearance during the normal interaction process and as another appearance after the agent determines the teaching bottleneck, and the latter has a more prominent display effect. As Figure 2F shown, the reference hint 52 may include a direct answer to the question in the second content (such as "the correct content of formula A"), and / or a detailed introduction to the knowledge point involved in the question, etc. Similar to the second reply, the user can perform an input operation on the reference hint 52 to obtain help in answering the question and promote the progress of the teaching.
[0038] Alternatively, in some embodiments, in response to the second reply having a correct correspondence with the question, a solution to the question is generated based on the second content and the second reply as the third content. When the user correctly answers the question in the second content 5, that is, the user has mastered the knowledge point involved in the question, the agent can continue to output the subsequent content of the explanation process. The agent simulates the "blackboard writing" in the real teaching process and can combine the previous question as the second content 5 and the user's second reply to generate a written answer to the question, which is displayed as the third content in the interaction interface. In a non-limiting embodiment, referring to Figure 2E , as described above, the question is about the formula A corresponding to the variable a. When the user correctly answers the content of the formula A in the second reply, the generated third content 51 may include a declarative written expression of "the formula A of the variable a". Further, the third content 51 also includes the next question of the agent to continue the explanation process. As Figure 2E shown, after multiple rounds of question-and-answer interactions between the user and the agent, multiple blackboard writing contents can be gradually formed, and the key points can be recorded during the process of guiding the user to think about the solution to the problem points for the user to understand and review.
[0039] Alternatively, in some embodiments, in response to the second reply containing the specified end semantics, the interaction between the user and the agent in solving the problem is summarized as the third content. When the user gets the inspiration for solving the problem after receiving the question in the second content 5 and can already complete the answer independently, there is no need for the agent to gradually guide and prompt the subsequent explanation process, and the current interaction in solving the problem can be ended. In a non-limiting embodiment, the user can express the semantics of wanting to end the interaction in the second reply, such as "I already understand here" or "I don't need the explanation", etc. The model of the agent can be set with the specified end semantics during training, such as the aforementioned "no need for explanation", etc., to identify the reply of the user wanting to end the interaction process of solving the problem. Alternatively, the specified negative semantics may also include the user expressing general closing remarks such as "close the conversation" or "stop" in the second reply.
[0040] Alternatively, in some embodiments, in response to the second reply being irrelevant to the question, a prompt message is displayed on the interaction interface, which is used to prompt the user to determine whether to continue the question-explanation interaction. When the user receives the second content 5 and replies with content that has no association with the question, it indicates that the user's thinking may have deviated, or there may be a situation of inattentiveness. In a non-limiting embodiment, the second reply of the user neither contains the aforementioned specified negative semantics or specified ending semantics, nor is it a correct or incorrect answer to the question, but other irrelevant content. Then the agent outputs the prompt message "Your answer is irrelevant to this question. Do you want to continue the explanation or switch to discussing other content?" The agent can determine whether to continue the previous question-explanation interaction process based on the user's answer to this prompt message. For example, if the user wishes to continue the explanation, the agent can repeat the question of the second content 5 to guide the user back to the original explanation interaction; if the user wishes to switch to discussing the new content that is "irrelevant to the question" in the second reply, the agent can create a new interaction and prompt the user to explain about this new content.
[0041] Return reference Figure 2B , in some embodiments, in response to the question-explanation method selected by the user being idea guidance, a second machine learning model is used to generate a mind map of question 1, which is output to the interaction interface as the first content. Generally, a mind map is a knowledge association network constructed by visual elements such as graphics and keywords in a hierarchical and divergent structure. The agent uses the mind map for idea guidance in the explanation, including but not limited to quickly sorting out the question logic and knowledge point context, systematically disassembling the problem, and intuitively presenting the problem-solving idea, etc. Further, in some embodiments, in response to the question-explanation method selected by the user being idea guidance, the interaction between the user and the mind map is received as the first reply. In a non-limiting embodiment, the mind map output by the agent can include a display form that expands or folds hierarchically. The user can conduct personalized learning according to their own knowledge system. For example, the user can selectively fold the parts that they are not interested in and focus on the key content that the user needs to understand, thereby enhancing the depth of understanding and memory effect of the question. The interaction between the user and the mind map includes the user's selection operation and / or questions raised by the user regarding the mind map, etc., providing personalized explanation ideas and assistance for the user.
[0042] Next, please refer to Figures 3A to 3C , which shows a schematic diagram of the agent generating question-explanation content according to some embodiments of the present disclosure. First, as Figure 3AAs shown, the user selects a question and inputs it into the interaction interface in step S3, and the agent makes a speakable determination of the question based on the attribute information (step S32). In a non-limiting embodiment, the attribute information includes, but is not limited to, basic attributes, explanation attributes, question type attributes, etc. The basic attributes of the question such as whether the question contains a graph, whether it exceeds the specified number of words, etc., the explanation attributes include the subject and academic stage corresponding to the question, etc., and the question type attributes include the specific examination form of the question, whether it belongs to the type of questions in the question bank, etc. The speakable determination of the question by the agent such as includes: for a text-based agent, a question containing a graph is not speakable; a question exceeding the specified number of words of the semantic recognition upper limit of the agent is not speakable; for an agent exclusive to a subject or academic stage, a question with a mismatched explanation attribute is also not speakable. If the question fails to pass the speakable determination, the agent can output a prompt such as unable to explain, indicating that the user fails to enter the subsequent interaction process. Correspondingly, for a question that passes the speakable determination by the agent, a question-explanation entry is displayed in the interaction interface (step S33), and the question-explanation interaction between the user and the agent starts from this.
[0043] Further, referring to Figure 3B , for the question selected by the user in step 31, the agent calls a model in step S34 based on the question and makes a determination of specific question type explanation in combination with the question type attributes of the question (step S35). The models called by the agent in the explanation interaction include, but are not limited to, VLM, LLM, etc. As mentioned above, the specific examination forms in the question type attributes include multiple-choice questions, fill-in-the-blank questions, solution questions, etc., and whether the question belongs to the type of questions in the question bank mainly includes the aggregated features generated by the extraction based on the knowledge point dimension, such as problems of distance and speed, problems of resistance and power in physics, etc. These "type of questions" integrate related knowledge points and have relatively fixed problem-solving ideas that can be summarized and analogized, so they are summarized as specific question types when using the question bank to train the agent. Thus, if it is determined in step S35 that the question to be explained belongs to a specific question type (determined as "yes"), then it enters step S36-1 to match the generation logic of the specific question type explanation plan; if it is determined that the question to be explained does not belong to any specific question type (determined as "no"), then it enters step S36-2 to match the generation logic of the general question type explanation plan.
[0044] Next, in step S37, the agent outputs an explanation plan for the question to be explained based on the explanation plan generation logic matched in step S36. The explanation plan includes, but is not limited to, all the necessary steps to complete the question explanation. For example, for a multiple-choice question with multiple options, in its explanation plan, the agent should at least complete the knowledge point analysis of the question stem and the matching situation analysis of each option, that is, no matter how many rounds of interaction the agent has with the user during the explanation process, the agent needs to output the explanation steps in the plan to the user in sequence. In a non-limiting embodiment, in combination with Figure 2B andFigure 2C As shown, multiple steps of the lecture planning can be listed as key content for the user when the user determines that the lecture method is key-point explanation, so that the user can determine the problem points to be explained in the first response.
[0045] Additionally, in the associated step S37’, a version number is generated for the lecture plan, cached, and incrementally updated. In response to the agent's version upgrade as it continuously trains and learns, the lecture plan generated by the agent for the same question may be updated accordingly. In step S37, it can be determined whether the lecture plan needs to be regenerated based on the version number. When the version number remains unchanged, the generation result cached can be directly reused, thereby reducing the computation amount and lowering the resource invocation of the model.
[0046] Furthermore, in response to the agent outputting the lecture plan, in step S38, the model is called to match the blackboard writing generation logic (step S39), and then in step S310, the conclusion blackboard writing is generated, mainly including outputting the lecture plan as the lecture content displayed on the interaction interface based on the blackboard writing generation logic, such as the step-by-step problem-solving process, etc. Similarly, in the associated step S310’, a version number is generated for the conclusion blackboard writing, cached, and incrementally updated, so that the generation result cached can be directly reused when the version number remains unchanged.
[0047] In some embodiments, the process of the agent's dialogue-based lecture is as Figure 3C shown. Among them, in step S31, the user selects the question to be lectured and inputs it into the interaction interface. The agent asks the user in the first content output what lecture method the user hopes to adopt. The user then selects one of the overall explanation, key-point explanation, and thought circuit as the lecture method in step S311. The agent matches the corresponding "lecture method" engineering link in step S312 according to the user's selection, and calls the model in step S313 to perform differential processing on the input parameters of the model and system presets, etc., to generate the first-round lecture content determined based on the question and the lecture method (step S314). Among them, the first-round lecture content can include a declarative overall analysis of the question to be lectured, that is, generally introducing the scope of knowledge points examined by the question; it can also include the breakdown of the key content in the question, that is, simply listing the steps or key points involved in the lecture plan, etc. In particular, the first-round lecture content can be displayed as text on the interaction interface and / or output to the user in the form of voice. In a non-limiting embodiment, the agent can use text-to-speech (TTS) technology to explain the first-round lecture content, and through the integration of multiple disciplines such as acoustic models, language models, and prosody models, achieve natural and fluent voice output. Similar to steps S37’ and S310’, in the associated step S314’, the first-round lecture content is cached and incrementally updated.
[0048] Additionally or further, in step S315, the agent secondarily invokes the model based on the content of the first-round explanation to output a "next-round question blackboard writing". The blackboard writing is displayed in text form on the interaction interface to guide the user to think about the problem in the form of questions. The user replies to the question in step S316. In particular, return reference Figure 2C , for the case where the user selects the topic-explanation method as key-point explanation in step S311, in some embodiments, based on the attribute information of question 1, a multi-step explanation plan is determined; in response to the first reply 41 including the determined explanation starting point by the user, the starting step for the agent to conduct the explanation is determined among the multiple steps as the second content 5. Specifically, the multiple explanation steps in the explanation plan may correspond to the multiple key points of the first content 4 generated by the agent for key-point explanation. Then, when the user's first reply 41 includes a selection operation for the "second-step" problem point, it is determined that the desired explanation starting point by the user is the step corresponding to "second step" in the explanation plan as the starting step for the explanation. In other words, the first-round explanation content generated by the agent in step S314 may include listing the key points of the question, and the question blackboard writing output in step S315 may include asking the user to determine which step as the explanation starting point, and the user's reply to the question in step S316 may include the selection of the starting step.
[0049] Further, the agent identifies the user's reply intention in step S317. Specifically, the user's reply intention includes but is not limited to normal response, additional explanation, irrelevant questions, etc. In response to identifying the user's reply intention, the agent matches the "intention recognition" engineering link in step S318 and invokes the model in step S319 to generate a new round of explanation content (step S320) based on the question blackboard writing and the user's reply. Similar to the first-round explanation content generated in S314, the new round of explanation content can also use TTS technology to generate explanation voice. At this time, the agent outputs the conclusion blackboard writing of the previous round in step S321-1 and the question blackboard writing of the next round in step S321-2 respectively. That is, Figure 2E the first three shown in belong to the conclusion blackboard writing summarized in the previous rounds, and the fourth is the new question blackboard writing of the next round. After receiving the question returned by the agent in step S321, the user returns to execute step S316, and thus the explanation interaction between the user and the agent enters a multi-step loop and gradually completes the interaction round by round. It can be understood that the new round of explanation content in S320 is also the non-first-round explanation content and belongs to the explanation steps included in the question-explanation plan. Then the judgment condition for the agent to jump out of the above loop can be placed in step S320: in response to all the steps in the question-explanation plan being completed and the agent being unable to generate a new round of explanation content, the conclusion blackboard writing of the previous round is output in step S321-1 and the explanation interaction is prepared to end.
[0050] Specifically, when the user's reply intention is identified as additional explanation in step S317, the agent matches the "intention recognition" engineering link of the additional explanation in step S318. In a non-limiting embodiment, in response to the knowledge points involved in the user's reply (such as Figure 2D the second reply) being outside the lecture plan but associated with the question, a new round of lecture content is generated based on this knowledge point (i.e., Figure 2D the third content in
[0051] ). It should be understood that the lecture plan generated based on the lecture plan generation logic is the recognition and understanding of the question, and its generation result is related to the model called by the agent, and may not fully cover all the knowledge points related to the problem-solving idea of this question; therefore, it is very likely that the user's reply to the question is outside the steps of the lecture plan, such as involving the extension of the current question, or the discrimination and comparison between the current question and approximate knowledge points, that is, the user involves "additional knowledge points" in the reply. At this time, if the agent interacts according to the original lecture plan, it cannot help the user understand and master the additional knowledge points. Therefore, the lecture plan is temporarily not considered, but a new round of lecture content is generated based on this additional knowledge point, and after one or more rounds of interaction with the user for the additional knowledge point to complete the corresponding lecture, it returns to the original lecture plan and continues to execute the subsequent lecture steps. Figure 2E and Figure 2F , if the user's reply has a correct corresponding relationship with the question, the content of the question and the reply is summarized to output the corresponding conclusion blackboard writing; if the user's reply is a wrong answer to the question, the agent comments on the correctness of the reply, analyzes the lecture bottleneck, and / or provides reference hints for the question, etc. Alternatively, when the user's reply intention is identified as an irrelevant question in step S317, the agent matches the intention recognition engineering link of the irrelevant question in step S318. As mentioned above, if the user's reply has no association with the question, the agent determines that the user's thinking has deviated, and asks the user in the next round of question blackboard writing whether they need to continue the lecture, etc.
[0052] Furthermore, the end of the loop that the agent executes from step S316 to step S321 can be judged in step S320. In some embodiments, in response to the completion of the entire lecture plan, the lecture interaction between the user and the agent is summarized to output the last round of conclusion blackboard writing in step S321-1 (refer to Figure 2DThe third content in). Additionally or alternatively, the conclusion blackboard writing at the end of the lecture interaction can include the conclusion blackboard writing corresponding to the last explanation step, and can also include a summary description of all completed rounds. In some other embodiments, in response to the user's reply to the question in step S316 (such as Figure 2D The second reply in) contains the specified end semantics, interrupt the lecture planning, start a new interaction between the user and the intelligent agent, that is, return to Figure 3C The initial step S31 in, reselect the topic, select the lecture method, and regenerate the first-round explanation content (as the third content), and similarly enter the loop of subsequent lecture interactions.
[0053] Next, please refer to Figures 4A to 4D , which shows a schematic diagram of a method for training the retrieval enhanced generation logic of an intelligent agent according to some embodiments of the present disclosure. Generally, Retrieval-Augmented Generation (RAG) is a technology that combines information retrieval with a generation model, aiming to improve the accuracy and richness of content generation in natural language processing tasks. Its core lies in enabling the generation model to dynamically obtain relevant information from an external knowledge base when generating text to assist the text generation process. Taking the AI lecture scenario as an example, when the user asks the intelligent agent about the solution to a certain question, during the process of generating interactive text, the intelligent agent dynamically calls the retrieval system based on the content already generated in the interaction context and / or the blackboard writing already output, etc., and retrieves the most relevant information fragments from a verified knowledge base storing a large amount of data. Incorporating these objective information fragments into the interaction context, the intelligent agent combines the new information with the existing context to continue generating text and iterates multiple times until the complete text is generated. Due to the introduction of external knowledge base information, the final text output by the intelligent agent is often more accurate and rich.
[0054] In some embodiments, when training the model, the intelligent agent in the interaction method of the present disclosure can construct training data for the optimal lecture method of typical question types based on data such as lecture planning, blackboard writing, first-round lecture content, and non-first-round lecture content for different question type attributes and different lecture methods. Further, based on these training data, using the RAG logic of "question type discrimination + lecture planning + dialogue lecture", gradually construct and improve the model's ability to automatically discriminate the optimal lecture planning and dialogue lecture of test questions.
[0055] Specifically, Figure 4AShows the data collected by the agent in generating interactive content, including but not limited to the topic 41 input by the user, the selected topic presentation method 42, and the responses 47 to the agent's first-round lecture content 45 + the blackboard writing of the first-round questions 46 or the blackboard writing of non-first-round questions 410, etc.; the question type discrimination result 43 of the agent, the generated lecture plan 44, the first-round lecture content 45 determined based on the topic 41 and the lecture presentation method 42, the blackboard writing of the first-round questions 46 based on the first-round lecture content, the recognition result 48 of the intention recognition of the response 47, the next-round content of the dialogue lecture generated based on the recognition result, that is, the non-first-round lecture content 49, the blackboard writing of the questions in the same round generated based on the non-first-round lecture content 49, the overall blackboard writing 411 generated based on the lecture plan 44, and the conclusion blackboard writing 412 generated based on the non-first-round lecture content 49 and the overall blackboard writing 411, etc.
[0056] Furthermore, Figure 4B Shows the topic portrait data abstracted from the topic interaction between the agent and the user. Based on the topic 413, determine the topic type attribute 414 of this topic, such as determining whether the topic 413 belongs to a specific topic type or a general topic type, etc., and based on the topic type attribute 414, judge the commonly selected topic presentation method 415 of the user. Generate a lecture plan 415 based on the topic presentation method 415, and then generate the overall conclusion blackboard writing 417 based on the lecture plan 415, and enter the dialogue lecture 418 between the agent and the user. Generate the blackboard writing of the questions 419 based on the content in each round of the dialogue lecture 418, receive the user's response 420 to the blackboard writing of the questions 419, and generate the conclusion blackboard writing 421 of the current round in combination with the blackboard writing of the questions 419 and the user's response. Then return to the process of the dialogue lecture 418 to continue the lecture steps of the next round, and repeat the generation of the data of questions, responses, and conclusions.
[0057] Correspondingly, Figure 4C Shows the RAG training data for constructing the optimal lecture presentation method for typical question types. Refer to Figure 4B Regarding the topic portrait, first determine its topic type attribute 414' based on the typical topic 413', and based on this, judge the commonly used topic presentation method 415' for this topic type. Generate a movie lecture plan 416' based on the topic presentation method 415', and then generate the overall typical conclusion blackboard writing 417'. Enter the typical dialogue lecture 418', generate the typical blackboard writing of the questions 419' in each round, determine the common user responses 420' to this blackboard writing 419', and generate the typical conclusion blackboard writing 421' of the current round; and return to the dialogue lecture 418 to continue promoting the lecture plan. Furthermore, use the constructed optimal RAG to train the agent's lecture model to improve the lecture effect of the model.
[0058] Based on this, please refer to Figure 4D, which shows an application example of the problem-solving RAG model trained based on optimal RAG data. In response to a user inputting problem 401, the agent recommends one or more problem-solving methods 402 that match the attribute information of the problem to the user, and the user selects the desired problem-solving method 403 therefrom. Based on the problem-solving method selected by the user, the agent matches the "explanation plan" type 404 and outputs the first round of the dialogue explanation process, that is, the first-round explanation content 405, in the explanation interaction. The agent further outputs the question blackboard writing for the current round 406 based on the first-round explanation content, and the user makes a response 407 to the question in the blackboard writing 406. The agent outputs the next-round explanation content 408 in the explanation plan upon receiving the response 407, thereby outputting the conclusion blackboard writing 409, and returns to the step of outputting the question blackboard writing to guide the user into subsequent explanation steps in the case where the explanation plan is not completed.
[0059] Further reference is made to Figure 5 , which shows a schematic block diagram of an interaction device for an agent to solve problems according to some embodiments of the present disclosure. Specifically, the interaction method can be implemented by the interaction device 5. The interaction device 5 includes a processor and a memory (not shown), where the processor can refer to various implementations of digital circuit systems, analog circuit systems, or mixed-signal (a combination of analog and digital) circuit systems that perform functions in a computing system. The processing circuit can include, for example, circuits such as integrated circuits (ICs), application-specific integrated circuits (ASICs), parts or circuits of a single processor core, the entire processor core, a single processor, programmable hardware devices such as field-programmable gate arrays (FPGAs), and / or systems including multiple processors. Additionally, the memory of the interaction device 5 can store information generated by the processor and programs and data for the operation of the processor. The memory can be volatile memory and / or non-volatile memory. For example, the memory can include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Generally, the processor can be configured to execute instructions stored in the memory to implement the interaction method for an agent to solve problems in the present disclosure.
[0060] Specifically, as Figure 5As shown, in some embodiments, the interaction device 5 of the present disclosure may further include a receiving module 51, a guiding module 52, a first output module 53, a first receiving module 54, and a second output module 55. Specifically, the receiving module 51 is configured to receive the questions input by the user in the interaction interface of the intelligent agent; the guiding module 52 is configured to display guiding information on the interaction interface for guiding the user to determine the way of explaining the questions; the first output module 53 is configured to generate and output a first content on the interaction interface based on the questions and the way of explaining the questions; the first receiving module 54 is configured to receive the first reply of the user to the first content; and the second output module 55 is configured to generate and output a second content on the interaction interface based on the questions and the first reply.
[0061] The present disclosure also provides an interaction device, which may include a memory; and a processor coupled to the memory, the processor being configured to execute, based on instructions stored in the memory, the interaction method for the intelligent agent to explain questions according to any one of the foregoing embodiments of the present disclosure. The interaction device may refer to Figure 6 , which shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0062] The memory 61 is used to store one or more computer-readable instructions. The memory 61 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 61 may store, for example, an operating system, application programs, a boot loader (BootLoader), a database, and other programs, and may also store various application programs and various data, etc.
[0063] The processor 62 is used to run the computer-readable instructions to implement the song screening method described in any one of the foregoing embodiments or the method described in any one of the foregoing embodiments. The specific implementation of each step of the method may refer to the above embodiments, and the repeated parts will not be elaborated here.
[0064] The processor 62 may be configured to execute Figures 1 to 7 the steps in. The processor 62 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be of the X86 or ARM architecture, etc.
[0065] The processor 62 and the memory 61 can communicate with each other directly or indirectly. For example, the processor 62 and the memory 61 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 62 and the memory 61 can also communicate with each other through a system bus, which is not limited in this disclosure.
[0066] It should be noted that Figure 6 The components of the electronic device 6 shown are merely exemplary and not restrictive. According to actual application needs, the electronic device 6 may also have other components. The processor 62 can control other components in the electronic device 6 to perform desired functions.
[0067] The electronic device 6 can be implemented in software, firmware, and / or hardware, and can be integrated into a device installed with relevant application programs.
[0068] Figure 7 A block diagram of an electronic device according to other embodiments of the present disclosure is shown.
[0069] Figure 7 The electronic device 7 shown can be a computer system with a dedicated hardware structure, and can perform corresponding functions when installed with relevant application programs.
[0070] The electronic device includes but is not limited to mobile terminals such as smartphones, laptop computers, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0071] As Figure 7 shown, the central processing unit (CPU) 71 executes various processes according to the programs stored in the read-only memory (ROM) 72 or the programs loaded from the storage section 78 into the random access memory (RAM) 73. In the RAM 73, data required when the CPU 71 executes various processes, etc., is stored as needed. The central processing unit is merely exemplary, and it can also be other types of processors, such as the various processors described above. The ROM 72, the RAM 73, and the storage section 78 can be various forms of computer-readable storage media. It should be noted that although Figure 7 the ROM 72, the RAM 73, and the storage section 78 are shown separately, one or more of them can be combined, or located in the same or different memories or storage modules.
[0072] The CPU 71, ROM 72, and RAM 73 are connected to each other via a bus 74. An input / output interface 75 is also connected to the bus 74.
[0073] The following components are connected to the input / output interface 75: an input section 76, such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 77, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 78, including a hard disk, a magnetic tape, etc.; and a communication section 79, including a network interface card such as a LAN card, a modem, etc. The communication section 79 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 7 the parts shown in the electronic device 7 communicate through the bus 74, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0074] As needed, a drive 710 is also connected to the input / output interface 75. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 710 as needed, so that a computer program read therefrom is installed in the storage section 78 as needed.
[0075] In the case where the above series of processes are implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 711.
[0076] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the method described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, including program codes for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from a network through the communication section 79, or installed from the storage section 78, or installed from the ROM 72. When the computer program is executed by the CPU 71, the method of the embodiment of the present disclosure is executed.
[0077] It should be noted that, in the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0078] A computer-readable medium can be a computer-readable storage medium, or a computer-readable signal medium, or any combination of the above two.
[0079] A computer-readable storage medium includes, but is not limited to, systems, devices, or components of electricity, magnetism, optics, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component. Computer instructions are stored on the computer-readable storage medium, and when executed by a processor, the instructions implement the method described in any of the foregoing embodiments.
[0080] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0081] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device.
[0082] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the method described in any of the foregoing embodiments. For example, the instructions may be embodied as computer program code.
[0083] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0085] The functions described above may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.
[0086] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An interaction method, comprising: Receiving a question entered by a user in an interaction interface of an agent; Displaying guiding information on the interaction interface, where the guiding information is used to guide the user to determine a topic-explaining manner; Generating and outputting first content on the interaction interface based on the question and the topic-explaining manner; Receiving a first response of the user to the first content; Generating and outputting second content on the interaction interface based on the question and the first response.
2. The interactive method according to claim 1, wherein, Generating and outputting first content on the interaction interface based on the question and the topic-explaining manner includes: In response to the topic-explaining manner including key-point explanation, using a first machine learning model to perform semantic analysis on the question to obtain key content of the question as the first content; or In response to the topic-explaining manner including idea inspiration, using a second machine learning model to generate a mind map for answering the question as the first content.
3. The interactive method according to claim 2, wherein, Receiving a first response of the user to the first content includes: In response to the topic-explaining manner including key-point explanation, receiving problem points selected by the user from the key content as the first response; or In response to the topic-explaining manner including idea inspiration, receiving the interaction of the user with the mind map as the first response.
4. The interaction method according to claim 1, further comprising: In response to the second content including a question, receiving a second response of the user to the question; And Generating and outputting third content on the interaction interface based on the second content and the second response.
5. The interactive method according to claim 4, wherein: Generating and outputting third content on the interaction interface based on the second content and the second response includes: In response to the time from outputting the second content to receiving the second response exceeding a specified time interval, or the second response containing a specified negative semantics, determining an explanation bottleneck in the second content; and Generating a reference hint for the explanation bottleneck as the third content.
6. The interaction method according to claim 4, wherein, Generating and outputting third content on the interaction interface based on the second content and the second response includes: In response to the second response having a correct corresponding relationship with the question, generating an answer to the question based on the second content and the second response as the third content.
7. The interactive method according to claim 4, wherein, Generating and outputting third content on the interaction interface based on the second content and the second response includes: In response to the second response containing a specified end semantics, summarizing the topic-explaining interaction between the user and the agent as the third content.
8. The interactive method according to claim 4, wherein: Generating and outputting third content on the interaction interface based on the second content and the second response includes: In response to the second response having no relation to the question, displaying a prompt message on the interaction interface, where the prompt message is used to prompt the user to determine whether to continue the topic-explaining interaction.
9. The interactive method according to claim 4, wherein, Generating and outputting second content on the interaction interface based on the question and the first response includes: Determining an explanation plan including multiple steps based on the attribute information of the question; and In response to the first reply including the starting point of the explanation determined by the user, determine, among the multiple steps, the starting step for the agent to give an explanation as the second content.
10. The interactive method according to claim 9, wherein, Generating and outputting third content on the interaction interface based on the second content and the second reply includes: In response to the knowledge point involved in the second reply being outside the explanation plan but associated with the question, generating the third content based on the knowledge point.
11. The interactive method according to claim 9, wherein, Generating and outputting third content on the interaction interface based on the second content and the second reply includes: In response to the completion of the entire question explanation plan, summarizing the interaction between the user and the agent for the question explanation as the third content; or In response to the second reply containing a specified end semantics, interrupting the question explanation plan and taking the new interaction between the user and the agent as the third content.
12. An interaction device, comprising: A receiving module configured to receive a question input by a user in an interaction interface of an agent; A guiding module configured to display guiding information on the interaction interface, the guiding information being used to guide the user to determine an explanation mode; A first output module configured to generate and output first content on the interaction interface based on the question and the explanation mode; A first receiving module configured to receive a first reply of the user to the first content; And A second output module configured to generate and output second content on the interaction interface based on the question and the first reply.
13. An electronic device, comprising: A memory; And A processor coupled to the memory, the processor being configured to execute the interaction method according to any one of claims 1 to 11 based on instructions stored in the memory.
14. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor, executes the interaction method according to any one of claims 1 to 11.
15. A computer program product, comprising computer-executable instructions that, when executed by a processor, cause the processor to implement the interaction method according to any one of claims 1 to 11.