Blackboard-writing generation-driven multi-mode dialogue learning guide content generation method, device and system
By building a multi-agent system, combining large language models and cross-modal reasoning technology, synchronously update the content of the blackboard writing, the integration of oral response and blackboard writing generation in the existing technology is solved, and more efficient learning effects and student interaction promotion are achieved.
Patent Information
- Application Number
- CN202510184164.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The existing technology has failed to effectively integrate oral response and blackboard production, lacks a cross-modal reasoning mechanism, and has insufficient promotion of students' social learning.
By building a multi-agent system, using a large language model to generate text responses, and combining problem-solving process information and historical blackboard content, the blackboard content is synchronized through cross-modal reasoning, and dynamically generate blackboard content updated synchronously with oral responses.
Effectively reduce the cognitive load of learners, improve learning effect, enhance the ability to generalize new questions, and promote students' interaction and social construction learning.
Smart Images

Figure CN120124641A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent education, and particularly to a method, device and system for generating multi-modal conversational guidance content driven by blackboard writing generation. Background Art
[0002] Currently, the commonly used prior art in the industry is as follows:
[0003] With the rapid development of large language models in intelligent education, they have significantly improved students' problem-solving abilities. For example, parents can use large language models for Chinese composition tutoring, as discussed in the literature X. Zhuang et al., ‘TOREE: Evaluating Topic Relevance of Student Essays for Chinese Primary and Middle School Education’, in Findings of the Association for Computational Linguistics ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds., Bangkok, Thailand and virtual meeting: Association for Computational Linguistics, Aug. 2024, pp. 5749–5765. Accessed: Aug. 13, 2024. To enhance the ability of large language models to handle complex tasks, an agent workflow is introduced into large language models to strengthen their ability to handle complex tasks, as discussed in the literature Y.-H. Jiang, J. Shi, Y. Tu, Y. Zhou, W. Zhang, and Y. Wei, ‘For Learners: AI Agent is All You Need’, in Enhancing Educational Practices: Strategies for Assessing and Improving Learning Outcomes, Y. Wei, C. Qi, Y.-H. Jiang, and L. Dai, Eds., in Education in a Competitive and Globalizing World., New York, NY, USA: Nova Science Publishers, 2024, pp. 21–46. doi: https: / / doi.org / 10.52305 / RUIG5131.Research has found that multi-agent systems supported by agent workflows have been able to effectively solve many complex problems in teaching, including: 1) classroom scenario simulation, as discussed in the literature J. Shi, J. Zhao, Y. Wang, X. Wu, J. Li, and L. He, ‘CGMI: Configurable general multi-agent interaction framework’, ArXiv Prepr. ArXiv230812503, 2023; 2) contextualized mathematics problem generation, as discussed in the literature R. Li, Y. Wang, C. Zheng, Y.-H. Jiang, and B. Jiang, ‘Generating Contextualized Mathematics Multiple-Choice Questions Utilizing Large Language Models’, in Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium and Blue Sky, A.M. Olney, I.-A. Chounta, Z. Liu, O.C. Santos, and I.I. Bittencourt, Eds., Cham: Springer Nature Switzerland, Jul. 2024, pp. 494–501. doi: 10.1007 / 978-3-031-64315-6_48.
[0004] Learning in the real world is a multimodal behavior. Teachers often help students understand abstract concepts by synchronously dictating and writing on the board, as discussed in the literature C. Choudary and T. Liu, ‘Summarization of Visual Content in Instructional Videos’, IEEE Trans. Multimed., vol. 9, no. 7, pp. 1443–1455, Nov. 2007, doi: 10.1109 / TMM.2007.906602. This behavior of writing or annotating on a blackboard, whiteboard, or screen is called bansho, as discussed in the literature S. Tan, ‘Enhancing your teaching with traditional bansho board writing’, Teacher Magazine. Accessed: Aug. 28, 2024. Research has found that bansho can effectively promote collaborative behavior in teams, as discussed in the literature S. Mailes-Viard Metz, P. Marin, and E. Vayre, ‘The shared online whiteboard: An assistance tool to synchronous collaborative design’, Eur. Rev. Appl. Psychol., vol. 65, no. 5, pp. 253–265, Sep. 2015, doi: 10.1016 / j.erap.2015.08.001; it can also guide the emergence of higher-level metacognitive activities, as discussed in the literature M.F. Teng, ‘Interactive-whiteboard-technology-supported collaborative writing: Writing achievement, metacognitive activities, and co-regulation patterns’, System, vol. 97, p. 102426, Apr. 2021, doi: 10.1016 / j.system.2020.102426, and is widely used in many scenarios, including: 1) software design, as discussed in the literature N. Mangano, T.D. LaToza, M. Petre, and A. van der Hoek, ‘How Software Designers Interact with Sketches at the Whiteboard’, IEEE Trans.As discussed in Softw.Eng., vol.41, no.2, pp.135–156, Feb. 2015, doi:10.1109 / TSE.2014.2362924; 2) interactive teaching, as discussed in the literature K.R. Subramanian and T.Cassen, ‘Across-domain visual learning engine for interactive generation of instructional materials’, SIGCSE Bull, vol.40, no.1, pp.488–492, Mar. 2008, doi:10.1145 / 1352322.1352300; 3) distance education, as discussed in the literature E.A.M. Reguera and M. Lopez, ‘Using a digital whiteboard for student engagement in distance education’, Comput.Electr.Eng., vol.93, p.107268, Jul. 2021, doi:10.1016 / j.compeleceng.2021.107268.
[0005] Taking the mathematics discipline as an example, the generative ability of large language models plays an important role in mathematics learning. Especially when generating text answers, it helps students understand and solve problems. Research has found that BW can promote students' conceptualization, as discussed in the literature by F. Saltan, 'The New Generation of Interactive Whiteboards: How Students Perceive and Conceptualize?', Particip. Educ. Res., vol. 6, no. 2, Art. no. 2, Dec. 2019, doi: 10.17275 / per.19.15.6.2. Blackboard writing can also reduce students' cognitive load, as discussed in the literature by O. M. A. Aldalalah, 'The Effectiveness of Infographic via Interactive Smart Board on Enhancing Creative Thinking: A Cognitive Load Perspective', Int. J. Instr., vol. 14, no. 1, pp. 345–364, Jan. 2021, and has a positive impact on students' knowledge construction, as discussed in the literature by J. B. Pena-Shaff and C. Nicholls, 'Analyzing student interactions and meaning construction in computer bulletin board discussions', Comput. Educ., vol. 42, no. 3, pp. 243–265, Apr. 2004, doi: 10.1016 / j.compedu.2003.08.003. However, traditional mathematics teaching usually only relies on text content for explanation and ignores visual aids such as blackboard writing. Blackboard writing is very important for students' knowledge construction. Although there are now studies dedicated to improving the reasoning ability of LLMs, few studies have focused on how to generate both oral responses and blackboard writing content simultaneously to provide more comprehensive and intuitive support for mathematics learning. However, although existing research has made progress in reasoning ability, there is still no research that can generate and display oral responses and blackboard writing content simultaneously in teaching, which limits the application of multimodal guided learning methods.
[0006] In summary, the problems existing in the prior art are:
[0007] Failure to effectively integrate oral responses with blackboard writing generation: Most of the existing technologies focus on how to improve the performance of LLMs in mathematical reasoning. However, in multimodal teaching, how to effectively combine oral responses with blackboard writing generation to reduce students' cognitive load has not been fully addressed.
[0008] Lack of cross-modal reasoning mechanism: Although some multimodal learning models have been applied to the education field, most of them ignore the importance of cross-modal reasoning in the generation of blackboard writing content. How to update the blackboard writing content while generating oral responses to form a dynamic learning path remains an unsolved technical problem.
[0009] Insufficient promotion of students' social learning: Although multimodal large language models can improve students' performance in mathematical reasoning, their effect on promoting interaction and social construction learning among students is still insufficient. The existing technologies lack system designs that can effectively promote students' interaction and cooperation during the learning process. Summary of the Invention
[0010] To solve the problems of the existing technologies, the objective of the present invention is to provide a multimodal conversational guiding content generation method, device, and system combined with blackboard writing generation. The method uses a large language model to construct a multi-agent system to generate text responses, and combines problem-solving process information and historical blackboard writing content to synchronously update the blackboard writing content through cross-modal reasoning. In this way, the system can dynamically generate blackboard writing content that is synchronized with oral responses, effectively reducing the cognitive load of learners and improving learning effects.
[0011] The innovation of the present invention lies in introducing visual information in the form of blackboard writing, combining the self-summary and induction capabilities of the multi-agent system to perform schema distillation, and storing the obtained meta-thinking template and meta-blackboard writing template in the schema buffer. This design not only enhances the generalization ability for new question types but also significantly improves the generation ability of guiding content. Specifically, with the help of the multi-agent system, it is possible to continuously optimize and generate new blackboard writing content based on historical interaction data, making each interaction step more personalized and more in line with the needs of learners.
[0012] In addition, through the gradually generated blackboard writing information, learners can obtain more intuitive support during the problem-solving process, thereby maximizing the promotion of students' understanding and knowledge mastery in teaching interactions. Compared with the traditional teaching method that only relies on text, this method can effectively improve the interactivity of teaching and improve learners' cognitive load management, thus bringing better learning effects in practical applications. By combining blackboard writing generation with conversational teaching, the present invention significantly improves learners' learning experience and provides a new solution for multimodal teaching methods in the future intelligent education field.
[0013] The specific technical solution to achieve the objective of the present invention is as follows:
[0014] A method for generating multimodal conversational guidance content driven by blackboard writing generation, the method comprising the following steps:
[0015] S1: Based on a large language model, construct an agent workflow for generating multimodal conversational guidance content. The agent workflow for generating multimodal conversational guidance content includes a multi-agent system module, a schema buffer module, and a text-to-speech module;
[0016] S2: The user inputs a question, and the question distillation agent performs a distillation operation on the input question and extracts a mind schema matching the current question type from the schema buffer;
[0017] S3: Based on the matched mind schema, the instantiated reasoning agent instantiates the reasoning of the question to obtain data on the correct answer and solution steps of the question;
[0018] S4: Based on the question input by the user, the AI teacher agent actively initiates a conversation with the user, asking about the specific problems the user faces when solving the question;
[0019] S5: Based on the data of the correct answer and solution steps of the question, the AI teacher agent synchronously generates corresponding oral responses and blackboard writing content for each input of the user. The oral response is obtained by the agent workflow first generating text to reply to the user, and then converting the text to reply to the user into audio and playing it; the blackboard writing content is generated and drawn by the agent workflow and is provided to the user synchronously with the oral response to help the user understand the process of solving the question;
[0020] S6: Determine whether the user ends the multimodal conversational guidance. If not, go back to step S5 to continue generating guidance content; if so, go to step S7;
[0021] S7: After the multimodal conversational guidance ends, the mind schema distillation agent and the blackboard writing schema distillation agent respectively perform distillation operations on the oral response and the blackboard writing content.
[0022] To enhance the multimodal conversational guidance content generation ability of the agent workflow, in step S1, the agent workflow for generating multimodal conversational guidance content includes a multi-agent system module.
[0023] The multi-agent system module is developed based on a large language model and includes five parts: a question distillation agent, an instantiated reasoning agent, an AI teacher agent, a thought schema distillation agent, and a blackboard schema distillation agent. The question distillation agent is suitable for simplifying the questions input by the user and classifying the questions, so as to facilitate the agent workflow to understand the questions to be solved; the instantiated reasoning agent is suitable for performing reasoning operations by combining the thought schema of a specific type of question, so as to facilitate the agent workflow to correctly reason out the correct answers and the data of the solution steps of the questions input by the user; the AI teacher agent is suitable for tutoring the user to solve problems and synchronously drawing the blackboard content to help the user learn to solve the input questions; the thought schema distillation agent is suitable for distilling the text interaction content between the user and the AI teacher agent to enhance the text generation ability of the agent workflow; the blackboard schema distillation agent is suitable for distilling the dynamically changing blackboard content during the interaction between the user and the AI teacher agent to enhance the blackboard generation ability of the agent workflow.
[0024] In order to enable the agent workflow to perform differential question solving and multi-modal guided learning processes according to different question types, in step S1, the agent workflow for generating multi-modal conversational guided learning content includes a schema buffer module.
[0025] The schema buffer module stores two types of content: thought schemas and blackboard schemas. Among them, the thought schemas are stored in text form, and each thought schema describes three parts: the task description, solution instructions, and reasoning thought chain of a type of question. The task description contains the text description of the type of problem; the solution instructions briefly summarize the method for solving the type of problem; the reasoning thought chain includes the text description of the general solution steps executed in sequence specific to the type of question.
[0026] The blackboard schema consists of blackboard output style instructions. The output style instructions specify the generated blackboard content and its structured storage form. The blackboard content is generated and drawn by the agent workflow and is provided to the user synchronously with the oral response to help the user understand the process of solving the problem; the oral response is obtained by the agent workflow first generating the text to reply to the user, and then converting the text to reply to the user into audio and playing it.
[0027] To simplify the content of the questions input by the user and delete the text information irrelevant to the questions, in step S2, the question distillation agent performs distillation operations on the input questions according to the following steps:
[0028] Step 2-1: The user inputs a question to the agent workflow, and the question distillation agent obtains the text representation QUES of the question QUES input by the user from the agent workflow Text:
[0029] QUES Text = AW.GetQues(User.Input(QUES))
[0030] Among them, QUES represents the question input by the user; QUES Text represents the text representation of the question input by the user, and the formulas contained therein are stored in text form after being converted in LaTex style. The LaTex style is a conversion rule for structuring the description of formulas in text form. AW represents the agent workflow, GetQues(·) represents the operation function of the question text representation. User represents the user, Input(·) represents the operation function of user input, and AW.GetQues(User.Input(QUES)) represents the operation process of the question distillation agent obtaining the text representation of the question input by the user from the agent workflow;
[0031] Step 2-2: Perform distillation operation on the text representation QUES of the question input by the user Text The meaning of the distillation operation is to delete the content irrelevant to the question and simplify the text representation of the question, satisfying:
[0032] QUES Clean = PDAgent.Clean(QUES Text ) = {MATH Key , TEXT Key}
[0033] Among them, QUES Clean represents the simplified text representation of the question, which is obtained after performing the distillation operation on the text representation of the question input by the user. PDAgent represents the question distillation agent, and PDAgent.Clean(·) is the distillation operation function of the question distillation agent. Among them, Clean(·) is the distillation operation function, and the input of this function is QUES Text , and the QUES Text represents the text representation of the question input by the user. {MATH Key , TEXT Key} represents the set composed of two elements, MATH Key and TEXT Key . The MATH Key is the key formula content contained in the simplified text representation QUES of the question, and these formula contents are organized in LaTex style; the TEXT Clean is the key text content contained in the simplified text representation QUES of the question. Key is the key text content contained in the simplified text representation QUES of the question. Clean The
[0034] Step 2-3: Extract the task description part of all thinking schemas from the schema buffer. According to the simplified text representation QUES of the topic processed in Step 2-2 Clean , by matching the similarity between QUES Clean and the task description parts of all thinking schemas, retrieve the thinking schema corresponding to the current topic from the schema buffer. The QUES Clean represents the simplified text representation of the topic. This process satisfies:
[0035]
[0036] where, TSchema Match represents the matched thinking schema. represents the thinking schema with the maximum similarity to QUES Clean . Sim(·) is the similarity calculation function between the thinking schema and the simplified text representation of the topic. The result range of this function is [0,1], where 0 means completely dissimilar and 1 means completely matching. The QUES Clean represents the simplified text representation of the topic. i is the number pointer of the thinking schema, and TSN is the total number of thinking schemas in the schema buffer. TSchema represents the thinking schema, and TSchema i represents the i-th thinking schema in the schema buffer. s.t. appearing in the formula represents the meaning of the constraint condition. Sim Max is called the maximum similarity, which represents the maximum value in the real number set composed of the similarity values between all thinking schemas and QUES Clean . θ represents the similarity threshold, which is the preset minimum similarity requirement. If the maximum similarity Sim Max is less than θ, it is determined that the matching fails. TSchema COT represents the thinking schema for solving general problems based on the thinking chain strategy. When the similarities of all thinking schemas are lower than the similarity threshold θ, return this thinking schema TSchema for solving general problems based on the thinking chain strategy COT , where COT represents the thinking chain strategy, that is, the execution strategy that requires the large language model to split and solve complex tasks step by step. {TASK Desc ,SOL Explain ,LOGIC Chain} represents the set composed of the three elements TASK Desc , SOL Explain and LOGIC Chain . The TASK Desc represents the task description, which describes the topic type and the overall information of the task; the SOL ExplainIndicates the solution instructions, summarizing the main steps and methods for solving the problem; the LOGIC Chain Indicates the reasoning thought chain, describing the sequential logical reasoning steps required to solve this type of problem. Extract(·) represents the parsing operation function, Extract(TSchema Match ) represents performing the parsing operation function on the matched thought schema TSchema Match . Indicates the task description of the matched thought schema TSchema Match . Indicates the solution instructions of the matched thought schema TSchema Match . Indicates the reasoning thought chain of the matched thought schema TSchema Match .
[0037] In order to provide targeted guidance based on the user's input, in step S5, for each user input, the AI teacher agent synchronously generates corresponding oral responses and blackboard writing content according to the following steps:
[0038] Step 5-1: The AI teacher agent stores the currently displayed blackboard writing content BW Current and the dialogue history data CH History . In addition, the AI teacher agent retrieves the correct answer CR and the solution steps RSteps of the question stored in the instantiated reasoning agent. These data are integrated by the AI teacher agent into the short-term memory pool Memory Short :
[0039] Memory Short ={BW Current ,CH History ,CR,RSteps}
[0040] where, {BW Current ,CH History ,CR,RSteps} represents the set composed of BW Current ,CH History ,CR and RSteps. BW Current represents the currently displayed blackboard writing content, CH History represents the dialogue history data, CR represents the correct answer of the question, and RSteps represents the solution steps of the question. Memory Short represents the short-term memory pool storing the above data.
[0041] Step 5-2: The AI teacher agent passes through the data in the short-term memory pool Memory Short and the matched blackboard writing schema BWSchema Match, use the cross-modal reasoning function CMReasoning(·) to generate the next state BW of the blackboard writing Next :
[0042]
[0043] Among them, {BW Current , CH History , CR, RSteps} represents the set composed of BW Current , CH History , CR and RSteps. BW Current represents the current blackboard writing content being displayed, CH History represents the dialogue history data, CR represents the correct answer to the question, and RSteps represents the solution steps of the question. Memory Short represents the short-term memory pool for storing the above data. CMReasoning(·) represents the cross-modal reasoning function. BWSchema Match represents the matching blackboard writing schema. MapSchema(·) represents the mapping operation function, which is used to retrieve the matching blackboard writing schema BWSchema from the schema buffer according to the matching thinking schema TSchema Match . TSchema Match represents the matching thinking schema.
[0044] Step 5-3: The AI teacher agent updates the current blackboard writing content BW Next through cross-modal reasoning operations with the next state BW of the generated blackboard writing Current , satisfying the formula:
[0045] BW Current ← BW Next
[0046] Among them, BW Next represents the next state of the blackboard writing, which is generated by the cross-modal reasoning function CMReasoning(·) in Step 5-2. BW Current represents the current blackboard writing content being displayed, and the ← symbol represents the update operation.
[0047] Step 5-4: The AI teacher agent extracts the region of interest ROI Current of the blackboard writing by comparing the blackboard writing content BW Next before the update and the blackboard writing content BW BW after the update, and generates the text representation Response Text of the oral response corresponding to the current update operation Text :
[0048]
[0049] Among them, Response Text represents the text representation of the verbal response to the current update operation. F Response represents the inference function, which is used to generate the text representation of the verbal response Response BW based on the region of interest of the blackboard writing ROI Short and the short-term memory pool Memory Text . ROI BW represents the region of interest of the blackboard writing, which is obtained by calculating the difference between BW Next and BW Current . represents the difference calculation symbol, indicating the operation of calculating the difference of the changed part of the blackboard writing content from BW Current to BW Next . Memory Short represents the short-term memory pool, which contains the current blackboard writing content, conversation history, correct answer, and solution steps. {BW Current , CH History , CR, RSteps} represents the set composed of BW Current , CH History , CR, and RSteps. BW Current represents the currently displayed blackboard writing content, CH History represents the conversation history data, CR represents the correct answer to the question, and RSteps represents the solution steps to the question.
[0050] Step 5-5: The AI teacher agent calls the text-to-speech module Trans TTS in the agent workflow Text to convert the text representation Response Audio of the generated verbal response into audio Response
[0051] Response Audio = Trans TTS (Response Text )
[0052] Subsequently, this audio Response Audio is played to the user. Among them, Response Text represents the text representation of the input verbal response, which is generated in Step 5-4. Trans TTS (·) represents the text-to-speech operation function of the text-to-speech module, which is used to convert the text representation of the verbal response into audio. Response Audio is the generated audio content, which is used to play to the user.
[0053] To enhance the ability of the agent workflow to generate guiding content for new problem types, in step S7, the mind schema distillation agent and the blackboard schema distillation agent perform distillation operations on the oral response and the blackboard content respectively according to the following steps:
[0054] Step 7-1: Extract the matched mind schema TSchema from the problem distillation agent Match . At the same time, extract TSchema from the schema buffer COT . The TSchema Match represents the matched mind schema, and the TSchema COT represents the mind schema for solving general problems based on the chain of thought strategy, where COT represents the chain of thought strategy, that is, the execution strategy that requires the large language model to split complex tasks into steps for solution.
[0055] Step 7-2: Determine whether the matched mind schema TSchema Match is the same as TSchema COT , where TSchema Match represents the matched mind schema, and TSchema COT represents the mind schema for solving general problems based on the chain of thought strategy, where COT represents the chain of thought strategy, that is, the execution strategy that requires the large language model to split complex tasks into steps for solution. Only when the similarity of all mind schemas is lower than the similarity threshold θ will TSchema COT be used. θ represents the similarity threshold, which is the preset minimum similarity requirement. If the matched mind schema TSchema Match is the same as TSchema COT , it means that there is no mind schema and blackboard schema specific to the problem of the user's current input in the schema buffer. At this time, go to the subsequent steps; if the matched mind schema TSchema Match is not the same as TSchema COT , it means that there are already mind schemas and blackboard schemas specific to the problem of the user's current input in the schema buffer. At this time, directly end.
[0056] Step 7-3: The mind schema distillation agent extracts the conversation history data CH from the AI teacher agent History . The conversation history data is a set composed of the text messages sent by the student and the text representation Response of the oral response generated by the AI teacher agent Text . The mind schema distillation agent performs distillation operations on the obtained conversation history data CH History to obtain a new mind schema TSchema specific to the problem type of the user input New , satisfying:
[0057]
[0058] Among them, TSchema New represents a new thought schema specific to the problem type of user input. DistillTS(·) represents the distillation operation function of the thought schema distillation agent, and DistillTS(Response Text , UserResponse) represents the parsing operation function of the thought schema distillation agent on Response Text and UserResponse. The Response Text represents the text representation of the oral response generated by the AI teacher agent, and the UserResponse represents the text message sent by the student. represents the set composed of and These three elements. The represents the task description of the new thought schema specific to the problem type of user input, which describes the overall information of the problem type and the task; the represents the solution instructions of the new thought schema specific to the problem type of user input, which summarizes the main steps and methods of solving the problem; the represents the reasoning thought chain of the new thought schema specific to the problem type of user input, which describes the sequential logical reasoning steps required to solve this type of problem. Extract(·) represents the parsing operation function, and Extract(CH History ) represents the parsing operation function on the matched conversation history data CH History .
[0059] Step 7-4: The blackboard writing schema distillation agent extracts all the blackboard writing contents BW generated in its conversation from the AI teacher agent, and organizes a series of blackboard writing contents presented in chronological order of generation during the dialog-based guided learning process into a time series, called the blackboard writing time series BW Seq . Among them, BW represents the blackboard writing content, and BW Seq represents the blackboard writing time series sorted by the generation time of the blackboard writing content. The blackboard writing schema distillation agent performs a distillation operation on the obtained blackboard writing time series BW Seq to obtain a new blackboard writing schema BWSchema New specific to the problem type of user input, satisfying:
[0060]
[0061] Among them, BWSchema NewA new blackboard writing schema representing the problem type specific to the user input, and DistillBWS(·) represents the distillation operation function of the blackboard writing schema distillation agent. DistillBWS(BW Seq ) represents the parsing operation function of the blackboard writing schema distillation agent on the blackboard writing time series BW Seq . The BW Seq represents organizing a series of blackboard writing contents presented in chronological order of generation during the conversational guided learning process into a time series, which is called the blackboard writing time series. BW represents the blackboard writing content, i represents the sequence number pointer of the blackboard writing content, and NDR represents the dialogue turn. BW 1 represents the blackboard writing content generated by the AI teacher agent in the first round of dialogue, and BW 2 represents the blackboard writing content generated by the AI teacher agent in the second round of dialogue, and BW i represents the blackboard writing content generated by the AI teacher agent in the i-th round of dialogue, and BW NDR represents the blackboard writing content generated by the AI teacher agent in the NDR-th round of dialogue.
[0062] Step 7-5: Add the new mind map schema TSchema New specific to the problem type of the user input obtained through the distillation operation and the new blackboard writing schema BWSchema New specific to the problem type of the user input to the schema buffer of the agent workflow, so that the agent workflow can self-reflect and summarize for new question types that have not been answered, enhancing the scalability of the agent workflow, and satisfying:
[0063]
[0064] where SchemaBuffer represents the schema buffer, ← represents the assignment operation symbol, TSchema New represents the new mind map schema specific to the problem type of the user input, and BWSchema New represents the new blackboard writing schema specific to the problem type of the user input. TSN is the total number of mind map schemas in the schema buffer, and SchemaBuffer.TSN represents the value of the total number TSN of mind map schemas obtained from the schema buffer SchemaBuffer.
[0065] The present invention also provides a multimodal conversational guided learning content generation device driven by blackboard writing generation, including a terminal device. The terminal device uses an Internet terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor for the multimodal conversational guided learning content generation method described above.
[0066] The present invention also provides a multi-modal conversational learning guidance content generation system driven by blackboard writing generation, including a computer-readable storage medium, which stores multiple instructions and an agent workflow. The instructions are adapted to be loaded and executed by a processor of a terminal device for the multi-modal conversational learning guidance content generation method described above. The agent workflow includes a multi-agent system module, a schema buffer module, and a text-to-speech module.
[0067] Compared with the prior art, the multi-modal conversational learning guidance content generation method, device, and system provided by the present invention have the following beneficial effects:
[0068] The present invention belongs to the field of intelligent education, and discloses a multi-modal conversational learning guidance content generation method, device, and system. The multi-modal conversational learning guidance content generation method driven by blackboard writing generation includes: constructing a multi-agent system using a large language model to generate text response content for learners; combining problem-solving process information and historical blackboard writing information to synchronously update the blackboard writing content through cross-modal reasoning operations; enabling the multi-agent system to summarize and generalize itself, distilling the interaction process, and storing the obtained meta-thinking template and meta-blackboard writing template in the schema buffer to enhance the generalization ability for new question types and the generation ability of learning guidance content. This method can gradually generate the visual information of the blackboard writing, effectively reduce the cognitive load of learners by introducing visual information, and thus improve the learning effect in the conversational learning guidance process. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flowchart of the method of the present invention.
[0070] Figure 2 It is a schematic diagram of a user interface of a preferred embodiment of the learning guidance content generation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The present invention will be further described below with reference to the drawings and embodiments.
[0072] As Figure 1 shown, a multi-modal conversational learning guidance content generation method driven by blackboard writing generation, the method includes steps S1 to S7:
[0073] S1: Based on a large language model, construct an agent workflow for multi-modal conversational learning guidance content generation. The agent workflow for multi-modal conversational learning guidance content generation includes a multi-agent system module, a schema buffer module, and a text-to-speech module;
[0074] Further, in order to enhance the multi-modal conversational guided learning content generation ability of the agent workflow, in step S1, the agent workflow for multi-modal conversational guided learning content generation includes a multi-agent system module.
[0075] The multi-agent system module is developed based on a large language model and includes five parts: a question distillation agent, an instantiated reasoning agent, an AI teacher agent, a mind schema distillation agent, and a blackboard schema distillation agent. The question distillation agent is suitable for simplifying the questions input by the user and classifying the questions, so that the agent workflow can understand the questions to be solved; the instantiated reasoning agent is suitable for performing reasoning operations in combination with the mind schema of a specific type of question, so that the agent workflow can correctly reason out the correct answer and the data of the solution steps of the question input by the user; the AI teacher agent is suitable for tutoring the user to solve the problem and synchronously drawing the blackboard content to help the user learn to solve the input question; the mind schema distillation agent is suitable for distilling the text interaction content between the user and the AI teacher agent to enhance the text generation ability of the agent workflow; the blackboard schema distillation agent is suitable for distilling the dynamically changing blackboard content during the interaction between the user and the AI teacher agent to enhance the blackboard generation ability of the agent workflow.
[0076] Further, in order to enable the agent workflow to perform differential question solving and multi-modal guided learning processes according to different question types, in step S1, the agent workflow for multi-modal conversational guided learning content generation includes a schema buffer module.
[0077] The schema buffer module stores two types of content: mind schemas and blackboard schemas. Among them, the mind schemas are stored in text form, and each mind schema describes three parts: the task description, the solution description, and the reasoning thought chain of a type of question. The task description contains the text description of the type of problem; the solution description briefly summarizes the method for solving the type of problem; the reasoning thought chain includes the text description of the general solution steps executed in sequence specific to the type of question.
[0078] The blackboard schema consists of the blackboard output style description. The output style description specifies the generated blackboard content and its structured storage form. The blackboard content is generated and drawn by the agent workflow and is provided to the user synchronously with the oral response to help the user understand the process of solving the question; the oral response is obtained by the agent workflow first generating the text to reply to the user and then converting the text to reply to the user into audio and playing it.
[0079] S2: The user inputs a question, and the question distillation agent performs distillation operations on the input question and extracts the thinking schema matching the current question type from the schema buffer;
[0080] Further, to simplify the content of the question input by the user and delete the text information irrelevant to the question, in step S2, the question distillation agent performs distillation operations on the input question according to the following steps:
[0081] Step 2-1: The user inputs a question into the agent workflow, and the question distillation agent obtains the text representation QUES of the question input by the user from the agent workflow Text :
[0082] QUES Text = AW.GetQues(User.Input(QUES))
[0083] where QUES represents the question input by the user; QUES Text represents the text representation of the question input by the user, in which the formulas contained are stored in text form after being converted in LaTex style. The LaTex style is a conversion rule for structurally describing formulas in text form. AW represents the agent workflow, GetQues(·) represents the operation function for the text representation of the question. User represents the user, Input(·) represents the operation function of user input, and AW.GetQues(User.Input(QUES)) represents the operation process of the question distillation agent obtaining the text representation of the question input by the user from the agent workflow;
[0084] Step 2-2: Perform distillation operations on the text representation QUES of the question input by the user Text The meaning of the distillation operation is to delete the content irrelevant to the question and simplify the text representation of the question, satisfying:
[0085] QUES Clean = PDAgent.Clean(QUES Text ) = {MATH Key , TEXT Key}
[0086] where QUES Clean represents the simplified text representation of the question, which is obtained after performing distillation operations on the text representation of the question input by the user. PDAgent represents the question distillation agent, PDAgent.Clean(·) is the distillation operation function of the question distillation agent, where Clean(·) is the distillation operation function, and the input of this function is QUES Text , and the QUES TextThe text representation of the title input by the user. {MATH Key ,TEXT Key} represents MATH Key and TEXT Key The set composed of two elements, where the MATH Key is the simplified text representation of the title QUES Clean The key formula content contained therein, and these formula contents are organized in LaTex style; the TEXT Key is the simplified text representation of the title QUES Clean The key text content contained therein.
[0087] Step 2-3: Extract the task description part of all thinking schemas from the schema buffer. According to the simplified text representation QUES of the title processed in Step 2-2 Clean , by matching QUES Clean with the similarity of the task description parts of all thinking schemas, retrieve the thinking schema corresponding to the current title from the schema buffer. The QUES Clean represents the simplified text representation of the title. This process satisfies:
[0088]
[0089] where, TSchema Match represents the matched thinking schema. represents the thinking schema with the maximum similarity to QUES Clean , Sim(·) is the similarity calculation function between the thinking schema and the simplified text representation of the title, and the result range of this function is [0,1], where 0 means completely dissimilar and 1 means completely matching. The QUES Clean represents the simplified text representation of the title. i is the number pointer of the thinking schema, and TSN is the total number of thinking schemas in the schema buffer. TSchema represents the thinking schema, and TSchema i represents the i-th thinking schema in the schema buffer. s.t. appearing in the formula represents the meaning of the constraint condition. Sim Max is called the maximum similarity, which represents the maximum value in the real number set composed of the similarity values between all thinking schemas and QUES Clean . θ represents the similarity threshold, which is the preset minimum similarity requirement. If the maximum similarity Sim Max is less than θ, it is determined that the matching fails. TSchema COT represents the thinking schema for solving general problems based on the thinking chain strategy. When the similarity of all thinking schemas is lower than the similarity threshold θ, return this thinking schema for solving general problems based on the thinking chain strategy TSchemaCOT , where COT represents the chain-of-thought strategy, which is an execution strategy that requires the large language model to decompose complex tasks into steps for solving. {TASK Desc , SOL Explain , LOGIC Chain} represents the set composed of the three elements TASK Desc , SOL Explain and LOGIC Chain . The TASK Desc represents the task description, which describes the type of the question and the overall information of the task; the SOL Explain represents the solution description, which summarizes the main steps and methods for solving the problem; the LOGIC Chain represents the chain of reasoning thoughts, which describes the sequential logical reasoning steps required to solve this type of problem. Extract(·) represents the parsing operation function, and Extract(TSchema Match ) represents performing the parsing operation function on the matched thought schema TSchema Match . represents the task description of the matched thought schema TSchema Match , represents the solution description of the matched thought schema TSchema Match , represents the chain of reasoning thoughts of the matched thought schema TSchema Match .
[0090] S3: Based on the matched thought schema, instantiate the inference agent to instantiate the inference of the question, and obtain the data of the correct answer and the solution steps of the question;
[0091] S4: Based on the question input by the user, the AI teacher agent actively initiates a conversation with the user and asks about the specific problems the user faces when solving the question;
[0092] S5: Based on the data of the correct answer and the solution steps of the question, the AI teacher agent synchronously generates corresponding oral responses and blackboard writing contents for each input of the user. The oral response is obtained by the intelligent agent workflow first generating the text to reply to the user and then converting the text to reply to the user into audio and playing it; the blackboard writing content is generated and drawn by the intelligent agent workflow and is used to be provided to the user synchronously with the oral response, so as to help the user understand the process of solving the question;
[0093] For ease of understanding step S5, please refer to Figure 2 in combination. Figure 2 is a schematic diagram of the user interface of a preferred embodiment of the learning guidance content generation provided by the embodiment of the present invention. As Figure 2As shown, for each user input, the AI teacher agent synchronously generates a corresponding oral response and blackboard writing content. The text representation of the generated oral response is displayed in Figure 2 the dialogue area of the user interface of a preferred embodiment shown, and at the same time, the oral response is converted into audio and played for the user. At the same time, the generated blackboard writing content will also be synchronously updated in Figure 2 the blackboard writing content of the user interface of a preferred embodiment shown.
[0094] Furthermore, in order to provide targeted guidance based on the user's input, in step S5, for each user input, the AI teacher agent synchronously generates a corresponding oral response and blackboard writing content according to the following steps:
[0095] Step 5-1: The AI teacher agent stores the currently displayed blackboard writing content BW Current and the dialogue history data CH History . In addition, the AI teacher agent retrieves the correct answer CR and the solution steps RSteps of the question stored in the instantiated reasoning agent. These data are integrated by the AI teacher agent into the short-term memory pool Memory Short :
[0096] Memory Short ={BW Current , CH History , CR, RSteps}
[0097] where {BW Current , CH History , CR, RSteps} represents the set composed of BW Current , CH History , CR, and RSteps. BW Current represents the currently displayed blackboard writing content, CH History represents the dialogue history data, CR represents the correct answer of the question, and RSteps represents the solution steps of the question. Memory Short represents the short-term memory pool storing the above data.
[0098] Step 5-2: The AI teacher agent uses the data in the short-term memory pool Memory Short and the matching blackboard writing schema BWSchema Match to generate the next state BW Next of the blackboard writing by using the cross-modal reasoning function CMReasoning(·):
[0099]
[0100] where {BW Current , CHHistory , CR, RSteps} represents the set composed of BW Current , CH History , CR and RSteps. BW Current represents the blackboard writing content currently being displayed, CH History represents the dialogue history data, CR represents the correct answer to the question, and RSteps represents the solution steps of the question. Memory Short represents the short-term memory pool for storing the above data. CMReasoning(·) represents the cross-modal reasoning function. BWSchema Match represents the matching blackboard writing schema. MapSchema(·) represents the mapping operation function, which is used to retrieve the matching blackboard writing schema BWSchema from the schema buffer according to the matching thinking schema TSchema Match Match . TSchema Match represents the matching thinking schema.
[0101] Step 5-3: The AI teacher agent updates the currently displayed blackboard writing content BW Next through cross-modal reasoning operations with the next state BW Current of the generated blackboard writing, satisfying the formula:
[0102] BW Current ← BW Next
[0103] where BW Next represents the next state of the blackboard writing, generated by the cross-modal reasoning function CMReasoning(·) in Step 5-2. BW Current represents the currently displayed blackboard writing content, and the ← symbol represents the update operation.
[0104] Step 5-4: The AI teacher agent extracts the region of interest of the blackboard writing ROI Current by comparing the blackboard writing content BW Next before the update and the blackboard writing content BW BW after the update, and generates the text representation Response Text of the oral response corresponding to the current update operation:
[0105]
[0106] where Response Text represents the text representation of the oral response for the current update operation. F Response represents the reasoning function, which is used to generate the text representation Response of the oral response based on the region of interest of the blackboard writing ROI BW and the short-term memory pool Memory Short Generate the text representation of the oral response Response Text . ROI BW Represents the area of the blackboard writing of interest, obtained by Next performing a difference calculation with BW Current . Represents the difference calculation symbol, indicating the operation of calculating the difference of the changed part of the blackboard writing content from BW Current to BW Next . Memory Short Represents the short-term memory pool, which contains the current blackboard writing content, conversation history, correct answer, and solution steps. {BW Current , CH History , CR, RSteps} represents the set composed of BW Current , CH History , CR, and RSteps. BW Current represents the currently displayed blackboard writing content, CH History represents the conversation history data, CR represents the correct answer to the question, and RSteps represents the solution steps to the question.
[0107] Step 5-5: The AI teacher agent calls the text-to-speech module Trans in the agent workflow TTS , and converts the text representation Response Text of the generated oral response Audio into an audio Response
[0108] : Audio Response TTS = Trans Text (Response
[0109] Subsequently, this audio Response Audio is played to the user. Among them, Response Text represents the text representation of the input oral response, generated by step 5-4. Trans TTS (·) represents the text-to-speech operation function of the text-to-speech module, which is used to convert the text representation of the oral response into an audio. Response Audio is the generated audio content, which is used to play to the user.
[0110] S6: Determine whether the user ends the multi-modal conversational tutoring. If not, go back to step S5 to continue generating tutoring content; if so, go to step S7;
[0111] S7: After the multi-modal conversational tutoring ends, the mind schema distillation agent and the blackboard writing schema distillation agent respectively perform distillation operations on the oral response and the blackboard writing content.
[0112] Further, to enhance the ability of the agent workflow to generate guiding content for new problem types, in step S7, the mind schema distillation agent and the blackboard writing schema distillation agent perform distillation operations on the oral response and the blackboard writing content respectively according to the following steps:
[0113] Step 7-1: Extract the matched mind schema TSchema from the problem distillation agent Match . At the same time, extract TSchema from the schema buffer COT . The TSchema Match represents the matched mind schema, and the TSchema COT represents the mind schema for solving general problems based on the chain of thought strategy, where COT represents the chain of thought strategy, that is, the execution strategy that requires the large language model to split complex tasks into steps for solution.
[0114] Step 7-2: Determine whether the matched mind schema TSchema Match is the same as TSchema COT , where TSchema Match represents the matched mind schema, and TSchema COT represents the mind schema for solving general problems based on the chain of thought strategy, where COT represents the chain of thought strategy, that is, the execution strategy that requires the large language model to split complex tasks into steps for solution. Only when the similarity of all mind schemas is lower than the similarity threshold θ will TSchema COT be used. θ represents the similarity threshold, which is the preset minimum similarity requirement. If the matched mind schema TSchema Match is the same as TSchema COT , it means that there are no mind schemas and blackboard writing schemas specific to the problem of the user's current input in the schema buffer. At this time, go to the subsequent steps; if the matched mind schema TSchema Match is not the same as TSchema COT , it means that there are already mind schemas and blackboard writing schemas specific to the problem of the user's current input in the schema buffer. At this time, directly end.
[0115] Step 7-3: The mind schema distillation agent extracts the conversation history data CH from the AI teacher agent History . The conversation history data is a set composed of the text messages sent by the student and the text representation Response of the oral response generated by the AI teacher agent Text . The mind schema distillation agent performs distillation operations on the obtained conversation history data CH History to obtain a new mind schema TSchema specific to the problem type of the user input New , satisfying:
[0116]
[0117] Among them, TSchema New represents a new thought schema specific to the problem type of user input. DistillTS(·) represents the distillation operation function of the thought schema distillation agent, and DistillTS(Response Text , UserResponse) represents the parsing operation function of the thought schema distillation agent on Response Text and UserResponse. The Response Text represents the text representation of the oral response generated by the AI teacher agent, and the UserResponse represents the text message sent by the student. represents a set composed of and these three elements. The represents the task description of the new thought schema specific to the problem type of user input, which describes the overall information of the question type and the task; the represents the solution description of the new thought schema specific to the problem type of user input, which summarizes the main steps and methods of solving the problem; the represents the reasoning thought chain of the new thought schema specific to the problem type of user input, which describes the sequential logical reasoning steps required to solve this type of problem. Extract(·) represents the parsing operation function, and Extract(CH History ) represents the parsing operation function on the matched conversation history data CH History .
[0118] Step 7-4: The blackboard writing schema distillation agent extracts all the blackboard writing contents BW generated in its conversation from the AI teacher agent, and organizes a series of blackboard writing contents presented in chronological order of generation during the conversational guidance process into a time series, called the blackboard writing time series BW Seq . Among them, BW represents the blackboard writing content, and BW Seq represents the blackboard writing time series sorted by the generation time of the blackboard writing content. The blackboard writing schema distillation agent performs a distillation operation on the obtained blackboard writing time series BW Seq to obtain a new blackboard writing schema BWSchema New specific to the problem type of user input, satisfying:
[0119]
[0120] Among them, BWSchema NewA new blackboard writing schema representing the problem type specific to the user input, DistillBWS(·) represents the distillation operation function of the blackboard writing schema distillation agent, and DistillBWS(BW Seq ) represents the blackboard writing schema distillation agent performing the parsing operation function on the blackboard writing time series BW Seq . The BW Seq represents organizing a series of blackboard writing contents presented in chronological order of generation during the conversational guided learning process into a time series, which is called the blackboard writing time series. BW represents the blackboard writing content, i represents the serial number pointer of the blackboard writing content, and NDR represents the conversation turn. BW 1 represents the blackboard writing content generated by the AI teacher agent in the first round of conversation, and BW 2 represents the blackboard writing content generated by the AI teacher agent in the second round of conversation, and BW i represents the blackboard writing content generated by the AI teacher agent in the i-th round of conversation, and BW NDR represents the blackboard writing content generated by the AI teacher agent in the NDR-th round of conversation.
[0121] Step 7-5: Add the new mind map schema TSchema specific to the problem type of the user input obtained through the distillation operation New and the new blackboard writing schema BWSchema specific to the problem type of the user input New to the schema buffer of the agent workflow, enabling the agent workflow to self-reflect and generalize for new question types that have not been answered, enhancing the scalability of the agent workflow, and satisfying:
[0122]
[0123] where SchemaBuffer represents the schema buffer, ← represents the assignment operation symbol, TSchema New represents the new mind map schema specific to the problem type of the user input, and BWSchema New represents the new blackboard writing schema specific to the problem type of the user input. TSN is the total number of mind map schemas in the schema buffer, and SchemaBuffer.TSN represents the value of the total number of mind map schemas TSN obtained from the schema buffer SchemaBuffer.
[0124] According to another aspect of one or more embodiments of the present disclosure, there is also provided a multimodal conversational guided learning content generation device driven by blackboard writing generation, including a terminal device. The terminal device uses an Internet terminal device and includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor for the multimodal conversational guided learning content generation method as described above.
[0125] According to another aspect of one or more embodiments of the present disclosure, there is also provided a blackboard writing generation-driven multi-modal conversational learning content generation system, including a computer-readable storage medium, in which multiple instructions and an agent workflow are stored. The instructions are adapted to be loaded and executed by a processor of a terminal device for the described blackboard writing generation-driven multi-modal conversational learning content generation method. The agent workflow includes a multi-agent system module, a schema buffer module, and a text-to-speech module.
[0126] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk memories, CD-ROMs, optical memories, etc.
[0127] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0128] The above are only embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for generating multi-modal conversational learning content driven by blackboard generation, characterized in that: The method comprises the following steps: S1: Based on the large language model, construct an intelligent agent workflow for generating multimodal conversational learning content, wherein the intelligent agent workflow for generating multimodal conversational learning content includes a multi-agent system module, a schema buffer module, and a text-to-speech module; S2: The user inputs a question, and the question distillation agent performs a distillation operation on the input question and extracts a thinking schema that matches the current question type from the schema buffer; S3: Based on the matched thinking schema, the instantiated reasoning agent instantiates the reasoning of the question and obtains the correct answer to the question and the data of the solution steps; S4: Based on the question input by the user, the AI teacher agent actively initiates a dialogue with the user, asking the user about the specific problems faced when solving the question; the AI means artificial intelligence; S5: Based on the data of the correct answer and solution steps of the question, the AI teacher agent synchronously generates the corresponding oral response and blackboard content for each user input; the oral response is obtained by the agent workflow first generating a text reply to the user, and then converting the text reply to the user into audio and playing it; the blackboard content is generated and drawn by the agent workflow, and is provided to the user synchronously with the oral response, so as to help the user understand the process of solving the question; S6: Determine whether the user has ended the multimodal conversational learning guidance; if not, go to step S5 to continue generating the learning guidance content; if it is ended, go to step S7; S7: After the multimodal conversational teaching is completed, the mind map distillation agent and the blackboard map distillation agent perform distillation operations on the verbal response and the blackboard content respectively; In step S1, the multi-agent system module is developed based on a large language model, including a problem distillation agent, an instantiated reasoning agent, an AI teacher agent, a mind map distillation agent and a blackboard map distillation agent; the problem distillation agent is used to simplify the questions input by the user and classify the questions so that the agent workflow can understand the questions to be solved; the instantiated reasoning agent is used to perform reasoning operations in combination with a mind map of a specific category of questions so that the agent workflow can correctly infer the correct answer to the question input by the user and the data of the solution steps; the AI teacher agent is used to coach the user to solve the problem and synchronously draw the blackboard content to help the user learn to solve the input question; the mind map distillation agent is used to distill the text interaction content between the user and the AI teacher agent to enhance the text generation ability of the agent workflow; the blackboard map distillation agent is used to distill the blackboard content that changes dynamically during the interaction between the user and the AI teacher agent to enhance the blackboard generation ability of the agent workflow.
2. According to the method for generating multi-modal conversational learning content driven by blackboard writing generation in claim 1, it is characterized in that: In step S1, the schema buffer module stores two types of content: mind schema and blackboard schema; wherein the mind schema is stored in text form, and each mind schema describes three parts: task description, solution instructions, and reasoning chain of a type of question, wherein the task description includes a text description of the type of question; the solution instructions briefly summarize the method for solving the type of question; and the reasoning chain of thought includes a text description of the general solution steps that are executed in sequence and are specific to the type of question; The blackboard writing diagram is composed of a blackboard writing output style description; the output style description specifies the generated blackboard writing content and its structured storage form; the blackboard writing content is generated and drawn by the intelligent body workflow, and is provided to the user synchronously with the oral response, so as to help the user understand the process of solving the problem; the oral response is obtained by the intelligent body workflow first generating a text to reply to the user, and then converting the text to reply to the user into audio and playing it.
3. The method for generating multi-modal conversational learning content driven by blackboard generation according to claim 1 is characterized in that: In step S2, the question distillation agent performs a distillation operation on the input question, specifically including: Step 2-1: The user inputs a question to the agent workflow, and the question distillation agent obtains the text representation of the user-input question QUES from the agent workflow Text : QUES Text =AW.GetQues(User.Input(QUES)) Among them, QUES represents the question entered by the user; Text The text representation of the question input by the user, wherein the formula contained therein is converted into text form and stored in LaTex style; the LaTex style is a conversion rule for describing formulas in a structured text form; AW represents the agent workflow, GetQues(·) represents the operation function of the question text representation; User represents the user, Input(·) represents the operation function of the user input, and AW.GetQues(User.Input(QUES)) represents the operation process of the question distillation agent obtaining the text representation of the question input by the user from the agent workflow; Step 2-2: QUES the text of the question entered by the user Text Perform a distillation operation; the distillation operation means deleting irrelevant content related to the topic and simplifying the text representation of the topic to meet the following requirements: QUES Clean =PDAgent.Clean(QUES Text )={MATH Key ,TEXT Key } Among them, QUES Clean The simplified text representation of the question is obtained by performing a distillation operation on the text representation of the question input by the user; PDAgent represents the problem distillation agent, and PDAgent.Clean(·) is the distillation operation function of the problem distillation agent, where Clean(·) is the distillation operation function, and the input of this function is QUES Text , the QUES Text Represents the text representation of the question entered by the user; {MATH Key ,TEXT Key } represents MATH Key and TEXT Key The set consisting of two elements, the MATH Key QUES is the simplified text representation of the question Clean The key formulas contained in it are organized in LaTex style; the above is TEXT Key QUES is the simplified text representation of the question Clean The key text content contained in it; Step 2-3: Extract the task description part of all mind schemas from the schema buffer; Simplify the text representation of the question according to step 2-2 to represent QUES Clean , by matching QUES Clean The similarity with the task description part of all mind schemas is used to retrieve the mind schema corresponding to the current question from the schema buffer; the QUES Clean Represents a simplified text representation of the title; this process satisfies: Among them, TSchema Match Indicates the matched mental schema; Representation and QUES Clean The thought pattern with the greatest similarity, Sim(·) is the similarity calculation function between the thought pattern and the simplified text representation of the question, and the result range of this function is [0,1], where 0 indicates complete dissimilarity and 1 indicates complete match; the QUES Clean Indicates the simplified text representation of the question; i is the number pointer of the mind schema, TSN is the total number of mind schemas in the schema buffer; TSchema represents the mind schema, TSchema i represents the i-th mental schema in the schema buffer; the st in the formula represents the meaning of the restriction condition; Sim Max It is called the maximum similarity, which means that all mental schemas are similar to QUES Clean The maximum value in the real number set composed of the similarity values; θ represents the similarity threshold, which is the preset minimum similarity requirement. If the maximum similarity Sim Max If it is less than θ, it is considered a match failure; TSchema COT Represents a thinking schema for solving general problems based on the thinking chain strategy. When the similarity of all thinking schemas is lower than the similarity threshold θ, the thinking schema TSchema for solving general problems based on the thinking chain strategy is returned. COT , where COT stands for the thought chain strategy, which is an execution strategy that requires a large language model to split complex tasks into steps for solution; {TASK Desc ,SOL Explain ,LOGIC Chain } represents the TASK Desc 、SOL Explain and LOGIC Chain These three elements constitute a set, the TASK Desc Represents the task description, which describes the topic type and the overall information of the task; the SOL Explain Indicates the solution instructions, summarizing the main steps and methods of solving the problem; the LOGIC Chain It represents the reasoning chain, describing the logical reasoning steps that need to be executed in sequence to solve this type of problem; Extract(·) represents the parsing operation function, Extract(TSchema Match ) represents the matched mindset TSchema Match Execute parsing operation function; Indicates the matched mindset TSchema Match Description of the task, Indicates the matched mindset TSchema Match The solution description, Indicates the matched mindset TSchema Match chain of reasoning thinking.
4. The method for generating multi-modal conversational learning content driven by blackboard writing generation according to claim 1 is characterized in that: In step S5, the AI teacher agent generates corresponding oral responses and blackboard content for each user input, including: Step 5-1: The AI teacher agent stores the current blackboard content BW Current And the conversation history data CH History ; In addition, the AI teacher agent retrieves the correct answer CR and solution steps RSteps of the question stored in the instantiated reasoning agent; these data are integrated into the short-term memory pool Memory by the AI teacher agent Short : Memory Short ={BW Current ,CH History ,CR,RSteps} Among them, {BW Current ,CH History ,CR,RSteps} represents the Current , CH History , CR and RSteps; BW Current Indicates the current blackboard content, CH History Represents the conversation history data, CR represents the correct answer to the question, and RSteps represents the steps to solve the question; Memory Short Represents a short-term memory pool that stores the above data; Step 5-2: AI teacher agent uses short-term memory pool Memory Short Data in and matching blackboard schema BWSchema Match , using the cross-modal reasoning function CMReasoning(·) to generate the next state BW of the blackboard Next : Among them, {BW Current ,CH History ,CR,RSteps} represents the Current , CH History , CR and RSteps; BW Current Indicates the current blackboard content, CH History Represents the conversation history data, CR represents the correct answer to the question, and RSteps represents the steps to solve the question; Memory Short represents the short-term memory pool storing the aforementioned data; CMReasoning(·) represents the cross-modal reasoning function; BWSchema Match Indicates the matching blackboard schema; MapSchema(·) indicates the mapping operation function, which is used to map the matching mind schema TSchema from the schema buffer Match Retrieve the matching blackboard schema BWSchema Match ;TSchema Match Indicates the matched mental schema; Step 5-3: The AI teacher agent uses the generated next state BW of the blackboard through cross-modal reasoning operation Next Update the current displayed blackboard content BW Current , satisfying the formula: BW Current ←BW Next Among them, BW Next represents the next state of the blackboard writing, which is generated by the cross-modal reasoning function CMReasoning(·) in step 5-2; BW Current Indicates the current blackboard content, and the ← symbol indicates an update operation; Step 5-4: The AI teacher agent compares the blackboard content BW before the update Current and updated blackboard content BW Next , extract the blackboard area of interest ROI BW , and generates a textual representation of the verbal response corresponding to the current update operation Response Text : Among them, Response Text A textual representation of the verbal response to the current update operation; F Response Represents the inference function, which is used to write on the blackboard area of interest ROI BW and short-term memory pool Memory Short Generates a textual representation of a spoken response Text ROI BW Indicates the blackboard area of interest, represented by BW Next With BW Current The difference is calculated; Indicates the difference calculation symbol, indicating that the calculation blackboard content is written by BW Current Change to BW Next Operation of the difference of the changed part; Memory Short Represents the short-term memory pool, which contains the current blackboard content, dialogue history, correct answers and solution steps; {BW Current ,CH History ,CR,RSteps} represents the Current , CH History , CR and RSteps; BW Current Indicates the current blackboard content, CH History Represents the conversation history data, CR represents the correct answer to the question, and RSteps represents the steps to solve the question; Step 5-5: The AI teacher agent calls the text-to-speech module Trans in the agent workflow TTS , the text representation of the generated verbal response is Response Text Convert to AudioResponse Audio : Response Audio =Trans TTS (Response Text ) Then, the audio response Audio is played to the user; Response Text The text representation of the verbal response to the input, generated by step 5-4; Trans TTS (·) represents the text-to-speech operation function of the text-to-speech module, which is used to convert the text representation of the oral response into audio; Response Audio It is the generated audio content for playing to the user.
5. The method for generating multi-modal conversational learning content driven by blackboard generation according to claim 1 is characterized in that: In step S7, the thought diagram distillation agent and the blackboard diagram distillation agent perform distillation operations on the oral response and the blackboard content respectively, specifically including: Step 7-1: Extract the matching mind schema TSchema from the problem distillation agent Match ; At the same time, extract TSchema from the schema buffer COT ; The TSchema Match Indicates the matched mind schema, the TSchema COT Represents a thought diagram for solving general problems based on the thought chain strategy, where COT stands for the thought chain strategy, which is an execution strategy that requires a large language model to split complex tasks into steps for solution; Step 7-2: Determine the matching mindset TSchema Match Is it related to TSchema COT Same, where TSchema Match Indicates the matched mindset, TSchema COT Represents a thought schema for solving general problems based on the thought chain strategy, where COT represents the thought chain strategy, which is an execution strategy that requires a large language model to split complex tasks into steps for solution; TSchema is only valid when the similarity of all thought schemas is lower than the similarity threshold θ COT will be used; θ represents the similarity threshold, which is the preset minimum similarity requirement; if the matching mind schema TSchema Match With TSchema COT If the same, it means that the schema buffer does not have a specific thinking schema or blackboard schema for the question entered by the user. At this time, go to the next step; if the matching thinking schema TSchema Match With TSchema COT If they are not the same, it means that the schema buffer already has a thinking schema and a blackboard schema specific to the question entered by the user. In this case, the system ends directly. Step 7-3: The mind map distillation agent extracts the conversation history data CH from the AI teacher agent History The dialogue history data is a text representation of the text messages sent by the students and the verbal responses generated by the AI teacher agent. Text The set formed by the mind map distillation agent obtains the conversation history data CH History Perform a distillation operation to obtain a new mind schema TSchema specific to the type of question entered by the user New ,satisfy: Among them, TSchema New represents a new mind map specific to the type of question input by the user; DistillTS(·) represents the distillation operation function of the mind map distillation agent, DistillTS(Response Text ,UserResponse) represents the mind map distillation agent’s response to Response Text and UserResponse to perform parsing operations; the Response Text a textual representation of a verbal response generated by the AI teacher agent, wherein the UserResponse represents a text message sent by the student; Represented by and The set of these three elements is A task description representing a new mind map specific to the type of question input by the user, describing the type of question and overall information of the task; A solution description of a new mind map specific to the type of question input by the user, summarizing the main steps and methods for solving the problem; The reasoning chain of a new mind map that represents a problem type specific to the user input describes the logical reasoning steps that need to be executed in sequence to solve this type of problem; Extract(·) represents the parsing operation function, Extract(CH History ) represents the matched conversation history data CH History Execute parsing operation function; Step 7-4: The blackboard writing schema distillation agent extracts all the blackboard writing content BW generated in the dialogue with the AI teacher agent, and organizes the series of blackboard writing content generated in the dialogue-based teaching process in the order of generation time into a time series, called blackboard writing time series BW Seq ; Among them, BW represents the blackboard content, BW Seq Represents the blackboard writing sequence sorted by the blackboard writing content in the order of generation time; the blackboard writing schema distillation agent obtains the blackboard writing sequence BW Seq Perform a distillation operation to obtain a new blackboard schema BWSchema specific to the question type entered by the user New ,satisfy: Among them, BWSchema New represents a new blackboard schema specific to the type of question input by the user, DistillBWS(·) represents the distillation operation function of the blackboard schema distillation agent, DistillBWS(BW Seq ) represents the blackboard writing time series BW of the blackboard writing schema distillation agent Seq Execute the parsing operation function; the BW Seq It means that a series of blackboard writing contents generated in the process of conversational teaching and presented in the order of generation time are organized into a time series, which is called blackboard writing sequence; BW represents the blackboard writing content, i represents the serial pointer of the blackboard writing content, and NDR represents the conversation round; BW1 represents the blackboard writing content generated by the AI teacher agent in the first round of conversation, BW2 represents the blackboard writing content generated by the AI teacher agent in the second round of conversation, and BW i represents the blackboard content generated by the AI teacher agent in the i-th round of dialogue, BW NDR Represents the blackboard content generated by the AI teacher agent in the NDR round of dialogue; Step 7-5: Create a new mind schema TSchema specific to the question type input by the user obtained through the distillation operation New and a new BWSchema specific to the type of question the user enters New Adding it to the schema buffer of the agent workflow enables the agent workflow to self-reflect and summarize new types of questions that have not been answered, enhancing the scalability of the agent workflow and satisfying: SchemaBuffer represents the schema buffer, ← represents the assignment operation symbol, and TSchema New New mind schema, BWSchema, to represent question types specific to user input New Represents a new blackboard schema specific to the type of question input by the user. TSN is the total number of thinking schemas in the schema buffer. SchemaBuffer.TSN represents the value of the total number of thinking schemas TSN obtained from the schema buffer SchemaBuffer.
6. A multi-modal conversational learning content generation device driven by blackboard generation, characterized in that: It includes a terminal device, which is an Internet terminal device, including a processor and a computer-readable storage medium, the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, and the instructions are used to be loaded by the processor and executed by a multi-modal conversational teaching content generation method driven by blackboard generation as described in any one of claims 1-6.
7. A multi-modal conversational learning content generation system driven by blackboard generation, characterized in that: It includes a computer-readable storage medium, in which multiple instructions and an agent workflow are stored. The instructions are used by a processor of a terminal device to load and execute a multi-modal conversational teaching content generation method driven by blackboard generation as described in any one of claims 1-6, and the agent workflow includes a multi-agent system module, a schema buffer module and a text-to-speech module.
Citation Information
Cited By
Teaching and research system and method based on large language model agent
CN120780835A