Method and device for processing education task based on multi-modal large model driven agent cooperation, equipment and medium
By adopting a multimodal large-scale model-driven intelligent agent collaboration method, the problems of data perception and process processing in AI education systems in multi-person collaborative scenarios are solved, realizing the effective execution and data perception of multimodal collaborative education tasks and improving the collaborative capabilities of the education system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FUYOU YILI TECHNOLOGY CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing AI education systems lack data perception, multimodal data processing, and workflow processing capabilities for collaborative tasks in multi-person collaborative scenarios, thus failing to effectively cultivate teamwork skills, and there is an upper limit to the human intervention of teachers.
A multimodal large model-driven agent collaboration method is adopted. The agent calls the multimodal large model to extract task parsing results, determine collaborators, obtain multimodal task execution information in real time, determine task completion based on execution information, and generate task summary, thereby realizing data perception and process processing of multimodal collaborative education tasks.
It enables data perception and multimodal data processing in multi-person collaborative education scenarios, enriches education scenarios, accurately executes collaborative task processes, reduces data silos, and improves the collaborative efficiency of the education system.
Smart Images

Figure CN122491760A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and medium for processing educational tasks based on multimodal large model-driven intelligent agent collaboration. Background Technology
[0002] With the rapid development of artificial intelligence technology, generative artificial intelligence has increasingly broad application prospects in the field of education.
[0003] Currently, various AI (Artificial Intelligence) education platforms and online education systems are widely used in scenarios such as assisting teaching, personalized learning path planning, intelligent homework grading, and learning behavior analysis.
[0004] However, existing AI education systems primarily operate on a one-to-one, single-point interaction model. Especially in collaborative scenarios, the system only provides educational services based on individual student information. In reality, teamwork skills also require development. Multi-person projects can only be addressed by increasing teacher involvement, but there are limits to teacher intervention. Furthermore, current AI education systems lack data perception, multimodal data processing, and workflow management capabilities for collaborative tasks. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for processing educational tasks based on multimodal large model-driven intelligent agent collaboration. It can provide a process for implementing collaborative educational tasks, increase the content of data perception in educational scenarios, improve the accuracy of multimodal data processing, and reduce data silos.
[0006] According to one aspect of the present invention, a method for processing educational tasks based on agent collaboration driven by a multimodal large model is provided, the method comprising: The project's intelligent agent invokes a multimodal large model to extract task parsing results from multimodal collaborative education tasks initiated by the target user; The project role intelligent agent invokes the large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information; When the confirmation command to start the task is detected from each participating user, the project role agent calls the large model to send the task content from the task parsing result to the target user and each of the collaborators; the participating users include the target user and the real users among the collaborators. The project role intelligent agent obtains multimodal task execution information provided by the target user and each of the collaborators in real time regarding the task content; The project role intelligent agent calls the multimodal big model in real time to determine the completion judgment result of the expected collaborative goal in the task parsing result based on the multimodal task execution information; When the completion determination result is that the task is completed, the project role agent calls the large model to obtain the collaboration logs of each participating user and generates task summary information.
[0007] According to one aspect of the present invention, an educational task processing device based on multimodal large model-driven agent collaboration is provided, the device comprising: The task parsing module is used to call the multimodal big model to extract task parsing results from multimodal collaborative education tasks initiated by the target user; The collaborator determination module is used to call the large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information; The task distribution module is used to send the task content from the task parsing result to the target user and each of the collaborators when a confirmation command to start the task is detected from each participating user; the participating users include the target user and the real users among the collaborators; The task execution module is used to acquire multimodal task execution information provided in real time by the target user and each of the collaborators regarding the task content; The target determination module is used to call the multimodal large model in real time to determine the completion determination result of the expected collaborative target in the task parsing result based on the multimodal task execution information; The task summary module is used to retrieve the collaboration logs of each participating user from the large model and generate task summary information when the completion determination result is that the task is completed.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal collaborative education task processing method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the multimodal collaborative education task processing method according to any embodiment of the present invention.
[0010] The technical solution of this invention determines the collaborators for a multimodal collaborative education task by obtaining task parsing results from a multimodal collaborative education task initiated by a target user and combining this with the target user's memory and learning information. When the actual user in the multimodal collaborative education task confirms the task initiation command, the task content is sent to each participant, and the multimodal task execution information of each participant is obtained. The multimodal task execution information is then compared with the expected collaborative goal to determine completion, resulting in task summary information. This achieves data perception and processing for the target user in the multimodal collaborative education task, and implements the collaborative process. It solves the technical problem in existing technologies that can only process data in single-person education scenarios, increases the application scenarios for collaboration, enriches education scenarios, increases the content of data perception and multimodal data in education scenarios, and can accurately process multimodal data, correctly execute collaborative task processes, and simultaneously interact with data from multiple participants, reducing data silos.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a multimodal collaborative education task processing method provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of a multimodal collaborative education task processing method provided by an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a multimodal collaborative education task processing device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device that implements the multimodal collaborative education task processing method of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Figure 1 This is a flowchart illustrating a multimodal collaborative education task processing method provided in an embodiment of the present invention. This embodiment is applicable to situations where an agent is invoked to execute a multimodal collaborative education task based on a large multimodal model. The method can be executed by a multimodal collaborative education task processing device, which can be implemented in hardware and / or software.
[0017] See Figure 1 The educational task processing method based on multimodal large model-driven agent collaboration, as shown, includes: S101, The project role intelligent agent calls the multimodal large model to extract task parsing results from the multimodal collaborative education task initiated by the target user.
[0018] Among them, the educational task processing method based on multimodal large-scale model-driven agent collaboration can be executed by an agent. The agent's function is to act as a project role, assisting the target user in completing a multimodal collaborative educational task. A multimodal collaborative educational task can be a learning scenario built based on an agent collaborative architecture during educational activities or teaching processes, requiring the joint participation, cooperation, and information sharing of two or more roles (such as teachers, students, parents, or AI assistants) to complete. The target user can refer to the real user who initiates the multimodal collaborative educational task. The task parsing result can refer to the key information extracted from the multimodal collaborative educational task. For example, the task parsing result should at least include: task content and expected collaborative goals. The structured information obtained by deep semantic parsing of the multimodal collaborative educational task using a multimodal large-scale model can be used as the task parsing result. The pedagogical attributes, interaction logic, and social dimensions of the multimodal collaborative educational task can be modeled using a multimodal large-scale model. For example, the task parsing result should at least include: task content, expected collaborative goals, role configuration requirements, and task pressure coefficient.
[0019] Task content can refer to the execution guidance information provided to participants (real people or virtual agents), allowing them to understand the task at hand, including the task theme, task type (such as debate, co-creation, and puzzle-solving), and at least one phased sequence of sub-tasks. In addition to the initiator and target user, participants in a multimodal collaborative education task must include at least one collaborator. All collaborators include real people and / or agents. That is, in a multimodal collaborative education task, besides the agent used by the target user (project role agent), real people can choose whether to use an agent (personal assistant role), and virtual people can be constructed through agents. In other words, when a real person selects an agent or a virtual person exists, multiple agents exist in the multimodal collaborative education task, which is completed through multi-agent collaborative assistance.
[0020] Collaboration goals serve as the criteria for determining whether a task has been completed, and as the logical basis for this determination. Task content may include a task topic, task type, and at least one sub-task. Correspondingly, collaboration goals can be sub-goals corresponding to each sub-task.
[0021] In some embodiments, the task parsing results may further include: collaborating user information and collaborative pressure. Collaborating user information may be attributes of the roles required for the multimodal collaborative education task. Role attributes can be determined based on task complexity. Role attributes include the number of roles, the functional positioning of each role (e.g., decision-maker, reviewer, or executor), and expected competency labels. Collaborative pressure is used to determine the type of pressure in the multimodal collaborative education task. Collaborative pressure may include: debate-based and / or co-creative.
[0022] In some embodiments, the system receives a request from a target user to initiate a multimodal collaborative education task. This request can be multimodal input such as text, voice, or image. Based on the task description information carried in the request, at least one of the following is extracted from the task description information: task content, expected collaborative goals, collaborative user information, and collaborative pressure. Alternatively, partial information (such as only the task topic and task type) can be extracted from the task description information, and a larger model can be invoked to generate the remaining information that needs to be completed (such as subtasks, expected collaborative goals, collaborative user information, and collaborative pressure). In some embodiments, the education platform system can preset multiple standard tasks. The target user can select one of these standard tasks as a multimodal collaborative education task. The target user can also modify the information in the standard task and use the modified standard task as the collaborative education task. Based on the target user's selection instruction or a combination of selection and modification instructions, the initiated multimodal collaborative education task is determined, and at least one of the following is extracted from the multimodal collaborative education task: task content, expected collaborative goals, collaborative user information, and collaborative pressure.
[0023] In one example, a target user initiates a multimodal collaborative educational task involving an English debate. The user's request specifies the task type as adversarial (an English debate scenario) and the task topic as: Should xx bring a mobile phone? The educational platform system can obtain a prompt template corresponding to the task type. This prompt template is used to extract the task parsing results. The prompt template is combined with the task topic to generate input data, which is then fed into a large model for processing. The output includes sub-tasks, expected collaborative goals, collaborative user information, and collaborative pressure. The multimodal large model can select the corresponding prompt template for parsing based on the task type. For example, in the 'English debate' scenario, it obtains the prompt corresponding to the adversarial task type; similarly, in the AI cat pet expert scenario, it obtains the prompt corresponding to the co-creation task type.
[0024] In one specific embodiment, target user A initiates a multimodal collaborative education task about English debate. User A inputs the topic: Should xx bring a mobile phone? The system calls the project manager, i.e., the project role intelligent agent M. After running, project role intelligent agent M determines that it needs to assist target user A in performing the multimodal collaborative education task. At this time, it obtains the Prompt, adds the relevant content of the multimodal collaborative education task to the Prompt to form input data, and calls the multimodal large model to provide the input data for processing, outputting the task parsing result.
[0025] The prompt template could be: You are M, a senior AI project architect. Please analyze the topic: [Should xx bring a mobile phone?]
[0026] Task requirements: Core Dimension Extraction: Deconstructing the debate entry points from dimensions such as security, learning interference, and social development.
[0027] Ability threshold setting: Based on the user's [target user's] memory and learning, set the language proficiency target for this debate (e.g., need to master concession and transition sentences such as 'xx').
[0028] Interaction Mode Determination: This task is determined to be a high-conflict, confrontational task, requiring the assignment of a meticulous opponent B and a guiding teammate C.
[0029] Phase Breakdown: The process is divided into three phases: argumentation, cross-examination, and closing argument.
[0030] Output format: Please output the subtasks, expected collaboration goals, role configuration requirements (personality parameter matrix) (collaboration user information) and task pressure coefficient (collaboration pressure) in structured JSON format.
[0031] The education system can transform vague educational tasks into a set of structured instructions that can be precisely executed by the education system. This solves the technical problems of traditional systems being unable to understand the depth of collaboration and unable to dynamically adjust strategies according to task type, laying a data foundation for the precise collaboration of intelligent agents in the future.
[0032] S102. The project role intelligent agent calls the big model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information.
[0033] The task analysis results are used to determine the collaborators required for the multimodal collaborative education task. Memory learning information refers to multidimensional feature vectors of the target user's ability dimensions, cognitive preferences, and social habits recorded in the education system. Typically, memory learning information can be learning-related information authorized by the target user or their guardian. The target user's memory learning information is used to determine the collaborators required by the target user. Combining the task analysis results and the target user's memory learning information to determine collaborators allows for a comprehensive consideration of both task and user dimensions, simultaneously adapting the task and user to determine collaborators. This achieves data-driven collaborator completion, avoids simple combinations, and improves the accuracy of the participation data required for the task, ensuring the correct and smooth execution of the multimodal collaborative education task. Collaborators can include real users or virtual users. Virtual users can be intelligent agents assigned specific roles, permissions, and abilities. The role settings of intelligent agents can be determined based on the memory learning information and the collaborating user information in the task analysis results. Virtual users can be intelligent agents specifically configured for the multimodal collaborative education task. Virtual users can provide interactive content required by the task content in the task analysis results and consistent with their role settings to interact with the target user, ensuring the sequential execution of the multimodal collaborative education task. In reality, real users who perform multimodal collaborative education tasks need to be online and interact in real time. Therefore, if there are no users online who can participate in multimodal collaborative education tasks, virtual users can be added to make the multimodal collaborative education tasks executable.
[0034] In some embodiments, at least one collaborator is determined to match the multimodal collaborative education task based on role attributes in the task analysis results and the target user's memory learning information. The project role agent invokes the large model to perform bidirectional matching in an online pool of real users and a pool of virtual agents based on role attributes and the target user's memory learning information. The virtual agents are virtual users dynamically assigned specific personality parameters by the large model based on the task analysis results.
[0035] The education system can utilize a multi-agent framework to generate multiple agents. For example, in a multimodal collaborative educational task involving English debate, the project-role agent M is responsible for global scheduling, monitoring interaction status, distributing instructions, and executing intervention logic. If collaborators include virtual users, these virtual users are agents. For instance, an AI adversary agent (B) can be constructed: agent B's personality parameters include rigor and argumentativeness, responsible for executing adversarial pressure tests to force the target user to logically counterattack; an AI teammate agent (C) can also be constructed: agent C's personality parameters include guidance and encouragement, responsible for providing scaffolding and supplementary arguments. Each agent corresponds to an independent LLM instance. The project-role agent M injects differentiated personality parameter matrices of role attributes into different instances based on task requirements. For example, the prompt input to agent B might inject logic focused on identifying logical flaws, while the prompt input to agent C might inject logic focused on semantic inspiration and emotional support. This multi-instance concurrent architecture ensures the realism of the interaction and the relevance of the educational objectives. When there are insufficient real online users to fill the role slots, the project role agent M dynamically instantiates the corresponding agent as a collaborator based on the target user's capability disadvantage information, thus ensuring the continuity of the collaboration flow.
[0036] S103. When the confirmation command to start the task is detected by each participating user, the project role agent calls the large model to send the task content in the task parsing result to the target user and each of the collaborators; the participating users include the target user and the real users among the collaborators.
[0037] The "Confirm Start Task" command confirms that all collaborators and the target user are ready to begin executing the multimodal collaborative education task. The task content can be a description of the task, including actions, quality requirements, and workflow. The target user and collaborators can input multimodal task execution information according to the task content and workflow. Participating users are real users. It is necessary to determine whether participating users can trigger the start of the multimodal collaborative education task. The task begins only after all real users have issued the "Confirm Start Task" command. In some embodiments, after the multimodal collaborative education task starts, initialization is required, calling the large model to assign roles to the target user and collaborators. For example, a data collection role is assigned to the target user, and a data annotation role is assigned to the collaborators. After role assignment, the task content corresponding to each role is sent to the target user and each collaborator. The project role agent can send role-appropriate task content to the target user and each collaborator.
[0038] In one example, the prompt template for role assignment is: You need to assign roles to a multimodal collaborative education task. The members are target user A, online user B, and virtual user C. Their respective attributes (personality and abilities, etc.) are xx. The collaborative user information is xx. Please assign roles to A, B, and C. Output the roles of A, B, and C.
[0039] In some embodiments, the task content distribution process is driven by the project role agent M: M generates differentiated distribution content for each role based on the current stage goals and multimodal context.
[0040] For example, the initialization prompt template for the project role agent M is: You are an AI project architect M, who needs to analyze user multimodal input and provide timely feedback. During project collaboration, you need to guide users to engage in more communication and sharing, ensuring the team moves closer to the expected collaborative goals and ultimately achieve project success. The task theme is xx, and the task type is xx. The team configuration is 4 participants. Other members A are the target user, with relevant information and a role function of xx; member B is a participating user, with relevant information and a role function of xx; member C is a virtual human, with a virtual personality of xx and a role function of xx.
[0041] The prompt template for determining the content to be distributed is as follows: Based on the context of the multimodal collaborative education task (xx), the current stage (xx), and the corresponding task content (xx), determine the content to be distributed to each member based on their respective role information. Output the content to be distributed to each member, as well as the transmission interface for each member.
[0042] The input data generated from the project manager's initial prompt template and the prompt template for the assigned content is fed into the large model to obtain the assigned content and transmission interface for each member. The transmission interface is then called to transmit the assigned content to the corresponding member. This agent-driven dynamic assignment mechanism ensures that the task instructions received by each member not only conform to the overall goal but are also highly aligned with their assigned role attributes. This role-instruction two-way binding model solves the problems of single task distribution and lack of personalized role interaction in traditional education platforms, significantly enhancing the accumulated value of collaborative tasks and user engagement.
[0043] S104. The project role intelligent agent obtains in real time the multimodal task execution information provided by the target user and each of the collaborators for the task content.
[0044] Multimodal task execution information refers to the content related to the execution operations of the task content input by the target user and collaborators. It can be obtained by collecting full multimodal interaction data from collaborators during the collaboration process. Multimodal task execution information can include data of at least one type (modality) such as text, audio, images, video, and operation behavior logs. The content related to the execution operations can include the execution content, execution process, and execution results. Multimodal task execution information needs to be synchronized with the target user and all collaborators. Multimodal task execution information can also include interaction information between the target user and all collaborators. The project role intelligent agent M monitors and synchronizes this multimodal information in real time, aligning execution intent and semantics through a large model. For example, if a task requires multiple people to work together, the target user waits for the collaborators to input multimodal task execution information, and then inputs their own multimodal task execution information based on the collaborators' input and the task content. Alternatively, the target user can interact with the collaborators, suggest adjustments, and input new (adjusted) multimodal task execution information. Based on this, the target user inputs their own multimodal task execution information, combining the new multimodal task execution information with the task content.
[0045] S105. The project role intelligent agent calls the multimodal big model in real time to determine the completion judgment result of the expected collaborative goal in the task parsing result based on the multimodal task execution information.
[0046] The completion determination result is used to determine whether the expected collaborative goal has been achieved, thereby judging whether the multimodal collaborative education task has been completed. In some embodiments, the multimodal collaborative education task includes at least one stage. In each stage, the education system invokes an intelligent agent acting as the collaboration facilitator to send at least one sub-task to be executed in the current stage to the target user and at least one of the collaborators. Each sub-task is sent to the target user and each collaborator sequentially according to the order and triggering conditions. For example, sub-tasks sent first are sent first, and when the first sub-task is determined to be completed, subsequent sub-tasks are sent. After each sub-task is sent, the target user or collaborator executes the sub-task and inputs response information. All response information sent by the target user or collaborator is determined as the multimodal task execution information. That is, the multimodal task execution information is continuously updated as the multimodal collaborative education task progresses. Based on the subtasks and corresponding sub-goals, the multimodal task execution information (response information) input by the target user and collaborators in real time is used to determine whether the subtask is completed. After the determination of completion, the next subtask is sent until all subtasks in the current stage are completed. At this time, the next stage is entered and the new stage's subtasks are sent until all stages are completed.
[0047] The project management agent determines goal completion in real time based on multimodal task execution information obtained through real-time monitoring. Task content can include at least one sub-task, which can be a phased task. The collaborative expected goal can include at least one sub-goal, which can be a phased goal. The agent can determine the completion status of the sub-tasks in the current phase based on the semantic matching degree between the multimodal task execution information and the sub-goals. When all sub-tasks in the current phase are completed, the agent automatically triggers the distribution of sub-task content for the next phase, continuing until all sub-goals are completed. When all sub-goals are completed, the collaborative expected goal is considered achieved.
[0048] The project's intelligent agent M acts as the evaluation center. M performs semantic comparisons between real-time perceived multimodal task execution information and the current sub-goals within the pre-defined collaborative expectation goals. The multimodal collaborative education task is dynamically broken down into several key phases, each corresponding to a set of sub-goals and corresponding ending rules. M obtains the prompt word templates for the goal completion judgment stage, generates input data, and calls the large multimodal model, enabling M to perform phase-based verification. You are a task evaluation expert. The current task stage is [Argumentation Stage], and the subtask content is [User A needs to propose two core arguments and obtain confirmation from C]. Current execution information: [Real-time perceived multimodal task execution information, such as the target user A's input text being 'xx', and the behavior log showing that they uploaded supporting documents]. Judgment requirements: 1. Calculate the semantic matching degree between the current text and the target; 2. Detect whether there is a logical loop (i.e., whether C has given effective feedback); 3. If the matching degree is greater than the threshold of 0.8, determine that the subtask is [Completed], otherwise trigger the [Guiding Branch]. Output: Stage completion status, semantic deviation analysis, and suggestions for triggering the next subtask.
[0049] When the completion determination result is "incomplete," M uses a large model to dynamically generate guidance content based on semantic deviation analysis, instructing participating users to complete the missing logic. When the completion determination result is "complete," M automatically executes a stage transition, distributing sub-task instructions for the next stage. This dynamic determination mechanism based on large model logical reasoning ensures the rigor and adaptability of the collaborative process.
[0050] S106. When the completion determination result is that the task is completed, the project role intelligent agent calls the large model to obtain the collaboration logs of each participating user and generates task summary information.
[0051] Collaboration logs refer to the learning behavior information input by real users during multimodal collaborative education tasks. Task summary information is used for task evaluation, feedback, and recording of each real user's learning behavior. The target user's relevant information in the task summary information can be added to the target user's memory learning information, enriching the target user's perceptual and multimodal data in the educational scenario, so as to adjust the task content of the multimodal collaborative education task based on the memory learning information. Collaboration logs and summary prompt templates can be combined to generate input data, which is then fed into a large model for processing to output task summary information. For example, the task summary information might be: Target user A contributed 15 high-quality images; User C corrected 3 logical errors. The collaboration efficiency between the two increased by 20%. The summary prompt template might be: You are an intelligent collaborative teaching assistant proficient in educational psychology and data analysis. You are generating a summary report for target user A for a multimodal collaborative education task. The recorded information for this task is {collaboration log}. Target user A's historical level is {memory learning information}. The evaluation rule is xx. Task Requirements: The summary must include: key statement quotations, quantified progress, specific behavioral performance, and results. Output format: xx.
[0052] By constructing a multi-agent collaborative network, a real social learning environment was simulated, addressing the technical pain points of existing AI education tools, such as the lack of social interaction and insufficient depth of interaction. Utilizing the generative capabilities of large models, dynamic generation of teaching scaffolding was achieved, enabling millisecond-level adjustment of guidance content based on real-time user feedback, thereby realizing personalized education.
[0053] The technical solution of this invention determines the collaborators for a multimodal collaborative education task by obtaining task parsing results from a multimodal collaborative education task initiated by a target user and combining this with the target user's memory and learning information. When the actual user in the multimodal collaborative education task confirms the task initiation command, the task content is sent to each participant, and the multimodal task execution information of each participant is obtained. The multimodal task execution information is then compared with the expected collaborative goal to determine completion, resulting in task summary information. This achieves data perception and processing for the target user in the multimodal collaborative education task, and implements the collaborative process. It solves the technical problem in existing technologies that can only process data in single-person education scenarios, increases the application scenarios for collaboration, enriches education scenarios, increases the content of data perception and multimodal data in education scenarios, and can accurately process multimodal data, correctly execute collaborative task processes, and simultaneously interact with data from multiple participants, reducing data silos.
[0054] In an optional embodiment, the step of determining the completion judgment result of the collaborative expected goal in the task parsing result based on the multimodal task execution information in real time includes: obtaining the interaction content of the current stage in the multimodal task execution information, at least one target content of the current stage in the collaborative expected goal, and the end judgment rule of each target content; for each target content, matching the interaction content of the current stage with the target content according to the end judgment rule of the target content to obtain the target completion detection result of the target content; when the target completion detection result of each target content in the current stage is complete and the current stage is the final stage of the multimodal collaborative education task, determining the completion judgment result of the collaborative expected goal as complete; when the target completion detection result of each target content in the current stage is complete, determining the completion judgment result of the collaborative expected goal as complete. When all content completion detection results are complete and the current stage differs from the final stage of the multimodal collaborative education task, the system proceeds to the next stage and sends the sub-task content of the next stage to the target user and each collaborator, while determining that the completion judgment result of the expected collaborative goal is incomplete. When the target completion detection result of some target content is incomplete, target guidance content is generated based on the difference between the interactive content of the current stage and the incomplete target content, and sent to the target user to enable the target user to provide new interactive content, while determining that the completion judgment result of the expected collaborative goal is incomplete. The incomplete target content is matched with the new interactive content, and the target completion detection result of the incomplete target content is updated.
[0055] The multimodal collaborative education task process can be divided into at least one stage, essentially breaking down a collaborative task into executable and evaluable steps. The expected collaborative goal includes the target content of all stages and the termination criteria for each target content. The interaction content of the current stage can refer to the input from the target user and collaborators in the current stage. The target content of the current stage can refer to the task objective of the current stage. The target content of the current stage can include at least one sub-task. The termination criteria can refer to the rules for determining whether the target user and collaborators have completed the target content of the current stage. Matching the interaction content with the target content is equivalent to determining whether the existing interaction content meets the task objective of the current stage. The goal completion detection result can refer to the detection result for determining whether the target content has been completed. If all target content of the current stage has been completed, it indicates that the multimodal collaborative education task can proceed to the next stage. If there is no next stage, i.e., the current stage is the final stage of the multimodal collaborative education task, the completion judgment result of the expected collaborative goal is determined as task completion. If there is a next stage, the completion judgment result of the expected collaborative goal is determined as incomplete. If all objectives in the current stage are not completed, the completion result of the expected collaborative objective is determined to be incomplete. In this case, further assistance is needed until all objectives in the current stage are completed. Only when all objectives in all stages show a completion result, is the expected collaborative objective determined to be complete. If any objective shows a failure to complete its objective, the expected collaborative objective is determined to be incomplete.
[0056] In some embodiments, a fault tolerance mechanism can also be configured. When the execution time of a certain stage is greater than or equal to the upper limit time and the number of auxiliary attempts is greater than or equal to the upper limit number of attempts, the target completion detection result of the target content can be determined as completed, and other target content in the current stage can be continued or the next stage can be entered to avoid getting stuck on a certain target content or a certain stage.
[0057] The difference between the current interactive content and the incomplete target content can refer to the data that participating users need to supplement to complete the target content. Target guidance content is provided to participating users to guide them in inputting the data required to address the difference, so that the target completion detection result for the incomplete target content is considered complete. Target guidance content may include: the content of the data that needs to be supplemented, the method of obtaining the supplemented data, the content of the data that needs to be modified, the modification method, and examples of correct interactive content. New interactive content can refer to new data input by participating users based on the target guidance content. For example, new interactive content may include newly provided operation data by participating users, modified operation data by participating users, newly input interactive content by participating users, and modified interactive content by participating users. Matching of incomplete target content continues based on the new interactive content, updating the target completion detection result for that incomplete target content. If the target completion detection result is complete, guidance content is generated based on other incomplete target content; if no incomplete target content exists, the process proceeds to the next stage. If the target completion detection result is incomplete, new target guidance content is generated based on the context and provided to participating users. This continues until the target completion detection result for the incomplete target content is considered complete. A fault-tolerance mechanism can be introduced to count the number of times the target guidance content for the same target content is generated. If the number of generation times, i.e. the number of auxiliary times, is greater than or equal to the upper limit, the correct interactive content for the target content is sent to the participating users, and the target completion detection result of the target content is determined to be completed, so as to continue the process of the multimodal collaborative education task and continue to the next stage or complete the expected collaborative goal.
[0058] In some embodiments, each received statement is added as multimodal task execution information to the context of the multimodal collaborative education task, and to the context of sub-tasks that can be further subdivided into stages. The received statements are combined with decision prompt templates to obtain input data, which is then fed into a larger model for processing, outputting the current stage and the decision result.
[0059] The prompt template for the judgment is as follows: Based on the context of the multimodal collaborative education task (xx), determine the current stage and target content, and query the termination judgment rule corresponding to the target content. Based on the termination judgment rule and the multimodal task execution information, determine the target completion detection result of the target content and the target completion detection result of the current stage. Output the current stage, target content, target completion detection result of the target content, and target completion detection result of the current stage.
[0060] The output needs to be added to the context of the multimodal collaborative education task.
[0061] Once the current stage's objective is determined to be complete and this is not the final stage, the content to be delivered in the next stage is determined. A prompt template for the delivered content can be used as a reference to generate the content for the next stage, and the corresponding interface can be called to transmit the content for the next stage.
[0062] When the current stage's objective completion detection result is determined to be completed and the current stage is the final stage, the completion judgment result of the collaborative expected objective is determined to be completed.
[0063] In a scenario involving a collaborative education task for training a cat recognition model, target user A uploads photos of cat behavior (eating, sleeping, and destroying things, etc.), using these photos as multimodal task execution information. This multimodal task execution information is synchronized to all members. Online user B labels the photos, sets recognition logic, and uploads relevant information. The received information from online user B is synchronized to all members. Target user A interacts with online user B's information. The system continuously receives interaction information from target user A and online user B, using it as multimodal task execution information. Based on the received multimodal task execution information and the termination rules for each unfinished objective in the current stage, the system determines whether the unfinished objective has been completed. For example, termination rules might include: a quantity threshold rule (the number of valid samples in each category reaches a preset value, such as 5); a semantic alignment rule (verifying whether the image content provided by A is logically consistent with the labels set by B, e.g., not labeling a cat biting the sofa as sleeping); and a collaboration verification rule (at least 3 interaction records between A and B regarding feature definitions).
[0064] In another scenario, virtual human C simulates learning based on training data jointly determined by A and B, and provides feedback on the initial accuracy. A and B provide test images and observe C's classification results. For example, termination rules include: Performance achievement rule: C's simulated classification accuracy reaches a preset threshold (e.g., >80%). Closed-loop feedback rule: If the accuracy is insufficient, A and C must complete at least one error attribution dialogue (e.g., "A, the lighting in this image is too dark; B can't recognize it. We need to reshoot"). Logical confirmation rule: Determining that the two have reached a consensus, confirming that the model can be delivered.
[0065] As can be seen, the entire process of multimodal collaborative education tasks is realized through goal decomposition, rule matching, response judgment, and incomplete guidance. This achieves a data-driven task flow, enabling task traceability, interruption, and adjustment. Furthermore, based on event-driven process control, each stage of the task can be executed precisely. At the same time, fine-grained data is collected from each subtask in each stage to enrich the user's learning data for locating learning problems. In addition, the addition of guidance content makes task execution more stable.
[0066] Figure 2 This is a flowchart illustrating a multimodal collaborative education task processing method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention calls a large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information. The method includes: when the number of participants in the multimodal collaborative education task is determined to be less than the standard number in the task parsing results, obtaining the target user's memory learning information; calling the large model to query the current online users for a first user corresponding to the multimodal collaborative education task based on the task parsing results and the target user's memory learning information; when a first user exists among the current online users, initiating a collaboration request to each of the first users; and determining the first user who confirms participation as a collaborator for the multimodal collaborative education task. It should be noted that parts not detailed in this embodiment of the present invention can be found in the descriptions of other embodiments.
[0067] See Figure 2 The multimedia presentation generation method shown includes: S201. The project role intelligent agent calls the multimodal large model to extract task parsing results from the multimodal collaborative education task initiated by the target user.
[0068] S202. When it is determined that the number of participants in the multimodal collaborative education task is less than the standard number of participants in the task analysis result, the project role agent obtains the memory learning information of the target user.
[0069] The standard number of participants can be the number configured in the collaborative user information within the task parsing results. After initiating a multimodal collaborative education task, the target user can send task invitation requests to at least one selected online user. If the invited user agrees to participate, a confirmation response is sent. Following the order of agreement, invited users are sequentially designated as collaborators, and the total number of participants in the multimodal collaborative education task is accumulated. When the number of participants equals the standard number, the process of sending task invitation requests stops, and new confirmed invited users are no longer designated as collaborators. If the number of agreed invited users is less than the standard number, additional users need to be added to the multimodal collaborative education task.
[0070] S203. The project role intelligent agent calls the large model to query the first user corresponding to the multimodal collaborative education task among the currently online users based on the task parsing results and the memory learning information of the target user.
[0071] In this process, the project's intelligent agent invokes a large model to extract features from the task parsing results, obtaining a first vector. It then extracts features from the memorized learning information, obtaining a second vector. The first and second vectors are fused to obtain a fused vector, which is then compared with vectors extracted from the attributes and / or memorized learning information of online users. At least one similar online user is selected as the first user. If an online user exists, the first user is queried from among them. If no online user exists, the first user is determined not to exist. If none of the online users correspond to the multimodal collaborative education task, the first user is determined not to exist. In some embodiments, a similarity value between the online user's vector and the fused vector greater than or equal to the similarity value indicates that the online user corresponds to the multimodal collaborative education task. A similarity value between the online user's vector and the fused vector less than the similarity value indicates that the online user does not correspond to the multimodal collaborative education task.
[0072] S204. When there is a first user among the currently online users, the project role agent initiates a collaboration request to each of the first users.
[0073] The collaboration request is used by the first user to determine whether they can participate in the multimodal collaborative education task.
[0074] S205. The project role intelligent agent will identify the first user who confirms participation as the collaborator for the multimodal collaborative education task.
[0075] The feedback confirming the participation of the first user indicates that the first user intends to participate in the multimodal collaborative education task.
[0076] S206. When the confirmation command to start the task is detected by each participating user, the project role agent calls the large model to send the task content in the task parsing result to the target user and each of the collaborators; the participating users include the target user and the real users among the collaborators.
[0077] S207. The project role intelligent agent obtains in real time the multimodal task execution information provided by the target user and each of the collaborators regarding the task content.
[0078] S208. The project role intelligent agent calls the multimodal big model in real time to determine the completion judgment result of the expected collaborative goal in the task parsing result based on the multimodal task execution information. S209. When the completion determination result is that the task is completed, the project role intelligent agent calls the large model to obtain the collaboration logs of each participating user and generates task summary information.
[0079] The technical solution of this invention, when the number of participants is less than the standard number of participants for a multimodal collaborative education task, queries the online user for a first user and initiates a collaboration request to the first user. The first user who confirms participation is identified as a collaborator. This can complete the collaborative participation of the multimodal collaborative education task. At the same time, querying the first user based on memory learning information and task parsing results can filter out users who are accurately matched to the multimodal collaborative education task, improve the accuracy of the filtering, and better assist the target user in completing the collaborative task.
[0080] In an optional embodiment, the step of querying the first user corresponding to the multimodal collaborative education task among the current online users based on the task parsing result and the target user's memory learning information includes: determining the target user's ability disadvantage information and social conflict information based on the target user's memory learning information, and using this as target matching information; querying at least one second user corresponding to the target matching information among the current online users; determining a third user corresponding to the task parsing result among each second user based on the learning path of each second user; and determining the first user based on each third user.
[0081] In this process, capability disadvantage information is used as attribute information for filtering the first user based on desired fulfillment. Social conflict information is used as relationship information with the target user for filtering the first user based on desired fulfillment. Capability disadvantage information can refer to attribute information regarding the target user's lack of ability in multimodal collaborative education tasks. Social conflict information can refer to information regarding negative interaction relationships with the target user. Vectors can be generated for capability disadvantage information and social conflict information, and these two vectors are concatenated to form target matching information. A second user with a similarity greater than or equal to a similarity threshold to the target matching information is then queried among online users. The second user can refer to online users whose capabilities are complementary to the first user's in multimodal collaborative education tasks and who have a good social fit.
[0082] The learning path refers to information such as the progress, type, sequence, and time of courses associated with multimodal collaborative education tasks that a user learns on the education platform. A third user refers to a second user whose learning path aligns with the target user's. For example, the target user initiates a multimodal collaborative education task to achieve a specific learning objective, such as after-class practice for a particular course. Alignment of the third user's and target user's learning paths indicates that the third user also has the same or similar learning objective within the same time period. Specifically, the associated learning objective is determined based on the task content in the task parsing results, and it is checked whether the associated learning objective matches the second user's current objective. For example, if the subject type of the associated learning objective matches the subject type of the current objective, and the progress of the associated learning objective matches the current objective, it is determined whether the learning objective associated with the multimodal collaborative education task matches the second user's current objective; if the subject type of the associated learning objective does not match the subject type of the current objective, or the progress of the associated learning objective does not match the current objective, it is determined whether the learning objective associated with the multimodal collaborative education task does not match the second user's current objective. The matched second user is then designated as the third user.
[0083] In a specific example, the education platform system acquires the target user's memory and learning information, and selects second users based on the principles of complementary abilities and social compatibility. For instance, if target user A excels at collecting images but is slightly weaker in logic, the system prioritizes online users with high labeling accuracy and strong logical attributes. Simultaneously, it avoids users with conflicting history (or conflicts with A), and prioritizes online users with tags indicating a tendency towards mutual assistance. For each second user, the system acquires their learning courses, compares the learning courses corresponding to the task content with the second user's learning courses, and identifies matching second users as third users, selecting them based on the principle of aligning learning paths.
[0084] In some embodiments, task negotiation requests can be sent to each third user, received by the learning assistant agent of the third user, and the agent determines whether the third user corresponds to the multimodal collaborative education task based on the task negotiation request. For example, the collaboration time of the multimodal collaborative education task can be extracted from the task negotiation request, and it can be detected whether the collaboration time conflicts with the third user's current learning task. The third user's current free time period can be determined based on the current learning task. If the collaboration time belongs to the current free time period, it is determined that there is no time conflict. If the collaboration time does not belong to the current free time period or the third user does not have a current free time period, it is determined that there is a time conflict. When there is no time conflict, the expected result is extracted from the task negotiation request, and it is detected whether the associated expected result matches the third user's current learning plan. For example, if the expected result is consistent with the ability improvement goal of the current learning plan and the expected result is consistent with the plan type of the current learning plan, it is determined that the associated expected result matches the third user's current learning plan; if the expected result is inconsistent with the ability improvement goal of the current learning plan or the expected result is inconsistent with the plan type of the current learning plan, it is determined that the associated expected result does not match the third user's current learning plan.
[0085] In one example, the negotiation process with the third user's personal agent is as follows: The education platform system, which can be specifically the target user's personal agent (or the aforementioned project manager M role in the large model or intelligent agent), sends a task negotiation request to the third user. The third user's personal agent receives the task negotiation request and extracts information about the multimodal collaborative education task. It obtains a task negotiation prompt template, combines the third user's memorized learning information and learning plan (current learning task) with the information of the multimodal collaborative education task indicated by the task negotiation request, generates input data, and inputs this input data into the large model for processing, outputting the negotiation result. The third user's personal agent then feeds back the negotiation result to the education platform system or the target user's personal agent. When the negotiation result shows that the third user corresponds to the multimodal collaborative education task, the third user is pre-approved, and at this point, the third user can be designated as the first user.
[0086] It is evident that by initially screening online users through the memorized learning information on the target user's ability disadvantages and social conflicts, a second user suitable for the target user is identified. The second user's learning path is then matched with the task analysis results to conduct a second screening, resulting in a third user. This process of multiple screenings of online users improves the accuracy of user selection.
[0087] In an optional embodiment, after determining the first user who confirms participation as a collaborator in the multimodal collaborative education task, the method further includes: when it is determined that the number of participants in the multimodal collaborative education task is less than the standard number in the task parsing result and the current time meets the collaboration invitation time limit, determining at least one virtual personality based on the task parsing result and the target user's memory learning information; and determining a virtual person based on each virtual personality, and using them as a collaborator in the multimodal collaborative education task.
[0088] The collaboration invitation time limit can refer to the maximum duration for waiting for real users to confirm their participation. The current time meeting the collaboration invitation time limit indicates that the waiting period for real users to confirm their participation in the multimodal collaborative education task is too long. The virtual personality can refer to the personality of a virtual user. Multiple candidate personalities can be preset. Based on the fusion vector of the task parsing results and the memorized learning information, the virtual personality corresponding to the fusion vector is selected from the candidate personalities. For each virtual personality, role information is generated for that virtual personality, and a virtual person is determined based on that role information. The role information of this virtual person is the role information of the virtual personality. A virtual person can be determined by an intelligent agent assigned that virtual personality. The intelligent agent assigned that virtual personality performs operations consistent with that virtual personality. When the intelligent agent corresponding to the virtual person outputs multimodal task execution information according to the assigned sub-task content, it can call the multimodal large model and add the virtual personality to the prompt word template used to generate multimodal task execution information, as well as inputting it to the multimodal large model, so that the multimodal large model generates multimodal task execution information consistent with that virtual personality based on the input data with the added virtual personality. In some embodiments, the virtual personality may include: Specifically, when a collaborator participates in a multimodal collaborative education task, their input data can be obtained by adding their role information and task content to a prompt template generated by the virtual human. This input data is then fed into a larger model for processing, outputting the collaborator's output. This output serves as the collaborator's multimodal task execution information based on the task content.
[0089] In one example, the collaborative pressure, expected collaborative goals, and capability disadvantages extracted from the target user's memory and learning information are obtained from the task analysis results and added to the prompt word template of the virtual personality. This generates input data, which is then fed into a larger model for processing, outputting the virtual personality. Additionally, when a real user is identified as a collaborator, their capability information can be added to the prompt word template. Prompt word template: Since a real user is currently unavailable, please generate a virtual personality for a virtual user. Requirements: Target user's capability disadvantages are: xx. Expected collaborative goals are: xx. Collaboration pressure is: xx. Output the virtual personality.
[0090] As in the previous example, in the prompt template for role assignment, a virtual personality is added to the attribute slot of virtual human C. For example, the role corresponding to the virtual personality (gentle guidance) (teammate of target user A) is: You are A's debate partner C. Core responsibility: Collaborate with A to express their views. Interaction strategy: 1. When A finishes expressing their views, repeat A's views in a more native way. 2. Be responsible for supplementary remarks: I agree with A, and I want to add... 3. Monitor A in real time and provide them with reference to their views. Another example is the role corresponding to the virtual personality (rigorous logic) (opponent of target user A): You are A's opponent B. Core responsibility: Rebut A's views. Difficulty control: 1. Identify A's CEFR level. 2. Throw a logical trap that is 5% higher than A's current level (e.g., a mobile phone can search for information, but how to prevent addiction to games?). Personality: Serious, objective.
[0091] When a virtual human receives task content, the process of determining multimodal task execution information is as follows: The task content, virtual personality, aforementioned role information, and prompt word templates for multimodal task execution information are combined to generate input data. This input data is then fed into a large model for processing, outputting the virtual human's multimodal task execution information. The input data for multimodal task execution information could be: You are a participant in a multimodal collaborative education task, and you need to fully understand the multimodal collaborative education task, as well as the input from other members during the collaboration process. The function of the assigned role is xx, and you communicate with members regarding stage results and feedback. The task topic is xx, the task type is xx, team members include target user A, relevant information about the target user, the function of the assigned role is xx, online user B, relevant information about B, the function of the assigned role is xx.
[0092] It is evident that by supplementing virtual participants when there are insufficient participants and the invitation period has expired, the number of participants in a multimodal collaborative education task can be quickly replenished, thereby improving the execution efficiency of the task.
[0093] In this embodiment of the invention, the priority of collaborators, from high to low, is as follows: target user manually invites friends, third-party user screening, and virtual personality generation.
[0094] In an optional embodiment, after sending the task content from the task parsing result to the target user and each of the collaborators, the method further includes: determining the current task execution state based on the multimodal task execution information provided by the target user and the collaborators; when the current task execution state meets the intervention scenario, outputting intervention content to the target user according to the intervention type corresponding to the current task execution state; when the current task execution state meets the guidance scenario, outputting guidance content to the target user according to the guidance type corresponding to the current task execution state.
[0095] Specifically, the project role intelligent agent calls the large model to determine the current task execution status based on the multimodal task execution information provided by the target user and collaborators; when the current task execution status meets the intervention scenario, it outputs intervention content to the target user according to the intervention type corresponding to the current task execution status; when the current task execution status meets the guidance scenario, it outputs guidance content to the target user according to the guidance type corresponding to the current task execution status.
[0096] The current task execution status is used to determine whether auxiliary functions are needed. Auxiliary functions include intervention or guidance. Intervention is used to correct the multimodal task execution information input by participating users, forcibly pulling the process back from its course and promoting the execution of the multimodal collaborative education task. Guidance is used to prompt participating users to input correct multimodal task execution information. The current task execution status can be periodically checked. Intervention scenarios, at least one intervention type of the intervention scenario, guidance scenarios, and at least one guidance type of the guidance scenario can be pre-configured, along with judgment rules for each intervention type and guidance type. When the current task execution status meets the judgment rules for the intervention type, intervention content is output to the target user to correct errors, resolve conflicts, and restore order, etc. When the current task execution status meets the judgment rules for the guidance type, guidance content is output to the target user to inspire the input of correct multimodal task execution information. Intervention content corresponds to intervention type. Guidance content corresponds to guidance type. In some embodiments, guidance content is determined based on the guidance type corresponding to the current task execution status and the target user's memory learning information.
[0097] For example, intervention types include arguing, going off-topic, and rule violations; guidance types include getting stuck (pausing), expressing oneself in a simplistic way, and logical breaks. Intervention content is content constrained by rules. Guidance content is content that inspires knowledge. In some embodiments, the project role agent invokes a multimodal large model, inputting multimodal data such as text and acoustic features to determine the intervention scenario and type. Multimodal data can increase the basis for judgment, achieving high-accuracy intervention triggering.
[0098] In some embodiments, if the multimodal task execution information includes uploaded data that is completely unrelated to the project (e.g., classifying plants but uploading pictures of wild animals and fruits), or discussing school anecdotes, then the multimodal task execution information deviates from the task topic in the task content. In this case, the current task execution state is determined to be off-topic, and the corresponding intervention type is "off-topic." If the multimodal task execution information consists of prolonged arguments or insults, then the current task execution state is determined to be "arguing," and the corresponding intervention type is "arguing."
[0099] In some embodiments, if the multimodal task execution information has not been updated for a long time, the current task execution state is determined to be silent, and the corresponding guidance type is a stuck / broken type. If the multimodal task execution information contains grammatical errors, the single-signature task execution state is determined to be grammatically incorrect, and the corresponding guidance type is an error type. Based on the target user's multimodal task execution information and memory learning information, the target user's ability evaluation result can be determined, and guidance content can be generated based on the ability evaluation result. The guidance content can change in real time as the target user's ability is updated. If guidance is successful: the target user inputs high-quality content, and the task execution process continues. If guidance fails (secondary stuck): the current task content is judged to be too difficult, and a dimensionality reduction branch is executed: the task content is broken down into smaller sub-steps, or a virtual human is used, or a correct answer example is directly output.
[0100] As can be seen, by determining the current task execution status in real time based on multimodal task execution information, when it is determined to be in an intervention scenario, the corresponding intervention content is output to correct errors, resolve conflicts, and restore order. When it is determined to be in a guidance scenario, the corresponding guidance content is output to inspire the input of correct multimodal task execution information. This can ensure that collaborative education tasks are executed correctly, avoid redundant computation caused by users abandoning the task midway, and thus save computing resources. At the same time, it can make the input data more standardized, thereby improving response speed and reducing the risk of crashes.
[0101] In an optional embodiment, determining the current task execution state based on the multimodal task execution information provided by the target user and the collaborator includes: obtaining the target user's input state based on the multimodal task execution information provided by the target user and the collaborator, the input state including: silence duration and input semantics; determining the target user's current task execution state as a stuck / broken type in the guidance scenario when the target user's input state meets the stuck / broken condition; calculating the semantic difference between the target user's input semantics and the task content; determining the target user's current task execution state as a deviating type in the intervention scenario when the target user's semantic difference meets the off-topic condition; obtaining the target user's input content based on the multimodal task execution information provided by the target user and the collaborator, and identifying the target user's emotional content; determining the target user's current task execution state as an arguing type in the intervention scenario when the target user's emotional content meets the negative emotional condition; performing contextual repetition detection on the target user's input content to obtain repetition detection results; determining the target user's current task execution state as a deadlock type in the intervention scenario when the target user's repetition detection results meet the repetition expression condition.
[0102] The process of determining the completion result of the expected collaborative goal in the task parsing results based on multimodal task execution information in real time can be entirely implemented by calling the multimodal large model. The silence duration is used to determine the cumulative time during which participating users, especially the target user, have no input. For example, the cumulative duration can begin at the start of the target user's turn, and end when the target user inputs any content, obtaining the silence duration and resetting the cumulative duration. The bottleneck condition is used to detect whether the silence duration exceeds a preset threshold and / or whether key information related to the task content is missing from the input semantics. The bottleneck type can be a scenario where the task is still executing but progress is paused or a scenario where content at a certain stage is missing. For example, a bottleneck condition could be a silence duration greater than 8 seconds, or the input semantics missing key semantics. Key semantics are usually semantics pre-configured for the target user's role. Guiding content: You can try starting from point xx.
[0103] Input semantics refers to the semantics of the multimodal task execution information input by the target user. Semantic difference can be the degree of difference between the input semantics and the task content. The cosine of the vector of input semantics and the semantic vector corresponding to the task topic in the task content can be calculated as the similarity between the input semantics and the task content. Off-topic conditions are used to determine whether the semantic difference is greater than a preset difference threshold. Off-topic types can be scenarios that deviate from the task topic. For example, an off-topic condition could be a semantic difference greater than 45°. Intervention content: We are discussing whether to bring mobile phones to class; please don't talk about tomorrow's PE class.
[0104] Input content can refer to the multimodal task execution information input by the target user. Emotional content can refer to the emotion-related content within the multimodal task execution information input by the target user. Negative emotion conditions are used to detect the presence of negative emotional words in the emotional content. Argument type can refer to scenarios where the target user has conflicts with other participating users. Negative emotion conditions are used to detect the presence of specific negative emotional words or phrases within the multimodal task execution information input by the target user. Intervention content: Pause the task and ask everyone to speak based on viewpoints rather than personal opinions.
[0105] Duplicate detection results can refer to duplicates in the input content and historical content of the multimodal task execution information. Similarity detection can be performed on statements from each round in the input and historical content, and the number of rounds containing similar statements can be counted as the duplicate detection result. A duplicate expression condition is used to determine if the duplicate detection result exceeds a preset threshold. Deadlock types can occur when participating users are unable to progress in the same stage. For example, a duplicate expression condition could be that the number of duplicates detected is greater than 2. Intervention content: Both parties have reached a consensus on this point; we will move on to the next topic.
[0106] Additionally, follow-up guidance can be provided for errors that are repeatedly made by the target user. For example, the error content and corresponding correct content of the target user can be added to the task summary information.
[0107] It is evident that by defining multiple types of current task execution states and judgment rules, the number of task execution scenarios can be increased, and corresponding differentiated processing methods can be provided for each scenario. This enables a scenario-to-response processing mechanism, improves task operation stability and efficiency, and enhances system fault tolerance in cases where the task process cannot proceed.
[0108] Figure 3 This is a schematic diagram of a multimodal collaborative education task processing device provided in an embodiment of the present invention. This embodiment of the present invention is applicable to situations where an intelligent agent is invoked to execute a multimodal collaborative education task based on a large multimodal model. The device can execute multimodal collaborative education task processing methods and can be implemented in hardware and / or software. The device can be configured in a server.
[0109] See Figure 3 The multimodal collaborative education task processing device shown includes: The task parsing module 301 is used to call the multimodal big model to extract task parsing results from the multimodal collaborative education task initiated by the target user; Collaborator determination module 302 is used to call the large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the memory learning information of the target user; The task distribution module 303 is used to send the task content from the task parsing result to the target user and each of the collaborators when a confirmation command to start the task is detected from each participating user; the participating users include the target user and the real users among the collaborators. Task execution module 304 is used to acquire multimodal task execution information provided by the target user and each of the collaborators in real time regarding the task content; The target determination module 305 is used to call the multimodal large model in real time to determine the completion determination result of the expected collaborative target in the task parsing result based on the multimodal task execution information; The task summary module 306 is used to call the large model to obtain the collaboration logs of each participating user and generate task summary information when the completion determination result is that the task is completed.
[0110] The technical solution of this invention determines the collaborators for a multimodal collaborative education task by obtaining task parsing results from a multimodal collaborative education task initiated by a target user and combining this with the target user's memory and learning information. When the actual user in the multimodal collaborative education task confirms the task initiation command, the task content is sent to each participant, and the multimodal task execution information of each participant is obtained. The multimodal task execution information is then compared with the expected collaborative goal to determine completion, resulting in task summary information. This achieves data perception and processing for the target user in the multimodal collaborative education task, and implements the collaborative process. It solves the technical problem in existing technologies that can only process data in single-person education scenarios, increases the application scenarios for collaboration, enriches education scenarios, increases the content of data perception and multimodal data in education scenarios, and can accurately process multimodal data, correctly execute collaborative task processes, and simultaneously interact with data from multiple participants, reducing data silos.
[0111] Optionally, the collaborator determination module 302 is specifically used for: When it is determined that the number of participants in the multimodal collaborative education task is less than the standard number of participants in the task analysis result, the memory learning information of the target user is obtained; The large model is invoked to query the first user corresponding to the multimodal collaborative education task among the current online users based on the task parsing results and the target user's memory learning information; When there is a first user among the currently online users, initiate a collaboration request to each of the first users; The first user to confirm participation will be identified as a collaborator in the multimodal collaborative education task.
[0112] Optionally, the collaborator determination module 302 is specifically used for: Based on the target user's memory learning information, determine the target user's ability disadvantage information and social conflict information, and use them as target matching information; Search among currently online users for at least one second user corresponding to the target matching information; Based on the learning path of each second user, identify the third user among the second users who corresponds to the task parsing result; The first user is determined based on each of the third users.
[0113] Optionally, the multimodal collaborative education task processing device also includes: The virtual user generation module is used for: After identifying the first user who confirmed participation as a collaborator in the multimodal collaborative education task, when it is determined that the number of participants in the multimodal collaborative education task is less than the standard number in the task analysis result and the current time meets the collaboration invitation time limit, at least one virtual personality is determined based on the task analysis result and the target user's memory learning information. Based on the virtual personalities described, virtual humans are identified and used as collaborators in the multimodal collaborative education tasks.
[0114] Optionally, the multimodal collaborative education task processing device also includes: Intervention guidance module, used for: After sending the task content to the target user and each of the collaborators, the current task execution status is determined based on the multimodal task execution information provided by the target user and the collaborators. When the current task execution state meets the intervention scenario, the intervention content is output to the target user according to the intervention type corresponding to the current task execution state; When the current task execution state meets the guidance scenario, guidance content is output to the target user according to the guidance type corresponding to the current task execution state.
[0115] Optional, intervention guidance module, specifically used for: Based on the multimodal task execution information provided by the target user and the collaborator, the input state of the target user is obtained, and the input state includes: silence duration and input semantics; When the input state of the target user meets the stuck and broken conditions, the current task execution state of the target user is determined to be a stuck and broken type in the guidance scenario; Calculate the semantic difference between the target user's input semantics and the task content; When the semantic difference of the target user meets the off-topic condition, the current task execution state of the target user is determined to be the off-topic type of the intervention scenario; Based on the multimodal task execution information provided by the target user and the collaborator, the input content of the target user is obtained, and the emotional content of the target user is identified; When the target user's emotional content meets the negative emotional condition, the target user's current task execution state is determined to be the argument type of the intervention scenario; Perform contextual duplicate detection on the input content of the target user to obtain duplicate detection results; When the duplicate detection result of the target user meets the duplicate expression condition, the current task execution state of the target user is determined to be a deadlock type in the intervention scenario.
[0116] Optionally, the target determination module 305 is specifically used for: Obtain the interaction content of the current stage in the multimodal task execution information, at least one target content of the current stage in the collaborative expected goal, and the end judgment rule for each target content; For each target content, the interaction content of the current stage and the target content are matched according to the end judgment rule of the target content to obtain the target completion detection result of the target content; When the target completion detection results of each target content in the current stage are all completed and the current stage is the final stage of the multimodal collaborative education task, the completion judgment result of the expected collaborative goal is determined to be completed. When the target completion detection results of each target content in the current stage are all completed and the current stage is different from the final stage of the multimodal collaborative education task, proceed to the next stage of the current stage, send the sub-task content of the next stage to the target user and each of the collaborators, and determine that the completion judgment result of the expected collaborative goal is not completed. When the target completion detection result of the target content is not completed, target guidance content is generated based on the difference between the current interactive content and the uncompleted target content, and sent to the target user so that the target user can provide new interactive content, and the completion judgment result of the expected collaborative goal is determined to be uncompleted; The incomplete target content is matched with the new interactive content, and the target completion detection result of the incomplete target content is updated.
[0117] The acquisition, storage, and application of data involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0118] The multimodal collaborative education task processing device provided in this embodiment of the invention can execute the multimodal collaborative education task processing method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0119] Figure 4 A schematic diagram of the structure of an electronic device 400 that can be used to implement an embodiment of the present invention is shown.
[0120] like Figure 4As shown, the electronic device 400 includes at least one processor 401 and a memory, such as a read-only memory 402 or a random access memory 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 402 or loaded from storage unit 408 into the random access memory 403. The random access memory 403 may also store various programs and data required for the operation of the electronic device 400. The processor 401, read-only memory 402, and random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.
[0121] Multiple components in electronic device 400 are connected to input / output interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0122] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as multimodal collaborative education task processing methods.
[0123] In some embodiments, the multimodal collaborative education task processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 400 via read-only memory 402 and / or communication unit 409. When the computer program is loaded into random access memory 403 and executed by processor 401, one or more steps of the multimodal collaborative education task processing method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the multimodal collaborative education task processing method by any other suitable means (e.g., by means of firmware).
[0124] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0127] To provide interaction with the user, the systems and techniques described herein can be implemented on an operational detection device. This call response processing device includes: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the call response processing device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0128] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0129] A computing system can include target user terminals and servers. Target user terminals and servers are generally geographically separated and typically interact via communication networks. The relationship between target user terminals and servers is created by computer programs running on the respective computers and establishing a target user-server relationship between them. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.
[0130] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0131] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for processing educational tasks based on multimodal large-scale model-driven agent collaboration, characterized in that, include: The project's intelligent agent invokes a multimodal large model to extract task parsing results from multimodal collaborative education tasks initiated by the target user; The project role intelligent agent invokes the large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information; When the confirmation command to start the task is detected from each participating user, the project role agent calls the large model to send the task content from the task parsing result to the target user and each of the collaborators; the participating users include the target user and the real users among the collaborators. The project role intelligent agent obtains multimodal task execution information provided by the target user and each of the collaborators in real time regarding the task content; The project role intelligent agent calls the multimodal big model in real time to determine the completion judgment result of the expected collaborative goal in the task parsing result based on the multimodal task execution information; When the completion determination result is that the task is completed, the project role agent calls the large model to obtain the collaboration logs of each participating user and generates task summary information.
2. The method according to claim 1, characterized in that, The invocation model, based on the task parsing results and the target user's memory learning information, determines at least one collaborator matching the multimodal collaborative education task, including: When it is determined that the number of participants in the multimodal collaborative education task is less than the standard number of participants in the task analysis result, the memory learning information of the target user is obtained; The large model is invoked to query the first user corresponding to the multimodal collaborative education task among the current online users based on the task parsing results and the target user's memory learning information; When there is a first user among the currently online users, initiate a collaboration request to each of the first users; The first user to confirm participation will be identified as a collaborator in the multimodal collaborative education task.
3. The method according to claim 2, characterized in that, The step of querying the first user corresponding to the multimodal collaborative education task among the currently online users based on the task parsing results and the target user's memory learning information includes: Based on the target user's memory learning information, determine the target user's ability disadvantage information and social conflict information, and use them as target matching information; Search among currently online users for at least one second user corresponding to the target matching information; Based on the learning path of each second user, identify the third user among the second users who corresponds to the task parsing result; The first user is determined based on each of the third users.
4. The method according to claim 2, characterized in that, After identifying the first user who confirmed their participation as a collaborator in the multimodal collaborative education task, the process also includes: When it is determined that the number of participants in the multimodal collaborative education task is less than the standard number in the task analysis result and the current time meets the collaboration invitation time limit, at least one virtual personality is determined based on the task analysis result and the target user's memory learning information. Based on the virtual personalities described, virtual humans are identified and used as collaborators in the multimodal collaborative education tasks.
5. The method according to claim 1, characterized in that, After sending the task content from the task parsing result to the target user and each of the collaborators, the process also includes: The current task execution status is determined based on the multimodal task execution information provided by the target user and the collaborator; When the current task execution state meets the intervention scenario, the intervention content is output to the target user according to the intervention type corresponding to the current task execution state; When the current task execution state meets the guidance scenario, guidance content is output to the target user according to the guidance type corresponding to the current task execution state.
6. The method according to claim 5, characterized in that, The step of determining the current task execution status based on the multimodal task execution information provided by the target user and the collaborator includes: Based on the multimodal task execution information provided by the target user and the collaborator, the input state of the target user is obtained, and the input state includes: silence duration and input semantics; When the input state of the target user meets the stuck and broken conditions, the current task execution state of the target user is determined to be a stuck and broken type in the guidance scenario; Calculate the semantic difference between the target user's input semantics and the task content; When the semantic difference of the target user meets the off-topic condition, the current task execution state of the target user is determined to be the off-topic type of the intervention scenario; Based on the multimodal task execution information provided by the target user and the collaborator, the input content of the target user is obtained, and the emotional content of the target user is identified; When the target user's emotional content meets the negative emotional condition, the target user's current task execution state is determined to be the argument type of the intervention scenario; Perform contextual duplicate detection on the input content of the target user to obtain duplicate detection results; When the duplicate detection result of the target user meets the duplicate expression condition, the current task execution state of the target user is determined to be a deadlock type in the intervention scenario.
7. The method according to claim 1, characterized in that, The real-time determination of the completion judgment result of the expected collaborative goal in the task parsing result based on the multimodal task execution information includes: Obtain the interaction content of the current stage in the multimodal task execution information, at least one target content of the current stage in the collaborative expected goal, and the end judgment rule for each target content; For each target content, the interaction content of the current stage and the target content are matched according to the end judgment rule of the target content to obtain the target completion detection result of the target content; When the target completion detection results of each target content in the current stage are all completed and the current stage is the final stage of the multimodal collaborative education task, the completion judgment result of the expected collaborative goal is determined to be completed. When the target completion detection results of each target content in the current stage are all completed and the current stage is different from the final stage of the multimodal collaborative education task, proceed to the next stage of the current stage, send the sub-task content of the next stage to the target user and each of the collaborators, and determine that the completion judgment result of the expected collaborative goal is not completed. When the target completion detection result of the target content is not completed, target guidance content is generated based on the difference between the current interactive content and the uncompleted target content, and sent to the target user so that the target user can provide new interactive content, and the completion judgment result of the expected collaborative goal is determined to be uncompleted; The incomplete target content is matched with the new interactive content, and the target completion detection result of the incomplete target content is updated.
8. A multimodal collaborative educational task processing device, characterized in that, The device includes: The task parsing module is used to call the multimodal big model to extract task parsing results from multimodal collaborative education tasks initiated by the target user; The collaborator determination module is used to call the large model to determine at least one collaborator matching the multimodal collaborative education task based on the task parsing results and the target user's memory learning information; The task distribution module is used to send the task content from the task parsing result to the target user and each of the collaborators when a confirmation command to start the task is detected from each participating user; the participating users include the target user and the real users among the collaborators; The task execution module is used to acquire multimodal task execution information provided in real time by the target user and each of the collaborators regarding the task content; The target determination module is used to call the multimodal large model in real time to determine the completion determination result of the expected collaborative target in the task parsing result based on the multimodal task execution information; The task summary module is used to retrieve the collaboration logs of each participating user from the large model and generate task summary information when the completion determination result is that the task is completed.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal collaborative education task processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multimodal collaborative education task processing method according to any one of claims 1-7.