Education scene-oriented AI agent process automation method and system

By constructing hierarchical knowledge graphs and using meta-learners, the problem of AI agents adapting to complex teaching content and individual differences in students in educational scenarios is solved, and the automated optimization of the agent process and personalized learning experience are realized.

CN120163422AInactive Publication Date: 2025-06-17SUZHOU INST OF TRADE & COMMERCE
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510363345.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing AI agents are difficult to adapt to complex teaching content and individual differences between students in educational scenarios, lack personalized learning experience, and have a single interaction method, making it difficult to capture the fine-grained information in students' learning process.

Method used

The question-and-answer vector matrix is ​​obtained through semantic segmentation, a hierarchical knowledge graph is constructed, and an agent encoder model is generated using graph attention network and contrast learning method, combining meta-learners and knowledge cache modules to realize task decomposition and decision optimization.

Benefits of technology

It improves the learning efficiency and performance of the agent, realizes the automated optimization of the agent process, enhances adaptability and robustness, and provides a more personalized and efficient educational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163422A_ABST
    Figure CN120163422A_ABST
Patent Text Reader

Abstract

The invention provides an AI agent process automation method and system for an education scene, and relates to the technical field of AI education, and the method comprises the steps: constructing a hierarchical knowledge graph through semantic segmentation, and constructing an agent encoder model through a graph attention network and comparative learning. Based on agent operation data and a feedback mechanism, operation parameters are dynamically adjusted, task decomposition is performed, sub-tasks are processed by using a meta-learner, and a knowledge cache module is established. And semantic analysis and association network construction are performed by using the knowledge cache module, a decision model is trained, and an execution process is optimized. The workflow evolution trend is predicted according to the agent state data, sub-task configuration is optimized, closed-loop optimization is formed, and therefore the automatic processing efficiency and adaptability of the AI agent in an education scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to AI education technology, and in particular to an AI agent process automation method and system for educational scenarios. Background Art

[0002] In the field of education, the demand for personalized and adaptive learning is increasing day by day, and the traditional teaching mode is difficult to meet the diverse learning needs of students. The rapid development of artificial intelligence technology provides new solutions for educational scenarios. In recent years, AI agents have been widely used in the field of education. For example, intelligent teaching assistants can assist teachers in correcting homework and answering questions; intelligent learning systems can provide personalized learning resources and learning paths according to students' learning situations; intelligent evaluation systems can evaluate students' learning achievements more comprehensively and objectively. These AI agents have improved teaching efficiency and learning effects to a certain extent.

[0003] Existing AI agents usually operate based on pre-set rules and templates, lacking the ability to adapt to complex educational scenarios. When faced with new teaching content or individual differences among students, it is difficult to flexibly adjust teaching strategies and unable to provide a truly personalized learning experience.

[0004] The interaction mode between existing AI agents and the learning environment is relatively single, making it difficult to capture fine-grained information in the learning process of students. For example, factors such as students' emotional states and learning motivations have important impacts on learning effects, but existing AI agents often have difficulty effectively identifying and utilizing this information.

[0005] Existing AI agents have limitations in task decomposition and collaboration. Complex teaching tasks usually need to be decomposed into multiple subtasks and processed collaboratively, while existing AI agents often lack an efficient task decomposition and collaboration mechanism and are difficult to handle complex teaching scenarios. Summary of the Invention

[0006] Embodiments of the present invention provide an AI agent process automation method and system for educational scenarios, which can solve the problems in the prior art.

[0007] In the first aspect of the embodiments of the present invention, An AI agent process automation method for educational scenarios is provided, including: Obtaining a question-answer pair vector matrix through semantic segmentation and constructing a hierarchical knowledge graph, reconstructing the vector matrix based on the graph attention network to calculate the association strength between nodes, constructing a two-tower encoding network using the contrastive learning method, fusing multi-layer features using a cross-layer attention module and updating network parameters to generate an agent encoder model; Collect the agent operation data based on the agent encoder model and generate a performance evaluation vector. Input the performance evaluation vector into the feedback collection mechanism to update the operation parameters. Calculate the task complexity according to the operation parameters and implement task decomposition. Process the subtasks through the meta-learner and establish a knowledge cache module; Use the knowledge cache module to perform semantic parsing on the task entity and construct an association network. Train a decision-making model based on the association network. Input the output result of the decision-making model into the workflow adjustment module to optimize the execution process, and use the agent collaborative control mechanism to generate a decision result; Collect the agent state data according to the decision result to construct a state vector. Calculate the state transition sequence based on the state vector to predict the workflow evolution trend. Combine the workflow evolution trend to extract scene features and optimize the subtask configuration scheme. Feed back the configuration scheme to the performance evaluation link to form an optimization closed-loop.

[0008] Obtain the question-answer pair vector matrix through semantic segmentation and construct a hierarchical knowledge graph. Calculate the association strength between nodes based on the graph attention network to reconstruct the vector matrix. Use the contrastive learning method to construct a two-tower encoding network. Use the cross-layer attention module to fuse multi-layer features and update the network parameters. The generated agent encoder model includes: Receive the question-answer pair training corpus, perform semantic segmentation on the question-answer pair training corpus to obtain a question segment set and an answer segment set, use the pre-trained language model to extract the representation vectors of the question segment set and the answer segment set respectively, and construct an initial question-answer pair vector matrix; Construct a hierarchical knowledge graph based on the initial question-answer pair vector matrix, map the question segment set to node vectors, map the answer segment set to edge vectors, and calculate the association strength matrix between the node vectors and the edge vectors through the graph attention network; Reconstruct the initial question-answer pair vector matrix according to the association strength matrix, use the contrastive learning method to calculate the mutual information loss between positive sample question-answer pairs and the contrast loss between negative sample question-answer pairs, and use the weighted sum of the mutual information loss and the contrast loss as the training objective function; Construct a two-tower encoding network based on the training objective function. The two-tower encoding network includes a question encoding tower and an answer encoding tower. The question encoding tower and the answer encoding tower adopt a transformer structure with shared parameters, and optimize the parameters of the two-tower encoding network through the gradient descent method; Perform adaptive fusion on the intermediate layer feature maps of the two-tower encoding network to construct a cross-layer attention module. The cross-layer attention module calculates the interaction weights between different layer features based on the association strength matrix to generate an enhanced feature representation; Input the enhanced feature representation into the decoder for reconstruction, calculate the reconstruction error, and update the parameters of the two-tower encoding network. Use the trained two-tower encoding network as the encoder model of the agent to achieve the vectorized representation of the question-answer pair.

[0009] Collect the agent operation data based on the agent encoder model and generate a performance evaluation vector. Input the performance evaluation vector into the feedback collection mechanism to update the operation parameters. Calculate the task complexity based on the operation parameters and perform task decomposition. Process the subtasks through a meta-learner and establish a knowledge cache module, including: Construct a performance evaluation module. The performance evaluation module collects the operation data of the agent, calculates performance evaluation metrics including response latency, decision accuracy, and resource utilization rate based on preset weight coefficients, and generates a performance evaluation vector. Use the performance evaluation vector to construct a feedback collection mechanism. The feedback collection mechanism records real-time feedback information, calculates the parameter adjustment gradient according to the real-time feedback information, and dynamically updates the operation parameters of the agent. Input the operation parameters into a task complexity evaluation function. The task complexity evaluation function analyzes the input task based on the task hierarchy depth, data dependency degree, and resource demand, decomposes it into multiple subtasks, and calculates the dependency relationship between multiple subtasks at the same time. Construct a meta-learner to receive the multiple subtasks. The meta-learner generates scene features through a scene feature extraction function, fuses the scene features with the basic model parameters to obtain adaptive parameters, and uses the adaptive parameters to optimize the subtasks. Construct a time-series feature vector based on the processing results of the subtasks and calculate the dynamic characterization value. Combine the characterization value with the co-occurrence frequency to generate the time-series correlation strength. Evaluate the importance of knowledge items according to the correlation strength and access frequency, and update the cache content according to the importance. Input the performance evaluation metrics, task execution efficiency, and adaptability metrics into a global optimization function. The output of the global optimization function is used to adjust the task decomposition strategy, feedback the task execution result to the meta-learner for optimizing the knowledge cache, and feedback the cache status to the performance evaluation module to form an optimization closed-loop.

[0010] Construct a time-series feature vector based on the processing results of the subtasks and calculate the dynamic characterization value. Combine the characterization value with the co-occurrence frequency to generate the time-series correlation strength. Evaluate the importance of knowledge items according to the correlation strength and access frequency, and update the cache content according to the importance, including: Construct a time-series feature vector. The time-series feature vector contains the feature dimension information of knowledge items at multiple time points, and calculate the dynamic characterization value of knowledge items based on the time-series feature vector. Input the dynamic characteristic value into the association strength calculation module. The association strength calculation module calculates the initial association strength based on the co-occurrence frequency and the time decay function, and obtains the time-series association strength by enhancing the similarity of the dynamic characteristic value; Construct a knowledge item importance evaluation module based on the time-series association strength. The knowledge item importance evaluation module fuses the time-weighted access frequency and the importance of associated neighbors to obtain the comprehensive importance of the knowledge item; Calculate the replacement probability between knowledge items according to the comprehensive importance, dynamically adjust the update threshold based on the change of the cache hit rate, and execute cache update according to the replacement probability and the update threshold; Input the dynamic characteristic value, the time-series association strength, and the comprehensive importance into the performance evaluation module. The performance evaluation module optimizes the system parameters based on the gradient of the performance index, and feeds back the optimized system parameters to the feature extraction and association calculation link.

[0011] Use the knowledge cache module to perform semantic parsing on the task entity and construct an association network. Train a decision model based on the association network, input the output result of the decision model into the workflow adjustment module to optimize the execution process, and adopt an agent collaboration control mechanism to generate decision results, including: Extract semantic entities in the education task, calculate the vector cosine similarity of the semantic entities through vector space mapping, calculate the point mutual information value by combining the co-occurrence frequency and the marginal probability of the semantic entities in the corpus, and generate a task semantic association network by weighted fusion of the vector cosine similarity and the point mutual information value; Construct a state transition scoring function based on the task semantic association network, calculate the state transition probability by exponential normalization of the state transition scoring function, calculate the decision value based on the immediate reward and the discount factor using the temporal difference method, and evaluate and optimize the decision sequence according to the state transition probability and the decision value; Obtain the decision point density by counting the number of decision nodes in the unit process according to the decision sequence, calculate the nested depth and the number of branches of the process branch to obtain the branch complexity, count the number and levels of the feedback loops to obtain the feedback link number, calculate the process complexity by weighted summation of the decision point density, the branch complexity, and the feedback link number, and obtain the workflow adjustment factor through gradient calculation; Receive the workflow adjustment factor, score multiple intelligent agents in three dimensions of knowledge reserve, decision-making ability, and execution efficiency to obtain an evaluation score; construct a weighted average function based on the evaluation score, and fuse the decision outputs of multiple intelligent agents according to the weighted average function to obtain a collaborative decision result.

[0012] Collect the intelligent agent state data according to the decision result to construct a state vector, calculate the state transition sequence based on the state vector to predict the workflow evolution trend, extract the scenario features in combination with the workflow evolution trend and optimize the subtask configuration scheme, and feedback the configuration scheme to the performance evaluation link to form an optimization closed-loop, including: Calculate the performance evaluation index based on the preset weight coefficient, and construct a state vector containing the performance evaluation index; perform time series sampling on the state vector to obtain the state transition sequence, calculate the mutual information and conditional entropy between adjacent states, and construct a Markov state transition matrix in combination with the state duration and transition frequency to predict the workflow evolution trend; Calculate the weighted combination of the task hierarchy depth, data dependence intensity and computing resource requirements according to the workflow evolution trend to obtain the task complexity, decompose the original task into multiple subtasks based on the task complexity, and calculate the dependence intensity matrix between subtasks; Adopt the transfer learning method to extract knowledge features from the historical execution data and scenario features, fuse the knowledge features with the current scenario parameters to obtain the scenario adaptability parameters, and optimize the configuration of subtasks based on the scenario adaptability parameters; Calculate the trend features of the real-time feedback data during the execution process through the exponential moving average method, calculate the signal reliability in combination with the fluctuation amplitude and variance of the feedback signal, and perform weighted fusion on the feedback data based on the signal reliability; Input the weighted fusion feedback data into the global optimization function, update the system parameters using the gradient descent method, and apply the system parameters to the performance evaluation, state prediction and task decomposition links respectively to achieve the closed-loop optimization of the workflow.

[0013] In the second aspect of the embodiments of the present invention, Provide an AI intelligent agent process automation system for the education scenario, including: The first unit is used to obtain the question-and-answer pair vector matrix through semantic segmentation and construct a hierarchical knowledge graph, calculate the association strength between nodes based on the graph attention network to reconstruct the vector matrix, construct a two-tower encoding network using the contrastive learning method, fuse multi-layer features using the cross-layer attention module and update the network parameters to generate an intelligent agent encoder model; The second unit is used to collect the intelligent agent operation data based on the intelligent agent encoder model and generate a performance evaluation vector, input the performance evaluation vector into the feedback collection mechanism to update the operation parameters, calculate the task complexity according to the operation parameters and achieve task decomposition, and process the subtasks through the meta-learner and establish a knowledge cache module; A third unit, configured to perform semantic parsing on task entities by using the knowledge cache module and construct an association network, train a decision-making model based on the association network, input the output result of the decision-making model into a workflow adjustment module to optimize the execution process, and generate a decision result by adopting an agent collaborative control mechanism; A fourth unit, configured to collect agent state data according to the decision result to construct a state vector, calculate a state transition sequence based on the state vector to predict the workflow evolution trend, extract scenario features in combination with the workflow evolution trend and optimize the subtask configuration scheme, and feedback the configuration scheme to the performance evaluation link to form an optimization closed loop.

[0014] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] The beneficial effects of the present application are as follows: 1. Improve the learning efficiency and performance of the agent: By constructing a hierarchical knowledge graph and a contrast learning method, the accuracy and generalization ability of the agent encoder model can be effectively improved. Combining the meta-learner and the knowledge cache module can quickly process subtasks and accumulate experience, thereby improving the learning efficiency and performance of the agent.

[0017] 2. Realize the automatic process optimization of the agent: Based on the performance evaluation vector and the feedback acquisition mechanism, the agent can automatically adjust the operation parameters and the task decomposition strategy. By using the decision-making model and the workflow adjustment module, the execution process can be optimized and the task completion efficiency can be improved. The prediction of the state transition sequence and the extraction of scenario features can dynamically adjust the subtask configuration scheme to realize the automatic optimization of the agent process.

[0018] 3. Enhance the adaptability and robustness of the agent: The agent collaborative control mechanism can effectively handle complex tasks and improve the adaptability and robustness of the agent. Through continuous performance evaluation and optimization closed loop, the agent can continuously learn and improve, enhancing its adaptability in different educational scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic flowchart of the AI agent process automation method for educational scenarios according to the embodiments of the present invention; Figure 2 This is the architecture diagram of the AI agent process automation system for the education scenario in the embodiments of the present invention; Figure 3 This is the flowchart of RAG retrieval enhancement in the embodiments of the present invention; Figure 4 This is the example table of sample query and large language model generation in the embodiments of the present invention; Figure 5 This is the comparison table graph of the performance indicators of the AI agent process automation method in the embodiments of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 This is the flowchart of the AI agent process automation method for the education scenario in the embodiments of the present invention. As Figure 1 shown, the method includes: Obtaining a question-answer pair vector matrix through semantic segmentation and constructing a hierarchical knowledge graph, reconstructing the vector matrix based on the graph attention network to calculate the association strength between nodes, constructing a two-tower encoding network using the contrastive learning method, fusing multi-layer features using the cross-layer attention module and updating the network parameters to generate an agent encoder model; Collecting agent operation data based on the agent encoder model and generating a performance evaluation vector, inputting the performance evaluation vector into a feedback collection mechanism to update the operation parameters, calculating the task complexity according to the operation parameters and implementing task decomposition, and processing subtasks through a meta-learner and establishing a knowledge cache module; Semantically parsing task entities using the knowledge cache module and constructing an association network, training a decision model based on the association network, inputting the output result of the decision model into a workflow adjustment module to optimize the execution process, and generating a decision result using an agent collaborative control mechanism; Collect the state data of the agent according to the decision result to construct a state vector, calculate the state transition sequence based on the state vector to predict the evolution trend of the workflow, extract scenario features in combination with the workflow evolution trend and optimize the subtask configuration scheme, and feedback the configuration scheme to the performance evaluation link to form an optimization closed loop.

[0023] In an alternative embodiment, a question-answer pair vector matrix is obtained through semantic segmentation and a hierarchical knowledge graph is constructed. The association strength between nodes is calculated based on the graph attention network to reconstruct the vector matrix. A contrastive learning method is used to construct a two-tower encoding network. A cross-layer attention module is used to fuse multi-layer features and update the network parameters. The generated agent encoder model includes: Receive the question-answer pair training corpus, perform semantic segmentation on the question-answer pair training corpus to obtain a question segment set and an answer segment set, use a pre-trained language model to extract the representation vectors of the question segment set and the answer segment set respectively, and construct an initial question-answer pair vector matrix; Construct a hierarchical knowledge graph based on the initial question-answer pair vector matrix, map the question segment set to node vectors, map the answer segment set to edge vectors, and calculate the association strength matrix between the node vectors and the edge vectors through a graph attention network; Reconstruct the initial question-answer pair vector matrix according to the association strength matrix, calculate the mutual information loss between positive sample question-answer pairs and the contrast loss between negative sample question-answer pairs using the contrastive learning method, and use the weighted sum of the mutual information loss and the contrast loss as the training objective function; Construct a two-tower encoding network based on the training objective function. The two-tower encoding network includes a question encoding tower and an answer encoding tower. The question encoding tower and the answer encoding tower adopt a transformer structure with shared parameters, and optimize the parameters of the two-tower encoding network through the gradient descent method; Perform adaptive fusion on the intermediate layer feature maps of the two-tower encoding network to construct a cross-layer attention module. The cross-layer attention module calculates the interaction weights between features of different layers based on the association strength matrix to generate an enhanced feature representation; Input the enhanced feature representation into a decoder for reconstruction, calculate the reconstruction error and update the parameters of the two-tower encoding network, and use the trained two-tower encoding network as the encoder model of the agent to realize the vectorized representation of the question-answer pair.

[0024] The purpose of this method is to construct an agent encoder model that can efficiently vectorize the question-answer pair, thereby improving the performance of the agent in applications such as question-answer systems and knowledge graphs. The core idea is to deeply fuse the semantic information and structured knowledge of the question-answer pair through semantic segmentation, graph attention network, contrastive learning, and cross-layer attention mechanism, and finally obtain a robust and discriminative vector representation.

[0025] Receiving the question-answer pair training corpus is the first step in constructing an agent encoder model. For example, a typical question-answer pair training corpus may contain the following content: Question: "What is deep learning?" Answer: "Deep learning is a machine learning method that uses multiple-layer neural networks to extract complex features from data." Perform semantic segmentation on these question-answer pairs, breaking down the questions and answers into smaller semantic units, namely the set of question fragments and the set of answer fragments. The segmentation process can employ various natural language processing techniques, such as rule-based segmentation, statistical model-based segmentation, or deep learning model-based segmentation. For example, for the above question-answer pair, the segmented results may be as follows: Set of question fragments: {"What is", "deep learning", "?"} Set of answer fragments: {"Deep learning", "is", "a", "machine learning", "method", "it", "uses", "multiple-layer", "neural networks", "to", "extract", "data", "in", "the", "complex", "features", "."} After obtaining the set of question fragments and the set of answer fragments, use a pre-trained language model (such as BERT, RoBERTa, etc.) to extract their representation vectors respectively. The pre-trained language model can map each fragment to a high-dimensional vector space, and this vector can capture the semantic information of the fragment. For example, using the BERT model, the "deep learning" fragment can be mapped to a 768-dimensional vector. Concatenate the vectors of all question fragments to obtain the representation vector of the set of question fragments. Using the same method, obtain the representation vector of the set of answer fragments, and construct the initial vector matrix of the question-answer pair. The initial vector matrix contains the basic semantic information of the question-answer pair, laying the foundation for subsequent knowledge fusion and relationship modeling.

[0026] Based on the obtained initial vector matrix of the question-answer pair, construct a hierarchical knowledge graph to explicitly represent the relationship between the question fragments and the answer fragments. Specifically, map the set of question fragments to node vectors in the knowledge graph, and map the set of answer fragments to edge vectors. The node vectors represent the semantic information of the question fragments, and the edge vectors represent the explanation or supplementary description of the answer fragments to the question fragments. For example, in the knowledge graph, the "What is" node may be connected to the "deep learning" node and is connected by an edge representing the "definition" relationship.

[0027] With the knowledge graph, the association strength matrix between node vectors and edge vectors can be calculated using the Graph Attention Network (GAT). The Graph Attention Network is a neural network model that can learn the importance between nodes. Through the Graph Attention Network, the attention weight of each node (question segment) to each edge (answer segment) can be calculated, and this weight reflects the association strength between the two. For example, the "Deep Learning" node may have a relatively high attention weight for the edge "Deep learning is a machine learning method" and a relatively low attention weight for the edge "It uses multi-layer neural networks". Through multi-layer iterative calculations, the Graph Attention Network can capture the complex dependencies between nodes and edges, thereby obtaining a more accurate association strength matrix.

[0028] Reconstruct the initial vector matrix of the question-answer pairs according to the association strength matrix. The purpose of reconstruction is to incorporate the relationship information learned from the knowledge graph into the vector representation of the question-answer pairs. Specifically, each element in the initial vector matrix can be weighted according to the association strength matrix, so as to highlight important semantic units and suppress irrelevant semantic units. The reconstructed vector matrix contains richer semantic information and structured knowledge, and can better represent the meaning of the question-answer pairs.

[0029] To train the model, a suitable training objective function needs to be defined. Here, the contrastive learning method is adopted, and the model parameters are optimized by calculating the mutual information loss between positive sample question-answer pairs and the contrastive loss between negative sample question-answer pairs. Positive sample question-answer pairs refer to question-answer pairs with semantic relevance, such as "What is deep learning?" and "Deep learning is a machine learning method". Negative sample question-answer pairs refer to question-answer pairs with semantic irrelevance, such as "What is deep learning?" and "What's the weather like in Beijing today?". The mutual information loss encourages the model to map positive sample question-answer pairs to nearby positions in the vector space, and the contrastive loss encourages the model to map negative sample question-answer pairs to distant positions in the vector space. The weighted sum of the mutual information loss and the contrastive loss is used as the final training objective function.

[0030] Construct a two-tower encoding network based on the training objective function. The two-tower encoding network includes a question encoding tower and an answer encoding tower, and both encoding towers adopt a transformer structure with shared parameters. The design of shared parameters can reduce the number of model parameters and improve the generalization ability of the model. The question encoding tower is responsible for encoding the representation vector of the question segment set into a fixed-length question vector, and the answer encoding tower is responsible for encoding the representation vector of the answer segment set into a fixed-length answer vector. The parameters of the two-tower encoding network are optimized by the gradient descent method, so that the distance between the question vector and the answer vector in the vector space can reflect their semantic similarity.

[0031] To further improve the performance of the model, the intermediate layer feature maps of the two - tower encoding network are adaptively fused to construct a cross - layer attention module. The cross - layer attention module calculates the interaction weights between features of different layers based on the association strength matrix and generates enhanced feature representations. Feature maps of different layers capture semantic information at different levels. For example, low - level features may capture lexical information, and high - level features may capture semantic information. Through the cross - layer attention module, features at different levels can be fused to obtain a more comprehensive semantic representation. Specifically, the contribution degree of each feature map to the final representation can be calculated according to the association strength matrix, and the feature maps can be weighted and fused according to the contribution degree.

[0032] The enhanced feature representation is input into the decoder for reconstruction. The reconstruction error is calculated and the parameters of the two - tower encoding network are updated. The reconstruction error reflects the model's ability to reconstruct the input data. By minimizing the reconstruction error, the model's ability to understand the input data can be improved. The decoder can restore the enhanced feature representation to the original question - answer pair representation and calculate the difference between the restored representation and the original representation. Adding the reconstruction error as a regularization term to the training objective function can prevent the model from overfitting and improve the model's generalization ability.

[0033] After training through the above steps, the encoder model of the agent is finally obtained. This model can map question - answer pairs into a high - dimensional vector space, and keep the distance between semantically similar question - answer pairs in the vector space relatively close, while the distance between semantically dissimilar question - answer pairs is relatively far. This model can be applied to various question - answering systems, knowledge graphs and other applications to improve the performance of the agent.

[0034] Through semantic segmentation and pre - trained language models, semantic information in question - answer pairs can be effectively extracted, laying a foundation for subsequent knowledge fusion and relationship modeling. Through graph attention networks and contrastive learning, the association relationships between question - answer pair segments can be learned, and structured knowledge can be incorporated into the vector representation, thereby improving the accuracy and robustness of the vector representation. Through the cross - layer attention module and reconstruction error optimization, semantic information at different levels can be fused, and the generalization ability of the model can be improved, thus obtaining a better encoder model.

[0035] In an alternative implementation, based on the agent encoder model, agent operation data is collected and a performance evaluation vector is generated. The performance evaluation vector is input into a feedback collection mechanism to update the operation parameters. The task complexity is calculated according to the operation parameters and task decomposition is achieved. Processing sub - tasks through a meta - learner and establishing a knowledge cache module includes: Construct a performance evaluation module. The performance evaluation module collects the operation data of the agent, calculates performance evaluation metrics including response latency, decision accuracy, and resource utilization rate based on preset weight coefficients, and generates a performance evaluation vector; Construct a feedback collection mechanism using the performance evaluation vector. The feedback collection mechanism records real-time feedback information, calculates parameter adjustment gradients based on the real-time feedback information, and dynamically updates the operating parameters of the agent. Input the operating parameters into a task complexity evaluation function. The task complexity evaluation function analyzes and decomposes the input task into multiple subtasks based on the task hierarchy depth, data dependency degree, and resource demand, and calculates the dependency relationships between the multiple subtasks at the same time. Construct a meta-learner to receive the multiple subtasks. The meta-learner generates scene features through a scene feature extraction function, fuses the scene features with the basic model parameters to obtain adaptive parameters, and uses the adaptive parameters to optimize the subtasks. Construct a time series feature vector based on the processing results of the subtasks and calculate the dynamic representation value. Combine the representation value with the co-occurrence frequency to generate the time series correlation strength. Evaluate the importance of knowledge items according to the correlation strength and access frequency, and update the cache content according to the importance. Input the performance evaluation metrics, task execution efficiency, and adaptability metrics into a global optimization function. The output of the global optimization function is used to adjust the task decomposition strategy, feedback the task execution results to the meta-learner for optimizing the knowledge cache, and feedback the cache status to the performance evaluation module to form an optimization closed-loop.

[0036] This module is responsible for collecting various data during the operation of the agent and converting them into quantifiable performance evaluation metrics. The collection of agent operation data can be achieved through various methods. For example, monitoring points can be embedded in the agent's code to record information such as its response time, decision-making results, and resource consumption; or through an external monitoring system to monitor the operation status of the agent in real time. The collected data will be preprocessed, such as removing outliers and smoothing noise, to ensure the accuracy of the evaluation results.

[0037] In terms of the calculation of performance evaluation metrics, we selected response latency, decision accuracy, and resource utilization rate as the core metrics. Response latency refers to the time required for the agent to receive a task request and give a response, reflecting the processing speed of the agent; decision accuracy refers to the proportion of correct decisions made by the agent when completing a task, reflecting the intelligence level of the agent; resource utilization rate refers to the degree of utilization of computing resources, storage resources, etc. by the agent when completing a task, reflecting the efficiency of the agent.

[0038] To comprehensively evaluate the performance of the agent, we set preset weight coefficients for each metric. These weight coefficients can be adjusted according to the actual application scenario. For example, in scenarios with high requirements for real-time performance, the weight of response latency can be increased; in scenarios with high requirements for accuracy, the weight of decision-making accuracy can be increased. By means of weighted summation, multiple performance evaluation metrics are converted into a performance evaluation vector.

[0039] Suppose during a task execution of the agent, the response latency is 0.5 seconds, the decision-making accuracy is 90%, and the resource utilization rate is 70%. If the preset weight coefficients are 0.4, 0.3, and 0.3 respectively, the calculation process of the performance evaluation vector is as follows: Response latency score: 0.5 seconds * 0.4 = 0.2 Decision-making accuracy score: 90% * 0.3 = 0.27 Resource utilization rate score: 70% * 0.3 = 0.21 Performance evaluation vector: [0.2, 0.27, 0.21] This vector contains the performance of the agent in this task execution, providing a data basis for subsequent parameter adjustment and optimization.

[0040] This mechanism is used to record the real-time feedback information of the agent and calculate the parameter adjustment gradient based on this information, so as to dynamically update the operating parameters of the agent. The core of the feedback acquisition mechanism lies in the recording of real-time feedback information and the calculation of the parameter adjustment gradient.

[0041] The recording of real-time feedback information can be achieved through methods such as log recording and database storage. The recorded information includes task ID, input data, output results, performance evaluation vector, operating parameters, etc. This information provides comprehensive data support for subsequent parameter adjustment.

[0042] The calculation of the parameter adjustment gradient is the key of the feedback acquisition mechanism. The gradient refers to the rate of change of a function at a certain point, reflecting the impact degree of parameter adjustment on performance improvement. We can use various optimization algorithms, such as gradient descent method, Adam algorithm, etc., to calculate the parameter adjustment gradient according to the real-time feedback information. These algorithms analyze historical data to find the direction and amplitude of parameter adjustment, so as to continuously improve the performance of the agent.

[0043] Suppose that during multiple task executions by an agent, the performance evaluation vector shows a downward trend, which means that the current operating parameters may not be ideal. By analyzing historical data, we find that there is a negative correlation between the response latency and a certain parameter, that is, the larger the value of this parameter, the lower the response latency. At this time, we can calculate the adjustment gradient of this parameter and adjust the parameter value in the direction of the gradient, so as to reduce the response latency and improve the overall performance of the agent.

[0044] The task complexity evaluation function aims to analyze the input task and decompose it into multiple subtasks. The basis for task decomposition is the hierarchical depth of the task, the degree of data dependence, and the resource requirements.

[0045] The hierarchical depth of a task refers to the level of how many subtasks a task can be decomposed into. The deeper the level, the more complex the task. The degree of data dependence refers to the data dependence relationship between subtasks. The stronger the dependence relationship, the more complex the task. The resource requirements refer to the computing resources, storage resources, etc. required to complete the task. The greater the resource requirements, the more complex the task.

[0046] The task complexity evaluation function will comprehensively consider the above three factors, analyze the input task, and decompose it into multiple subtasks. At the same time, the evaluation function will also calculate the dependence relationship between subtasks, providing a basis for subsequent subtask scheduling and execution.

[0047] A complex image recognition task can be decomposed into multiple subtasks such as image preprocessing, feature extraction, model training, and result classification. Image preprocessing requires operations such as denoising and enhancement on the original image; feature extraction requires extracting key features from the preprocessed image; model training requires using the extracted features to train the recognition model; result classification requires using the trained model to classify the image. There are data dependence relationships between these subtasks. For example, feature extraction depends on the results of image preprocessing, and model training depends on the results of feature extraction. The task complexity evaluation function will generate a subtask dependence graph based on these dependence relationships, providing guidance for subsequent subtask scheduling and execution.

[0048] The role of the meta-learner is to quickly adapt to new subtasks by learning task processing experiences in different scenarios. The core of the meta-learner includes a scene feature extraction function and a fusion of basic model parameters.

[0049] The scene feature extraction function is used to extract the feature information of subtasks, such as task type, data scale, data distribution, etc. These feature information reflect the characteristics of subtasks and provide a basis for subsequent parameter adjustment.

[0050] The fusion of the basic model parameters refers to the fusion of scene features with the parameters of the pre-trained basic model to obtain adaptive parameters. The pre-trained basic model contains a large amount of general knowledge, which can provide basic support for the processing of subtasks. By fusing scene features with the basic model parameters, the model can better adapt to specific subtasks.

[0051] For an image classification subtask, the scene feature extraction function can extract features such as the resolution, color distribution, and object categories of the image. Then, these features are fused with the parameters of the pre-trained image classification model to obtain adaptive parameters. Using these adaptive parameters, the image classification model can be fine-tuned to better adapt to the current subtask.

[0052] The meta-learner uses the adaptive parameters to optimize the subtask and outputs the processing result of the subtask.

[0053] The knowledge cache module is used to store and manage the knowledge learned by the agent, improving the efficiency of the agent in processing tasks. The core of the knowledge cache module includes the construction of temporal feature vectors, the calculation of dynamic representation values, the generation of temporal association strengths, and the update of cache content.

[0054] The construction of temporal feature vectors refers to converting the processing result of the subtask into a temporal feature vector. This vector contains information such as the processing time, resource consumption, and result quality of the subtask, reflecting the processing process of the subtask.

[0055] The calculation of dynamic representation values refers to calculating the dynamic representation value of the subtask based on the temporal feature vector. The dynamic representation value reflects the importance and degree of change of the subtask, providing a basis for the subsequent evaluation of the importance of knowledge items. For example, if a subtask has a short processing time but high result quality, its dynamic representation value is high, indicating that the subtask is relatively important.

[0056] The generation of temporal association strengths refers to calculating the temporal association strength between subtasks based on the dynamic representation value and co-occurrence frequency of the subtasks. The co-occurrence frequency refers to the number of times two subtasks co-occur within a certain period of time. The temporal association strength reflects the dependency relationship between subtasks, providing a basis for the subsequent update of cache content.

[0057] The update of cache content refers to updating the cache content based on the importance of knowledge items and access frequencies. The higher the importance of a knowledge item and the higher its access frequency, the more important the knowledge item is and should be preferentially retained in the cache.

[0058] Suppose the agent frequently uses a certain feature extraction method when processing multiple image classification tasks. At this time, the knowledge item of this feature extraction method has a high importance and should be retained in the cache so that subsequent tasks can directly call it to improve processing efficiency.

[0059] The global optimization function is used to comprehensively consider performance evaluation metrics, task execution efficiency, and adaptability metrics, adjust the task decomposition strategy, optimize the knowledge cache, and form a complete optimization loop.

[0060] The performance evaluation metrics reflect the overall performance of the agent, the task execution efficiency reflects the speed and efficiency of the agent in processing tasks, and the adaptability metrics reflect the agent's ability to adapt to new tasks. The global optimization function will comprehensively consider the above three metrics and adjust the task decomposition strategy to enable the agent to better complete tasks.

[0061] The task execution results will be fed back to the meta-learner for optimizing the knowledge cache. The meta-learner will update the importance of knowledge items and adjust the cache content according to the task execution results, so that the cache can better support subsequent task processing.

[0062] The cache status will be fed back to the performance evaluation module for evaluating the cache effect. The performance evaluation module will adjust the weights of the performance evaluation metrics according to the cache status to make the evaluation results more accurate.

[0063] Through the global optimization function, modules such as performance evaluation, task decomposition, meta-learning, and knowledge caching can be organically combined to form a complete optimization loop, continuously improving the performance of the agent.

[0064] If the agent finds that the task decomposition strategy is not reasonable when processing a certain task, resulting in low task execution efficiency. At this time, the global optimization function will adjust the task decomposition strategy, decompose the task into smaller subtasks, or adjust the dependencies between subtasks, thereby improving the task execution efficiency. At the same time, the task execution results will be fed back to the meta-learner for optimizing the knowledge cache, so that the cache can better support subsequent task processing. The cache status will be fed back to the performance evaluation module for evaluating the cache effect, thereby adjusting the weights of the performance evaluation metrics to make the evaluation results more accurate.

[0065] Agent performance improvement: Through the coordinated action of mechanisms such as performance evaluation, parameter adjustment, task decomposition, meta-learning, and knowledge caching, the agent has a faster response speed, more accurate decision-making, and higher resource utilization. Task processing efficiency improvement: Through task decomposition and knowledge caching, the agent can better utilize existing knowledge and experience, quickly adapt to new tasks, and improve task processing efficiency. Agent adaptability enhancement: Through the meta-learning mechanism, the agent can learn task processing experience in different scenarios, quickly adapt to new subtasks, and improve the agent's adaptability.

[0066] In an alternative embodiment, a time-series feature vector is constructed based on the processing results of the subtasks, and a dynamicity representation value is calculated. The time-series correlation strength is generated by combining the representation value and the co-occurrence frequency. The importance of knowledge items is evaluated based on the correlation strength and the access frequency. Updating the cache content according to the importance includes: Construct a time-series feature vector, which includes the feature dimension information of knowledge items at multiple time points, and calculate the dynamicity representation value of the knowledge items based on the time-series feature vector; Input the dynamicity representation value into the correlation strength calculation module. The correlation strength calculation module calculates the initial correlation strength based on the co-occurrence frequency and the time decay function, and obtains the time-series correlation strength by enhancing the similarity of the dynamicity representation value; Construct a knowledge item importance evaluation module based on the time-series correlation strength. The knowledge item importance evaluation module fuses the time-weighted access frequency and the importance of associated neighbors to obtain the comprehensive importance of the knowledge item; Calculate the replacement probability between knowledge items according to the comprehensive importance, dynamically adjust the update threshold based on the change of the cache hit rate, and perform cache update according to the replacement probability and the update threshold; Input the dynamicity representation value, the time-series correlation strength, and the comprehensive importance into the performance evaluation module. The performance evaluation module optimizes the system parameters based on the gradient of the performance index, and feeds back the optimized system parameters to the feature extraction and correlation calculation links.

[0067] Extract the time-series features of the knowledge items. The system records the multiple feature dimension information of each knowledge item at different time points, such as the number of accesses, the modification time, the number of associated knowledge items, etc. These information constitute the time-series feature vector of the knowledge item. For example, if the number of accesses of knowledge item A on Monday, Tuesday, and Wednesday are 10, 15, and 12 respectively, the modification times are 10:00, 11:00, and 14:00 respectively, and the number of associated knowledge items are 3, 4, and 3 respectively, then the time-series feature vector of knowledge item A can be represented as {(Monday, 10, 10:00, 3), (Tuesday, 15, 11:00, 4), (Wednesday, 12, 14:00, 3)}.

[0068] Calculate the dynamicity representation value of the knowledge items. Based on the extracted time-series feature vector, the system calculates the dynamicity representation value of the knowledge items. The dynamicity representation value reflects the degree of change of the knowledge item features over time. For example, a statistic of the change amplitude of the feature values, such as variance or standard deviation, can be used to quantify the dynamicity of the knowledge item. Assuming that the standard deviation of the number of accesses of knowledge item A is 2, then its dynamicity representation value is 2.

[0069] Calculate the temporal correlation strength between knowledge items. The system first calculates the initial correlation strength based on the co-occurrence frequency of knowledge items. The co-occurrence frequency refers to the number of times two knowledge items are accessed within the same time period. For example, if knowledge items A and B are accessed together 5 times in a week, their co-occurrence frequency is 5. The system introduces a time decay function to reduce the impact of co-occurrence frequencies at earlier times on the correlation strength. For example, an exponential decay function is used, such that co-occurrence frequencies in the recent period contribute more to the correlation strength than those in the distant past. On this basis, the system enhances the initial correlation strength according to the similarity of the dynamic characterization values of knowledge items to obtain the final temporal correlation strength. For example, if the dynamic characterization values of knowledge items A and B are both relatively high and similar, their temporal correlation strength will be further enhanced. Suppose the initial correlation strength between knowledge items A and B is 0.5, and due to the similarity of their dynamic characterization values, the final temporal correlation strength is enhanced to 0.8.

[0070] Evaluate the importance of knowledge items. The system constructs a knowledge item importance evaluation module that fuses the time-weighted access frequency with the importance of associated neighbors to obtain the comprehensive importance of knowledge items. The time-weighted access frequency means that the weight of the recent access frequency is higher than that of the early access frequency. The importance of associated neighbors refers to the importance of other knowledge items with a relatively high correlation strength with this knowledge item. For example, if the time-weighted access frequency of knowledge item A is 0.7, and the importance of its associated knowledge items B and C are 0.8 and 0.6 respectively, and the correlation strengths between A and B, C are 0.8 and 0.5 respectively, then the comprehensive importance of knowledge item A can be calculated by weighted average. Suppose the calculation result is 0.75.

[0071] Update the cache content. The system calculates the replacement probability between knowledge items based on their comprehensive importance. The knowledge item with a lower importance is more likely to be replaced. The system dynamically adjusts the update threshold to maintain the cache hit rate within a reasonable range. For example, when the cache hit rate drops, the system will increase the update threshold to reduce the frequency of cache updates. The system performs cache update operations based on the replacement probability and the update threshold. For example, if the replacement probability of knowledge item A is higher than the update threshold, it will be removed from the cache, and a knowledge item with a higher importance will be added to the cache.

[0072] The performance evaluation module collects performance metrics during the system operation, such as cache hit rate, access latency, etc. Then, based on the gradients of these performance metrics, system parameters are optimized, such as the parameters of the time decay function, the adjustment strategy of the update threshold, etc. Finally, the optimized system parameters are fed back to the feature extraction and correlation calculation processes, thereby continuously improving the system performance.

[0073] Improve cache hit rate: By considering the dynamics, relevance, and access frequency of knowledge items, it is possible to more accurately predict which knowledge items will be accessed in the future, thereby increasing the cache hit rate. Reduce access latency: The improvement in cache hit rate can directly reduce access latency and speed up the acquisition of information. Adaptive optimization: The system can dynamically adjust parameters based on the feedback of performance metrics, so as to adapt to different application scenarios and data characteristics and achieve continuous performance optimization.

[0074] In an optional implementation manner, the knowledge cache module is used to perform semantic parsing on the task entity and construct an association network, train a decision model based on the association network, input the output result of the decision model into the workflow adjustment module to optimize the execution process, and the decision results generated by the proxy collaborative control mechanism include: Extract semantic entities in the education task, calculate the vector cosine similarity of the semantic entities through vector space mapping, calculate the point mutual information value by combining the co-occurrence frequency and marginal probability of the semantic entities in the corpus, and generate a task semantic association network by weighted fusion of the vector cosine similarity and the point mutual information value; Construct a state transition scoring function based on the task semantic association network, calculate the state transition probability by exponential normalization of the state transition scoring function, calculate the decision value based on the immediate reward and discount factor using the temporal difference method, and evaluate and optimize the decision sequence according to the state transition probability and the decision value; According to the decision sequence, count the number of decision nodes in the unit process to obtain the decision point density, calculate the nested depth and the number of branches of the process branch to obtain the branch complexity, count the number and levels of the feedback loops to obtain the feedback link number, calculate the process complexity by weighted summation of the decision point density, the branch complexity, and the feedback link number, and obtain the workflow adjustment factor through gradient calculation; Receive the workflow adjustment factor, score multiple intelligent agents in three dimensions of knowledge reserve, decision-making ability, and execution efficiency to obtain an evaluation score; construct a weighted average function based on the evaluation score, and fuse the decision outputs of multiple intelligent agents according to the weighted average function to obtain a collaborative decision result.

[0075] Extract key semantic entities in the education task description, such as "homework", "exam", "course", "student", "teacher", etc. These entities represent the core elements of the education task. Subsequently, map these semantic entities to a high-dimensional vector space, for example, use a pre-trained word vector model (such as Word2Vec, GloVe, or BERT) to convert each entity into a vector representation.

[0076] The term "assignment" can be transformed into a vector containing hundreds of numerical values, which captures the semantic information of "assignment" in a large text corpus. Then, the semantic association degree between entities is measured by calculating the cosine similarity between these vectors. The value of cosine similarity ranges from -1 to 1, and the closer it is to 1, the more similar the two entities are semantically. If the cosine similarity of the vectors of "assignment" and "exercise" is high, it indicates that they have a strong semantic association.

[0077] To more accurately reflect the association between entities, the co-occurrence frequency and marginal probability of entities in the educational corpus are further considered to calculate the point mutual information value. The point mutual information value measures the degree of difference between the probability of two entities co-occurring and their respective independent occurrence probabilities. If two entities often co-occur, their point mutual information value is high. For example, "student" and "assignment" often co-occur in the educational corpus, so their point mutual information value will be relatively high.

[0078] The vector cosine similarity and the point mutual information value are weighted and fused to generate a task semantic association network. The purpose of weighted fusion is to comprehensively consider the semantic similarity and statistical association between entities to more comprehensively characterize the relationship between entities. The weights can be adjusted according to the actual situation. For example, a higher weight can be given to the point mutual information value to emphasize the co-occurrence relationship between entities. The generated task semantic association network can be visualized as a graph, where nodes represent semantic entities, edges represent the association relationships between entities, and the weights of the edges represent the association strength.

[0079] Based on the constructed task semantic association network, a state transition scoring function is constructed. This function is used to evaluate the quality of transitioning to the next state after taking a certain action in a given state. The state can be defined as the completion status of the current task, the learning state of the student, etc. The action can be defined as adopting a certain teaching strategy, adjusting the task difficulty, etc. The state transition scoring function comprehensively considers multiple factors. For example, the efficiency of task completion, the learning effect of the student, the utilization rate of teaching resources, etc. The higher the score, the more favorable the state transition.

[0080] The state transition scoring function is normalized through an exponential function to obtain the state transition probability. The state transition probability represents the likelihood of transitioning to the next state after taking a certain action in a given state. The purpose of exponential normalization is to convert the score into a probability for subsequent decision optimization. For example, if the score given by the state transition scoring function is higher, the corresponding state transition probability is also higher.

[0081] The temporal difference method is used to calculate the decision value based on immediate rewards and discount factors. The temporal difference method is a reinforcement learning algorithm for learning the strategy of taking optimal actions in different states. Immediate rewards represent the rewards obtained immediately after taking a certain action. For example, the rewards obtained for completing a task. The discount factor represents the degree of attenuation of future rewards. For example, future rewards are not as important as current rewards. The decision value represents the expected value of the long-term rewards that can be obtained after taking a certain action in a certain state. By continuously iterating and updating the decision value, the optimal decision-making strategy can be learned. For example, if adopting a certain teaching strategy can improve the learning effect of students, the decision value of this strategy will gradually increase.

[0082] The decision sequence is evaluated and optimized according to the state transition probability and decision value. The decision sequence refers to a series of continuous decision-making processes. For example, a combination of a series of teaching strategies. The purpose of evaluating the decision sequence is to find the optimal decision sequence to maximize the long-term rewards. The decision sequence can be optimized through various methods. For example, the greedy algorithm can be used to select the optimal action at each step, or the dynamic programming algorithm can be used to solve the global optimal solution. For example, by evaluating different combinations of teaching strategies, the most suitable teaching plan for the current students can be found.

[0083] The number of decision nodes in the unit process is counted according to the decision sequence to obtain the decision point density. The decision point density reflects the frequency of decisions in the process. The nested depth and the number of branches of the process branch are calculated to obtain the branch complexity. The branch complexity reflects the complexity of the process. The number and levels of feedback loops are counted to obtain the feedback link number. The feedback link number reflects the degree of iteration of the process. The decision point density, branch complexity, and feedback link number are weighted and summed to calculate the process complexity, and the workflow adjustment factor is obtained through gradient calculation. The higher the process complexity, the more complex the process and the need for adjustment. The workflow adjustment factor represents the direction and amplitude of the adjustment required for the workflow. For example, if the process complexity is too high, the decision point density needs to be reduced, the number of branches needs to be decreased, or the number of feedback links needs to be reduced.

[0084] The workflow adjustment factor is received, and multiple intelligent agents are scored in three dimensions: knowledge reserve, decision-making ability, and execution efficiency to obtain the evaluation score. Intelligent agents can represent different teaching roles. For example, teachers, teaching assistants, intelligent tutoring systems, etc. Knowledge reserve reflects the amount of knowledge possessed by the intelligent agent. Decision-making ability reflects the ability of the intelligent agent to make reasonable decisions. Execution efficiency reflects the speed and quality of the intelligent agent to complete tasks. The evaluation score is an evaluation of the comprehensive ability of the intelligent agent. For example, an experienced teacher may score higher in terms of knowledge reserve and decision-making ability, while an intelligent tutoring system may score higher in terms of execution efficiency.

[0085] Construct a weighted average function based on the evaluation scores, and fuse the decision outputs of multiple intelligent agents according to the weighted average function to obtain a collaborative decision result. The purpose of the weighted average function is to comprehensively consider the advantages of different intelligent agents and make more reasonable decisions. The weights can be adjusted according to the evaluation scores of the intelligent agents. For example, higher weights can be given to the intelligent agents with higher scores. The collaborative decision result is the final decision obtained by comprehensively considering the opinions of multiple intelligent agents. For example, a teacher can formulate a personalized teaching plan based on the suggestions of the intelligent tutoring system and combined with their own teaching experience.

[0086] Through semantic parsing and correlation network construction, the requirements of educational tasks can be understood more accurately, providing a more reliable basis for subsequent decisions. Through decision model training and workflow optimization, teaching strategies can be automatically adjusted to improve teaching efficiency and quality. Through the agent collaborative control mechanism, the advantages of multiple intelligent agents can be comprehensively utilized to make more reasonable decisions and achieve personalized teaching.

[0087] In an alternative implementation, intelligent agent state data is collected according to the decision result to construct a state vector, the state transition sequence is calculated based on the state vector to predict the workflow evolution trend, scene features are extracted in combination with the workflow evolution trend and the subtask configuration scheme is optimized, and the configuration scheme is fed back to the performance evaluation link to form an optimization closed loop, including: Calculate the performance evaluation index based on the preset weight coefficient, and construct a state vector containing the performance evaluation index; perform time series sampling on the state vector to obtain the state transition sequence, calculate the mutual information and conditional entropy between adjacent states, and construct a Markov state transition matrix in combination with the state duration and transition frequency to predict the workflow evolution trend; Calculate the weighted combination of the task hierarchy depth, data dependency strength and computing resource requirements based on the workflow evolution trend to obtain the task complexity, decompose the original task into multiple subtasks based on the task complexity, and calculate the dependency strength matrix between the subtasks; Adopt a transfer learning method to extract knowledge features from historical execution data and scene features, fuse the knowledge features with the current scene parameters to obtain scene adaptability parameters, and optimize the configuration of subtasks based on the scene adaptability parameters; Calculate the trend features of the real-time feedback data during the execution process through the exponential moving average method, calculate the signal reliability in combination with the fluctuation amplitude and variance of the feedback signal, and perform weighted fusion on the feedback data based on the signal reliability; Input the weighted fusion feedback data into the global optimization function, update the system parameters using the gradient descent method, and apply the system parameters to the performance evaluation, state prediction and task decomposition links respectively to achieve the closed-loop optimization of the workflow.

[0088] To achieve the intelligent optimization of the workflow, it is first necessary to collect and analyze the state data of the agent and construct a state vector. This process is like a doctor diagnosing a patient, where various physical signs need to be collected to accurately judge the condition. Specifically, we need to define a series of key indicators that can reflect the running state of the agent, such as CPU utilization rate, memory occupancy rate, network latency, task completion time, and so on. These indicators, like a patient's blood pressure, heart rate, and body temperature, can objectively reflect the health status of the agent. Then, arrange these indicator values in a predetermined order to form a state vector. For example, the format of the state vector can be defined as [CPU utilization rate, memory occupancy rate, network latency, task completion time]. Suppose at a certain moment, the state of the agent is CPU utilization rate of 50%, memory occupancy rate of 70%, network latency of 10 ms, and task completion time of 5 s, then the corresponding state vector is [0.5, 0.7, 0.01, 5].

[0089] Based on these state vectors, it is necessary to construct a state transition sequence and then predict the evolution trend of the workflow. This is like a doctor needing to observe a patient for a period of time to understand the development trend of the condition. Specifically, we need to perform time-series sampling on the state vectors, that is, record the state of the agent at regular time intervals. For example, record the state vector once every 1 second, then a sequence of state vectors can be obtained. Then, we need to analyze this sequence to find the transition relationships between states. For example, if the CPU utilization rate continues to increase, it may mean that the workload is increasing; if the network latency suddenly increases, it may mean that there is a problem with the network. By analyzing these transition relationships, we can predict the evolution trend of the workflow, such as whether the load will continue to increase, whether the network will become more congested, and so on.

[0090] Based on the prediction of the workflow evolution trend, we need to extract scenario features and optimize the subtask configuration scheme. This is like a doctor needing to consider various situations of the patient comprehensively to formulate the best treatment plan. Specifically, we need to extract some key scenario features according to the evolution trend of the workflow. For example, if it is predicted that the load will continue to increase, then the scenario feature is "high load"; if it is predicted that the network will become more congested, then the scenario feature is "network congestion". Then, we need to optimize the configuration of subtasks according to these scenario features. For example, in a high-load scenario, the allocation ratio of the CPU can be increased; in a network congestion scenario, the task scheduling strategy can be optimized to reduce the network traffic volume.

[0091] To form an optimization closed-loop, it is necessary to feedback the optimization configuration plan to the performance evaluation link. This is like a doctor who needs to observe the treatment effect of a patient in order to continuously adjust the treatment plan. Specifically, we need to calculate the performance evaluation indicators based on the preset weight coefficients. For example, the performance evaluation indicators can be defined to include task completion time, resource utilization rate, error rate, etc., and a weight coefficient is set for each indicator. Then, apply the optimization configuration plan to the workflow and record the values of various performance evaluation indicators. For example, if the optimization configuration plan can reduce the task completion time, improve the resource utilization rate, and reduce the error rate, it means that this plan is effective.

[0092] In the performance evaluation link, it is necessary to construct a state vector containing performance evaluation indicators, perform time series sampling to obtain the state transition sequence, calculate the mutual information and conditional entropy between adjacent states, and construct a Markov state transition matrix by combining the state duration and transition frequency to predict the workflow evolution trend. This process is like weather forecasting, which requires collecting meteorological data and analyzing the changing rules to predict the future weather. Specifically, we need to use performance indicators such as task completion time, resource utilization rate, and error rate as elements of the state vector, and then perform time series sampling on these state vectors to obtain the state transition sequence. By calculating the mutual information and conditional entropy between adjacent states, we can understand the uncertainty of state transitions. For example, if the mutual information between adjacent states is very high, it means that the regularity of state transitions is very strong; if the conditional entropy is very high, it means that the uncertainty of state transitions is very high. Then, by combining the state duration and transition frequency, construct a Markov state transition matrix, and we can predict the transition probability between different states of the workflow, thus predicting the workflow evolution trend.

[0093] Based on the workflow evolution trend, we need to calculate the task complexity, decompose the original task into multiple subtasks, and calculate the dependency strength matrix between subtasks. This is like a software engineer developing a large software project, which needs to decompose the project into multiple modules and analyze the dependency relationships between modules. Specifically, we need to calculate the weighted combination of task hierarchy depth, data dependency strength, and computing resource requirements according to the workflow evolution trend to obtain the task complexity. The task hierarchy depth refers to the hierarchical relationship of tasks in the workflow, the data dependency strength refers to the degree of data dependency between tasks, and the computing resource requirements refer to the amount of computing resources required for tasks. Then, based on the task complexity, decompose the original task into multiple subtasks. For example, a complex task can be decomposed into multiple subtasks such as data preprocessing, feature extraction, model training, and result evaluation. Then, calculate the dependency strength matrix between subtasks to describe the dependency relationship between subtasks. For example, if the output of subtask A is the input of subtask B, then there is a dependency relationship between subtask A and subtask B.

[0094] To better optimize the subtask configuration, we need to adopt transfer learning methods to extract knowledge features from historical execution data and scenario features, fuse the knowledge features with the current scenario parameters to obtain scenario adaptation parameters, and optimize the subtask configuration based on the scenario adaptation parameters. This is like an experienced chef who can adjust the cooking method according to the characteristics of the ingredients and the cooking environment to make a delicious dish. Specifically, we can collect historical execution data, including task type, input data, execution time, resource consumption, and so on. Then, adopt transfer learning methods to transfer the knowledge in the historical data to the current scenario. For example, if the historical data shows that a certain type of task performs well in a certain scenario, then this task configuration plan can be transferred to the current scenario. At the same time, it is also necessary to fuse the knowledge features with the current scenario parameters to obtain scenario adaptation parameters. For example, the resource allocation of subtasks can be adjusted according to parameters such as the load situation and network status of the current scenario. Finally, based on the scenario adaptation parameters, optimize the subtask configuration, such as adjusting the priority of the subtasks, allocating more computing resources, and so on.

[0095] During the workflow execution, we need to calculate the trend features of the real-time feedback data in the execution process through the exponential moving average method, calculate the signal reliability by combining the fluctuation amplitude and variance of the feedback signal, and perform weighted fusion on the feedback data based on the signal reliability. This is like a stock analyst who needs to analyze the trend of stock prices to make investment decisions. Specifically, we can collect the real-time feedback data during the execution process, such as CPU utilization, memory occupancy, network latency, and so on. Then, use the exponential moving average method to calculate the trend features of these data. For example, if the trend of CPU utilization is rising, it means that the workload is increasing. At the same time, it is also necessary to combine the fluctuation amplitude and variance of the feedback signal to calculate the signal reliability. For example, if the fluctuation amplitude of a certain signal is large and the variance is high, it means that the reliability of this signal is low. Finally, based on the signal reliability, perform weighted fusion on the feedback data. For example, a higher weight can be assigned to signals with higher reliability, and a lower weight can be assigned to signals with lower reliability.

[0096] The feedback data after weighted fusion is input into the global optimization function, and the gradient descent method is used to update the system parameters. Then, the system parameters are respectively applied to the performance evaluation, state prediction, and task decomposition links to achieve the closed-loop optimization of the workflow. This is like an autonomous driving system that needs to continuously learn and adjust parameters to better adapt to complex road conditions. Specifically, we can define a global optimization function to measure the overall performance of the workflow. Then, the feedback data after weighted fusion is input into the global optimization function, and the gradient descent method is used to update the system parameters. For example, the weight coefficient of the performance evaluation link, the Markov state transition matrix of the state prediction link, the sub-task dependence intensity matrix of the task decomposition link, etc. can be updated. Finally, the updated system parameters are respectively applied to the performance evaluation, state prediction, and task decomposition links to achieve the closed-loop optimization of the workflow. Through continuous iteration and optimization, the performance of the workflow can reach the best state.

[0097] By collecting the intelligent agent state data and predicting the workflow evolution trend, potential performance bottlenecks can be predicted in advance, thus avoiding task failures caused by insufficient resources or improper configuration. Optimizing the sub-task configuration plan according to the prediction results can achieve intelligent allocation of resources, improve resource utilization rate, reduce task completion time, and thus improve the overall efficiency of the workflow. Adopting a closed-loop optimization mechanism can continuously learn and adjust system parameters to make the performance of the workflow reach the best state, so as to adapt to the changing application scenarios and requirements.

[0098] As Figures 2 - 4 shown, the method further includes: In this example, the educational scenario knowledge base K for only constructing knowledge is a collection of documents from various sources doc1... dlocN, where the number of documents N is very large. Each document dn is divided into multiple paragraphs plocn,1... plocn,Mn, where Mn (and m below) represents the paragraph index of the nth document, and each paragraph is embedded into a multi-dimensional embedding elocn,m through the neural encoder fkey. It should be noted that in addition to the existing vector-based knowledge representation, the present invention also adopts inverted index and graph index technologies to more easily and accurately find context-related data.

[0099] To improve the quality of AI agent process automation and reduce manual intervention, the present invention implements prompt engineering at the beginning stage of each subtask. However, simply exchanging responses cannot achieve effective multi-turn task-oriented communication because it inevitably faces instruction repetition and false responses. Therefore, it cannot advance a productive communication process and hinders the realization of meaningful solutions. Therefore, an initial prompt mechanism is adopted to initiate, maintain, and end the agent's communication to ensure a robust and efficient workflow. This mechanism consists of a guidance system prompt PI and an assistant system prompt PA. Then, through PI and PA, a guidance I and an assistant A are instantiated through a large language model:

[0100] where ρ is a process customization operation implemented through system message allocation.

[0101] The limited context length of common large language models usually restricts their ability to maintain a complete communication history among all AI agents and stages. To address this issue, the present invention accordingly segments the agent's context memory according to the sequential stages of the agent, thus forming two functionally different memory types: short-term memory and long-term memory. Short-term memory is used to maintain the continuity of the conversation within a single stage, while long-term memory is used to maintain context awareness among stages. Short-term memory records the current stage utterances of the AI agent to assist in context-aware decision-making. At time t in stage Pi, we denote the instructor's instruction as Iit and the assistant's response as Ait. The short-term memory M collects the utterances up to time t: Mit = [(Ii1, Ai1), (Ii2, Ai2) . . . (Ii_t, Ait)] At the next time step t + 1, a new instruction Iit+1 is generated using the current memory and then conveyed to the assistant to produce a new response Ait+1. The short-term memory is iteratively updated until the communication quantity reaches the upper limit |Mi|: Iit+1 = I(Mit), Ait+1 = A(Mit, Iit+1) Mit+1 = Mit ∪ (Iit+1, Ait+1) To perceive the conversation in the previous stage, the chat chain only transmits the solutions of the previous stage as long-term memory ˜M and integrates them at the beginning of the next stage, thus achieving cross-stage long-term conversation transmission: Ii+11 = ˜Mi ∪ Pi+1I, ˜Mi = τ(Mj |Mj|) Among them, P represents a predefined prompt that uniquely appears at the beginning of each stage.

[0102] By sharing only the solutions of each subtask instead of the entire communication history, the risk of the robotic process being overwhelmed by a large amount of information is minimized, thereby enhancing the focus on each task, encouraging more targeted cooperation, and promoting cross-stage context continuity.

[0103] In this example, the AI agent can execute multiple steps through integrated plugins (also known as tools) to collect relevant information instead of directly answering questions. Different from general plugins, the plugins in the example are mainly based on the database interaction mode. This design facilitates querying the database through natural language, simplifies the user query expression, and at the same time strengthens the query understanding and execution capabilities of the large language model. The database interaction mode includes two components: a schema analyzer that parses the schema into a representation of a structure that the large language model can understand, and a query executor that executes SQL queries on the database according to the natural language response of the large language model. In addition, the AI agent is also integrated with third-party services, such as the web search proposed in WebGPT, to execute tasks on another platform without leaving the chat. With these plugins, the AI agent proposed by the present invention can execute several end-to-end data analysis problems with powerful generation capabilities.

[0104] In this example, the AI agent system performs a sample query on response generation. K search results are sorted according to their cosine similarity with the query, and then the top J results (where J ≤ K) are inserted into the context part of the predefined prompt template. Finally, the large language model generates a response. The sample query improves the large language model's ability to process context information by adding additional context during the training or inference stage. The sample query enables the language model to have enhanced context understanding capabilities, improved reasoning and inference skills, and customized problem-solving capabilities. Since the performance of the sample query is sensitive to specific settings, including the prompt template, the selection of context examples, and the example order, etc., in the AI agent proposed by the present invention, several strategies are provided to formulate the prompt template. In addition, the AI agent applies privacy protection measures to mask personal information.

[0105] In this example, it solves the problem that in the educational scenario, users use natural language to query the database, generate complex SQL queries, and conduct multi-source knowledge base Q&A, thereby reducing the usage threshold and improving the user experience in the educational scenario. By transferring the AI intelligent agent for the construction and execution process to the proxy, the limitations of traditional large language models are overcome. This method combines a retrieval-enhanced generation knowledge system, self-learning technology, and a service-oriented multi-model framework, enabling the AI intelligent agent to automatically construct processes according to human instructions and identify the parts that require dynamic decision-making in the process. During the process execution, the AI intelligent agent will monitor the process and intervene in the processing of the dynamic decision-making part, thus achieving more flexible and intelligent automation, with innovativeness, practicality, and broad application prospects. The paradigm shift in the data interaction method proposed by the present invention provides a more natural, efficient, and secure data interaction method. This self-optimizing ability not only improves the flexibility and adaptability of the workflow but also reduces the dependence on manual intervention, thereby reducing operating costs and improving overall efficiency.

[0106] Figure 5 This is a comparison table graph of the performance indicators of the AI intelligent agent process automation method in the embodiments of the present invention: This figure is a comparison table of performance indicators, showing the comparison data and improvement ranges of the traditional method, the reinforcement learning-based method, and this method in seven key performance indicators. Specifically: In terms of the task completion rate, this method reaches 94.7%, which is a 20.6% increase compared to 78.5% of the traditional method and is also significantly better than 85.3% of the reinforcement learning-based method. The average response time is significantly reduced, from 325 ms of the traditional method to 143 ms, a decrease of 56.0%. The resource utilization rate is increased from 64.2% of the traditional method to 87.5%, with an overall increase of 36.3%. The state prediction accuracy rate reaches 92.4%, which is a 37.3% increase compared to 67.3% of the traditional method.

[0107] In terms of the scenario adaptability score, this method obtains 4.7 points, which is a 46.9% increase compared to 3.2 points of the traditional method. The sub-task execution efficiency is increased from 0.65 to 0.89, with an increase range of 36.9%. In terms of the closed-loop optimization convergence speed, this method only needs 5.3 rounds to converge, which is a 57.3% reduction compared to 12.4 rounds of the traditional method, indicating a faster optimization efficiency.

[0108] Overall, this method has achieved significant improvements in all seven indicators. The improvements in response time and convergence speed are the most obvious, with decreases of 56.0% and 57.3% respectively, indicating that this method has obvious advantages in terms of efficiency and performance optimization. Core indicators such as task completion rate and prediction accuracy rate have also achieved more than 35% improvements, fully demonstrating the effectiveness of this method.

[0109] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

[0110] In the second aspect of the embodiments of the present invention, an AI intelligent agent process automation system for the education scenario is provided, including: A first unit for obtaining a question-answer pair vector matrix through semantic segmentation and constructing a hierarchical knowledge graph, reconstructing the vector matrix based on the graph attention network to calculate the association strength between nodes, constructing a two-tower encoding network using the contrastive learning method, fusing multi-layer features using a cross-layer attention module and updating network parameters to generate an intelligent agent encoder model; A second unit for collecting intelligent agent operation data based on the intelligent agent encoder model and generating a performance evaluation vector, inputting the performance evaluation vector into a feedback collection mechanism to update operation parameters, calculating task complexity according to the operation parameters and realizing task decomposition, and processing subtasks through a meta-learner and establishing a knowledge cache module; A third unit for semantically parsing task entities using the knowledge cache module and constructing an association network, training a decision model based on the association network, inputting the output result of the decision model into a workflow adjustment module to optimize the execution process, and generating a decision result using an agent collaborative control mechanism; A fourth unit for collecting intelligent agent status data according to the decision result to construct a status vector, predicting the workflow evolution trend based on the status vector, extracting scenario features in combination with the workflow evolution trend and optimizing the subtask configuration scheme, and feeding back the configuration scheme to the performance evaluation link to form an optimization closed loop.

[0111] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0112] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0113] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0114] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. The AI ​​agent process automation method for educational scenarios is characterized by: include: Through semantic segmentation, we obtain the question-answer pair vector matrix and construct a hierarchical knowledge graph. Based on the graph attention network, we calculate the correlation strength between nodes and reconstruct the vector matrix. We use the contrastive learning method to build a dual-tower encoding network. We use the cross-layer attention module to fuse multi-layer features and update network parameters to generate an intelligent agent encoder model. Based on the agent encoder model, the agent operation data is collected and a performance evaluation vector is generated, the performance evaluation vector is input into the feedback collection mechanism to update the operation parameters, the task complexity is calculated according to the operation parameters and the task is decomposed, the subtasks are processed by the meta-learner and a knowledge cache module is established; The knowledge cache module is used to perform semantic analysis on the task entity and construct an association network, a decision model is trained based on the association network, the output result of the decision model is input into the workflow adjustment module to optimize the execution process, and the agent collaborative control mechanism is used to generate the decision result; According to the decision results, the agent state data is collected to construct a state vector, and the state transition sequence is calculated based on the state vector to predict the workflow evolution trend. In combination with the workflow evolution trend, the scene features are extracted and the subtask configuration plan is optimized, and the configuration plan is fed back to the performance evaluation link to form an optimization closed loop.

2. The method according to claim 1, characterized in that The vector matrix of question-answer pairs is obtained through semantic segmentation and a hierarchical knowledge graph is constructed. The vector matrix is ​​reconstructed based on the graph attention network to calculate the correlation strength between nodes. The dual-tower encoding network is constructed using the contrastive learning method. The cross-layer attention module is used to fuse multi-layer features and update network parameters to generate the agent encoder model, including: Receiving a question-answer pair training corpus, performing semantic segmentation on the question-answer pair training corpus to obtain a question segment set and an answer segment set, respectively extracting representation vectors of the question segment set and the answer segment set using a pre-trained language model, and constructing an initial vector matrix for the question-answer pair; Constructing a hierarchical knowledge graph based on the initial vector matrix of the question-answer pair, mapping the question fragment set to a node vector, mapping the answer fragment set to an edge vector, and calculating the association strength matrix between the node vector and the edge vector through a graph attention network; Reconstructing the initial vector matrix of the question-answer pair according to the association strength matrix, calculating the mutual information loss between the positive sample question-answer pairs and the contrast loss between the negative sample question-answer pairs by using a contrastive learning method, and taking the weighted sum of the mutual information loss and the contrast loss as a training objective function; Based on the training objective function, a dual-tower coding network is constructed, wherein the dual-tower coding network includes a question coding tower and an answer coding tower, wherein the question coding tower and the answer coding tower adopt a transformer structure with shared parameters, and the parameters of the dual-tower coding network are optimized by a gradient descent method; Adaptively fuse the intermediate layer feature maps of the dual-tower encoding network to construct a cross-layer attention module, wherein the cross-layer attention module calculates the interaction weights between features of different layers based on the association strength matrix to generate an enhanced feature representation; The enhanced feature representation is input into the decoder for reconstruction, the reconstruction error is calculated and the parameters of the dual-tower coding network are updated, and the trained dual-tower coding network is used as the encoder model of the intelligent agent to realize the vectorized representation of question-answer pairs.

3. The method according to claim 1, characterized in that Collecting agent operation data based on the agent encoder model and generating a performance evaluation vector, inputting the performance evaluation vector into a feedback collection mechanism to update operation parameters, calculating task complexity and implementing task decomposition according to the operation parameters, processing subtasks through a meta-learner and establishing a knowledge cache module include: Constructing a performance evaluation module, wherein the performance evaluation module collects the operation data of the intelligent agent, calculates the performance evaluation indicators including response delay, decision accuracy and resource utilization based on preset weight coefficients, and generates a performance evaluation vector; A feedback collection mechanism is constructed using the performance evaluation vector, the feedback collection mechanism records real-time feedback information, calculates parameter adjustment gradients according to the real-time feedback information, and dynamically updates the operating parameters of the intelligent agent; Input the operating parameters into a task complexity evaluation function, which analyzes the input task based on the task hierarchy depth, data dependency, and resource demand, and decomposes it into multiple subtasks, while calculating the dependency relationship between the multiple subtasks; Constructing a meta-learner to receive the plurality of sub-tasks, wherein the meta-learner generates scene features through a scene feature extraction function, fuses the scene features with basic model parameters to obtain adaptive parameters, and optimizes the sub-tasks using the adaptive parameters; Based on the processing results of the subtasks, a temporal feature vector is constructed and a dynamic characterization value is calculated, the temporal association strength is generated by combining the characterization value with the co-occurrence frequency, the importance of the knowledge item is evaluated according to the association strength and the access frequency, and the cache content is updated according to the importance; The performance evaluation index, task execution efficiency and adaptability index are input into the global optimization function, the output of the global optimization function is used to adjust the task decomposition strategy, the task execution result is fed back to the meta-learner for optimizing the knowledge cache, and the cache status is fed back to the performance evaluation module to form an optimization closed loop.

4. The method according to claim 3, characterized in that Based on the processing results of the subtasks, a time series feature vector is constructed and a dynamic characterization value is calculated. The time series association strength is generated by combining the characterization value and the co-occurrence frequency. The importance of the knowledge item is evaluated according to the association strength and the access frequency. The cache content is updated according to the importance, including: Constructing a time series feature vector, wherein the time series feature vector includes feature dimension information of the knowledge item at multiple time points, and calculating a dynamic representation value of the knowledge item based on the time series feature vector; The dynamic characterization value is input into an association strength calculation module, the association strength calculation module calculates the initial association strength based on the co-occurrence frequency and the time decay function, and obtains the temporal association strength according to the similarity enhancement of the dynamic characterization value; Based on the temporal association strength, a knowledge item importance evaluation module is constructed, wherein the knowledge item importance evaluation module fuses the time-weighted access frequency with the importance of associated neighbors to obtain the comprehensive importance of the knowledge item; Calculating the replacement probability between knowledge items according to the comprehensive importance, dynamically adjusting the update threshold based on the change of cache hit rate, and performing cache update according to the replacement probability and the update threshold; The dynamic characterization value, the temporal correlation strength and the comprehensive importance are input into a performance evaluation module, and the performance evaluation module optimizes system parameters based on the gradient of performance indicators, and feeds back the optimized system parameters to the feature extraction and correlation calculation link.

5. The method according to claim 1, characterized in that The knowledge cache module is used to perform semantic analysis on the task entity and construct an association network, a decision model is trained based on the association network, the output result of the decision model is input into the workflow adjustment module to optimize the execution process, and the agent collaborative control mechanism is used to generate the decision result, including: Extracting semantic entities in educational tasks, calculating vector cosine similarity of the semantic entities through vector space mapping, calculating point mutual information value by combining co-occurrence frequency and edge probability of semantic entities in corpus, and weighted fusion of vector cosine similarity and point mutual information value to generate task semantic association network; A state transition scoring function is constructed based on the task semantic association network, the state transition scoring function is normalized by exponential to calculate the state transition probability, a time series difference method is used to calculate the decision value based on the immediate reward and the discount factor, and the decision sequence is evaluated and optimized according to the state transition probability and the decision value; According to the decision sequence, the number of decision nodes in the unit process is counted to obtain the decision point density, the nesting depth and the number of branches of the process branches are calculated to obtain the branch complexity, the number and level of feedback loops are counted to obtain the number of feedback links, the decision point density, the branch complexity and the number of feedback links are weightedly summed to calculate the process complexity, and the workflow adjustment factor is obtained by gradient calculation; The workflow adjustment factor is received, and multiple intelligent agents are scored in three dimensions of knowledge reserve, decision-making ability and execution efficiency to obtain an evaluation score; a weighted average function is constructed based on the evaluation score, and the decision outputs of multiple intelligent agents are integrated according to the weighted average function to obtain a collaborative decision result.

6. The method according to claim 1, characterized in that According to the decision result, the agent state data is collected to construct a state vector, the state transition sequence is calculated based on the state vector to predict the workflow evolution trend, the scene features are extracted and the subtask configuration scheme is optimized in combination with the workflow evolution trend, and the configuration scheme is fed back to the performance evaluation link to form an optimization closed loop, including: Calculate the performance evaluation index based on the preset weight coefficient, and construct a state vector containing the performance evaluation index; perform time-series sampling on the state vector to obtain a state transition sequence, calculate the mutual information and conditional entropy between adjacent states, and construct a Markov state transition matrix based on the state duration and transition frequency to predict the workflow evolution trend; According to the workflow evolution trend, a weighted combination of task hierarchy depth, data dependency strength and computing resource requirements is calculated to obtain task complexity; based on the task complexity, the original task is decomposed into multiple subtasks, and a dependency strength matrix between the subtasks is calculated; A transfer learning method is used to extract knowledge features from historical execution data and scene features, the knowledge features are integrated with current scene parameters to obtain scene adaptability parameters, and subtasks are optimized based on the scene adaptability parameters; The trend characteristics of the real-time feedback data during the execution process are calculated by the exponential moving average method, the signal reliability is calculated by combining the fluctuation amplitude and variance of the feedback signal, and the feedback data is weighted and fused based on the signal reliability; The weighted fused feedback data is input into the global optimization function, the system parameters are updated by the gradient descent method, and the system parameters are respectively applied to the performance evaluation, state prediction and task decomposition links to achieve closed-loop optimization of the workflow.

7. An AI agent process automation system for educational scenarios, used to implement the method described in any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain the question-answer pair vector matrix through semantic segmentation and construct a hierarchical knowledge graph, reconstruct the vector matrix based on the graph attention network to calculate the correlation strength between nodes, construct a dual-tower encoding network using the contrastive learning method, and use the cross-layer attention module to fuse multi-layer features and update network parameters to generate an intelligent agent encoder model; The second unit is used to collect agent operation data based on the agent encoder model and generate a performance evaluation vector, input the performance evaluation vector into the feedback collection mechanism to update the operation parameters, calculate the task complexity according to the operation parameters and realize task decomposition, process subtasks through the meta-learner and establish a knowledge cache module; The third unit is used to use the knowledge cache module to perform semantic analysis on the task entity and construct an association network, train a decision model based on the association network, input the output result of the decision model into the workflow adjustment module to optimize the execution process, and use the agent collaborative control mechanism to generate a decision result; The fourth unit is used to collect the agent state data according to the decision result to construct a state vector, calculate the state transition sequence based on the state vector to predict the workflow evolution trend, extract the scene characteristics and optimize the subtask configuration plan in combination with the workflow evolution trend, and feed back the configuration plan to the performance evaluation link to form an optimization closed loop.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Power transmission and distribution production task cooperation system and method based on intelligent agent

    CN120338452A

  • Virtual workplace training task allocation method and system based on multi-agent collaboration

    CN120355204A

  • Education Internet of Things intelligent collaboration method and system based on multi-modal protocol self-adaption

    CN120475081A

  • Workflow dynamic reconstruction method and system based on multi-source architecture change perception

    CN120634485A

  • Operation and maintenance technology service remote guidance interaction method and system combined with agent assistance

    CN120670563A