Education agent personalized learning path dynamic planning method and system

By constructing an educational intelligent body state matrix and collaborative learning groups, the learning path is dynamically optimized, which solves the problem of low learning efficiency in existing technologies, realizes real-time adjustment and optimization of personalized learning paths, and improves learning efficiency and resource utilization efficiency.

CN120689171AInactive Publication Date: 2025-09-23BEIJING XUEHAI YUNYING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510751257.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies are unable to adapt to the dynamic changes in the state of knowledge mastery during the learning process, ignore the promoting role of collaborative learning, resulting in low learning efficiency and lack of real-time optimization mechanism.

Method used

By acquiring the learners' knowledge mastery data, constructing the state matrix of the educational intelligent agent, calculating the learning eigenvector, generating the state transition network, calculating the transition probability and cost, dynamically programming the search, marking the edges to be optimized, forming a collaborative learning group based on the intelligent agent collaborative matrix, adjusting the transfer cost, and realizing dynamic optimization of the personalized learning path.

Benefits of technology

It achieves precise planning of learning paths for educational intelligent entities, improves learning efficiency and quality of knowledge acquisition, solves knowledge bottleneck problems, adapts to changing needs in the learning process, and improves the efficiency of educational resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689171A_ABST
    Figure CN120689171A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic planning method and system for a personalized learning path of an education agent, and relates to the technical field of artificial intelligence education, and the method comprises the steps: calculating a learning feature vector through obtaining knowledge mastering data of a learner, constructing a state transition network and a transition matrix, and executing dynamic planning search to obtain a state transition path. And marking the transfer edge lower than a threshold value as a to-be-optimized edge, building a collaborative learning group based on the agent collaborative matrix, calculating the combined knowledge gain optimization transfer cost, and re-executing the dynamic planning optimization learning path. According to the method, accurate individuation and dynamic optimization of the learning path are realized, and the learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence education technology, and in particular to a method and system for dynamically planning personalized learning paths for educational intelligent agents. Background Art

[0002] With the rapid development of online education, adaptive learning path planning tailored to learners' individual needs has attracted widespread attention. Existing technologies primarily generate fixed learning paths by statically analyzing learners' knowledge mastery levels and learning goals. However, this approach cannot adapt to the dynamic changes in knowledge mastery during the learning process, resulting in low learning efficiency.

[0003] Traditional learning path planning methods often model learners as independent individuals, ignoring the role of collaborative learning in promoting knowledge acquisition. Furthermore, due to the lack of real-time monitoring and evaluation mechanisms for the rate of knowledge acquisition during the learning process, learning bottlenecks cannot be identified and optimized in a timely manner, affecting overall learning outcomes.

[0004] Existing technologies for dynamically adjusting learning paths generally employ simple path replanning strategies, failing to fully leverage the knowledge spillover effects of collaborative learning to optimize learning paths. Therefore, a method for dynamically optimizing learning paths based on collaborative learning effects is urgently needed, enabling real-time adjustment and optimization of personalized learning paths for educational agents. Summary of the Invention

[0005] The embodiments of the present invention provide a method and system for dynamically planning personalized learning paths for educational agents, which can solve the problems in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a method for dynamically planning a personalized learning path for an educational agent, comprising:

[0007] Obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix;

[0008] A state transition network is constructed based on the learning feature vector, where nodes represent learning states and transition edges between nodes represent learning behaviors. The state transition probability and transition cost are calculated based on the learning feature vector to generate a state transition matrix.

[0009] Perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edges with acquisition rates lower than the preset rate threshold as edges to be optimized;

[0010] Generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form a collaborative learning group based on the agent collaboration matrix with the agents with the highest degree of complementarity;

[0011] Calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to achieve dynamic optimization of the personalized learning path of the educational agent.

[0012] In an optional embodiment,

[0013] Obtaining the learner's knowledge mastery data, calculating the educational agent state matrix, and calculating the learning feature vector of each educational agent based on the educational agent state matrix include:

[0014] Obtaining learners' answer data on knowledge points, including the accuracy rate, completion time, and number of requests for help during the answering process; standardizing the answer data to generate a knowledge point mastery matrix;

[0015] Calculating the correlation coefficients between knowledge points based on the knowledge point mastery matrix, constructing a knowledge association network, calculating the migration difficulty coefficient of each knowledge point based on the node connectivity of the knowledge association network, and generating a knowledge migration capability parameter;

[0016] Collecting the learning time series corresponding to the answer data, calculating the mastery change rate per unit time based on the knowledge point mastery matrix, and generating the knowledge absorption capacity parameter based on the mastery change rate;

[0017] Performing regular backtesting during the question answering data collection process to obtain backtest data, constructing a forgetting curve based on the difference between the backtest data and the knowledge point mastery matrix, and generating a knowledge retention ability parameter based on the decay characteristics of the forgetting curve;

[0018] The knowledge transfer ability parameters, knowledge absorption ability parameters and knowledge retention ability parameters are combined to generate an ability parameter matrix, the ability parameter matrix is ​​combined with the knowledge point mastery matrix to construct an educational agent state matrix, principal component analysis is performed on the educational agent state matrix to generate a learning feature vector of the educational agent.

[0019] In an optional embodiment,

[0020] A state transition network is constructed based on the learning feature vector, where nodes represent learning states and transition edges between nodes represent learning behaviors. The state transition probability and transition cost are calculated based on the learning feature vector, and the state transition matrix is ​​generated, including:

[0021] Performing density cluster analysis on the learning feature vector to obtain a cluster center, and determining the cluster center as a state node, wherein the state node is used to represent the learning state;

[0022] Collecting learning feature vectors in a continuous time series, calculating the direction and magnitude of vector changes at adjacent time points based on the learning feature vectors, and constructing transfer edges between state nodes based on the change direction and magnitude, wherein the transfer edges are used to characterize the learning behavior;

[0023] Obtain historical transfer records between state nodes, calculate the characteristic vector distance between state nodes, calculate the transfer success rate between state nodes based on the characteristic vector distance and learner ability parameters, and generate the transfer probability between state nodes based on the transfer success rate;

[0024] Extracting the transfer time, knowledge mastery change, and capability parameter change between state nodes; calculating the transfer cost between state nodes based on the transfer time, knowledge mastery change, and capability parameter change;

[0025] The state nodes, transition edges, transition probabilities and transition costs are combined to generate a state transition matrix.

[0026] In an optional embodiment,

[0027] Perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edges with acquisition rates lower than the preset rate threshold as edges to be optimized.

[0028] Constructing a bimodal value function including knowledge difficulty cost and time consumption cost based on the state transition matrix, setting dynamic adjustment weights for the knowledge difficulty cost and the time consumption cost, and calculating the comprehensive cost of state transition according to the dynamic adjustment weights;

[0029] Performing a dynamic programming search based on the comprehensive cost, recording the optimal decision sequence for each state during the search process, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence;

[0030] Obtaining learning trajectory data for each transition edge on the state transition path, the learning trajectory data including a knowledge assessment score and a learning duration, identifying effective learning segments from the learning trajectory data, extracting changes in knowledge mastery and actual learning duration based on the effective learning segments, and calculating the knowledge acquisition rate on the state transition path;

[0031] Track and record the learner's performance stability in each knowledge state, determine the rate fluctuation range based on the performance stability, dynamically generate a rate threshold based on the rate fluctuation range, and mark the transfer edge with a knowledge acquisition rate lower than the rate threshold as an edge to be optimized.

[0032] In an optional embodiment,

[0033] Performing a dynamic programming search based on the comprehensive cost, recording the optimal decision sequence for each state during the search, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence includes:

[0034] Constructing a predictive evaluation function based on the comprehensive cost, applying the predictive evaluation function to candidate decision sequences in the state space, calculating expected benefits of the candidate decision sequences, sorting the candidate decision sequences based on the expected benefits, and selecting the state transition strategy with the highest expected benefit from the sorted results;

[0035] Setting a search window and adjusting the search window size based on a state transition strategy, performing a dynamic programming search within the search window to obtain a state node set, recording the optimal decision information of each node in the state node set, and connecting the optimal decision information in a state transition order to form a decision sequence;

[0036] Extracting state transition nodes from the decision sequence, obtaining transition directions and transition costs of the state transition nodes, calculating connection weights between nodes based on the transition directions and transition costs, and reconstructing the decision sequence according to the connection weights;

[0037] The reconstructed decision sequence is used to perform state backtracking, an initial state transition path is generated based on the state backtracking result, the initial state transition path is optimized, and a state transition path from the current state to the target state is output.

[0038] In an optional embodiment,

[0039] Generate an agent collaboration matrix based on the incremental knowledge difference of the transfer edge to be optimized. Based on the agent collaboration matrix, the agents with the highest degree of complementarity are grouped into collaborative learning groups, including:

[0040] Obtaining knowledge evaluation data of the agent on the transfer edge to be optimized, calculating the temporal changes of the knowledge evaluation data to obtain a knowledge increment sequence, extracting the change pattern of the knowledge increment sequence to construct a learning feature vector;

[0041] Obtain learning feature vectors of the first agent and the second agent respectively and perform difference calculation, repeatedly perform the difference calculation until all pairs of agents are traversed, combine the differences to generate a feature difference matrix, calculate the learning pattern distance between the agents based on the feature difference matrix, evaluate the degree of complementarity between the agents based on the learning pattern distance, and perform normalization to obtain an initial complementarity matrix;

[0042] Extracting the knowledge dependency strength of the transfer edge to be optimized to generate a weight coefficient, multiplying the weight coefficient with the initial complementary matrix to obtain a weighted complementary matrix, and normalizing the weighted complementary matrix to generate an agent collaboration matrix;

[0043] In the agent cooperation matrix, an agent pair with the largest complementary value is identified to form a cooperation group, cooperation matrix values ​​of the remaining agents and the cooperation group are calculated, and the agent with the largest cooperation matrix value is selected to join the collaborative learning group.

[0044] In an optional embodiment,

[0045] Calculating the combined knowledge gain of the collaborative learning group on the edge to be optimized and adjusting the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain include:

[0046] Acquire learning test data for collaborative learning group members on the edge to be optimized. The learning test data includes each member's independent answer results and collaborative answer results. Compare each member's collaborative answer results with their independent answer results to generate answer result difference markers. The difference markers indicate questions that were answered correctly in the collaborative learning but incorrectly in the independent learning.

[0047] Extracting corresponding knowledge point distribution features based on the difference marks, analyzing the knowledge mastery level of each member of the collaborative learning group, generating a set of mastered knowledge points for each member, comparing the mastered knowledge point sets to identify knowledge complementarity relationships between members, and determining a member combination method that can produce the greatest complementary effect;

[0048] Counting the overall knowledge level of the collaborative learning group based on the member combination method, calculating the improvement of the overall knowledge level compared to independent learning, and using the improvement as the combined knowledge gain of the edge to be optimized;

[0049] The combined knowledge gain is compared with a preset gain threshold; when the combined knowledge gain is greater than the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​reduced according to a preset ratio; when the combined knowledge gain is less than or equal to the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​kept unchanged.

[0050] A second aspect of an embodiment of the present invention provides a dynamic planning system for personalized learning paths of an educational agent, comprising:

[0051] The first unit is used to obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix;

[0052] The second unit is used to construct a state transition network based on the learning feature vector, where the nodes represent the learning state and the transition edges between nodes represent the learning behavior. The state transition probability and transition cost are calculated based on the learning feature vector to generate the state transition matrix;

[0053] The third unit is used to perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edge with an acquisition rate lower than a preset rate threshold as an edge to be optimized;

[0054] The fourth unit is used to generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form the agents with the highest degree of complementarity into a collaborative learning group based on the agent collaboration matrix;

[0055] The fifth unit is used to calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to realize the dynamic optimization of the personalized learning path of the educational intelligent agent.

[0056] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0057] processor;

[0058] a memory for storing processor-executable instructions;

[0059] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0060] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0061] In this embodiment, by acquiring the learner's knowledge mastery data, constructing a state transition network and performing a dynamic programming search, the precise planning of the learning path of the educational intelligent agent is achieved. The optimal learning path can be tailored according to the individual characteristics of the learner, significantly improving the learning efficiency and the quality of knowledge acquisition. By calculating the knowledge acquisition rate and marking the edges to be optimized, combining the intelligent agent collaborative matrix to form a highly complementary learning group, the knowledge bottleneck problem in traditional learning methods is solved, enabling learners to obtain targeted support at difficult points and effectively avoid learning stagnation. A dynamic adjustment mechanism is adopted to recalculate the state transition cost based on the combined knowledge gain of the collaborative learning group, achieving real-time optimization of the learning path, adapting to the changing needs in the learning process, improving the efficiency of educational resource utilization, and having important value in promoting the development of personalized education and intelligent learning systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flow chart of a method for dynamically planning personalized learning paths for an educational agent according to an embodiment of the present invention;

[0063] Figure 2 This is a schematic diagram of dynamic programming search simulation based on the state transition matrix;

[0064] Figure 3 A diagram showing the performance comparison of dynamic programming searches based on predictive evaluation functions. DETAILED DESCRIPTION

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0066] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0067] Figure 1 Schematic diagram of the process of dynamic planning method for personalized learning path of educational agent according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0068] Obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix;

[0069] A state transition network is constructed based on the learning feature vector, where nodes represent learning states and transition edges between nodes represent learning behaviors. The state transition probability and transition cost are calculated based on the learning feature vector to generate a state transition matrix.

[0070] Perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edges with acquisition rates lower than the preset rate threshold as edges to be optimized;

[0071] Generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form a collaborative learning group based on the agent collaboration matrix with the agents with the highest degree of complementarity;

[0072] Calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to achieve dynamic optimization of the personalized learning path of the educational agent.

[0073] In an optional embodiment, obtaining the learner's knowledge mastery data, calculating the educational agent state matrix, and calculating the learning feature vector of each educational agent based on the educational agent state matrix includes:

[0074] Obtaining learners' answer data on knowledge points, including the accuracy rate, completion time, and number of requests for help during the answering process; standardizing the answer data to generate a knowledge point mastery matrix;

[0075] Calculating the correlation coefficients between knowledge points based on the knowledge point mastery matrix, constructing a knowledge association network, calculating the migration difficulty coefficient of each knowledge point based on the node connectivity of the knowledge association network, and generating a knowledge migration capability parameter;

[0076] Collecting the learning time series corresponding to the answer data, calculating the mastery change rate per unit time based on the knowledge point mastery matrix, and generating the knowledge absorption capacity parameter based on the mastery change rate;

[0077] Performing regular backtesting during the question answering data collection process to obtain backtest data, constructing a forgetting curve based on the difference between the backtest data and the knowledge point mastery matrix, and generating a knowledge retention ability parameter based on the decay characteristics of the forgetting curve;

[0078] The knowledge transfer ability parameters, knowledge absorption ability parameters and knowledge retention ability parameters are combined to generate an ability parameter matrix, the ability parameter matrix is ​​combined with the knowledge point mastery matrix to construct an educational agent state matrix, principal component analysis is performed on the educational agent state matrix to generate a learning feature vector of the educational agent.

[0079] For example, to obtain learners' knowledge mastery data, online learning platforms collect information about their performance on various knowledge points. This data includes three core dimensions: accuracy reflects the learner's mastery of the knowledge point, completion time reflects proficiency, and the number of requests for help indicates the learner's independent problem-solving ability. For example, for mathematical functions, Learner A achieved an accuracy rate of 0.75 on quadratic functions, an average completion time of 120 seconds, and two requests for help. For linear functions, Learner A achieved an accuracy rate of 0.90, an average completion time of 80 seconds, and zero requests for help.

[0080] When normalizing the answer data, we used the minimum and maximum normalization method to map each dimension to a range of 0 to 1. The accuracy rate data was used directly as is. Completion time was normalized using the inverse conversion to ensure that larger values ​​represent better performance. The number of requests for help was negatively converted. The normalized data formed a knowledge point mastery matrix, with each row representing a learner and each column representing a comprehensive mastery score for a knowledge point. Learner A's comprehensive mastery score for quadratic functions was 0.68, and for linear functions was 0.85.

[0081] When calculating correlation coefficients between knowledge points, we analyzed the distribution of mastery across all learners across different knowledge points. Using the Pearson correlation coefficient method, we found a correlation coefficient of 0.72 between the quadratic function and the linear function, indicating a strong correlation between the two knowledge points. We also constructed a knowledge association network, in which each node represents a knowledge point, and the edge weights indicate the strength of the association between the knowledge points. When the correlation coefficient exceeds the preset threshold of 0.6, a connection is established between the corresponding nodes.

[0082] Based on the topological structure of the knowledge association network, the node connectivity of each knowledge point is calculated. A knowledge point with a high connectivity indicates that it is associated with multiple other knowledge points, making it easier for learners to transfer knowledge to other related knowledge points after mastering the knowledge point. The linear function, as a foundational knowledge point, has a connectivity of 8, while the quadratic function has a connectivity of 5. Therefore, the transfer difficulty coefficients for the linear function and the quadratic function are set to 0.2 and 0.4, respectively. The smaller the transfer difficulty coefficient, the easier it is to transfer the knowledge point to other knowledge points. This is used to generate the knowledge transfer ability parameter.

[0083] When collecting learning time series corresponding to answer data, the changes in the learner's knowledge mastery over consecutive time periods are recorded. For example, for Learner A learning trigonometric functions, their mastery was 0.3 in the first week, increased to 0.6 in the second week, and reached 0.8 in the third week. The rate of change in mastery per unit time was calculated: the rate of change from the first week to the second week was 0.3 per week, and the rate of change from the second week to the third week was 0.2 per week. This rate of change reflects the learner's knowledge absorption rate. Based on the historical rate of change data, the learner's absorptive capacity curve was fitted to generate a knowledge absorptive capacity parameter. Learner A's trigonometric function absorptive capacity parameter was 0.75, indicating strong knowledge absorptive capacity.

[0084] During regular backtesting, learners are retested at fixed intervals after completing a particular knowledge point. For example, learner A's initial mastery of logarithmic functions dropped from 0.85 to 0.78 after one week on the backtest, then further to 0.72 after two weeks, before stabilizing at 0.70 after three weeks. A systematic analysis of the patterns of change between the backtest data and initial mastery revealed that the forgetting process is characterized by rapid growth at first and then slowing down. By fitting the forgetting data points, a personalized forgetting curve is constructed. The curve's decay rate and stability reflect the learner's knowledge retention characteristics. Learner A's logarithmic function knowledge retention parameter is 0.65, indicating moderate knowledge retention.

[0085] The ability parameter matrix is ​​constructed by combining parameters for knowledge transfer, absorption, and retention according to specific weights. Based on educational psychology research, weights for the three abilities are assigned: transfer ability with a weight of 0.4, absorption ability with a weight of 0.4, and retention ability with a weight of 0.2. Learner A's comprehensive ability parameter for function-related knowledge points is calculated to be 0.67. The ability parameter matrix and the knowledge point mastery matrix are concatenated row by row to form the educational agent state matrix. The matrix dimension is the number of learners multiplied by the number of knowledge points plus the number of ability parameters.

[0086] During principal component analysis, the state matrix of the educational agent is subjected to dimensionality reduction to extract key features. The covariance matrix of the state matrix is ​​calculated, and the main direction of change is determined through eigenvalue decomposition. The first few principal components with a cumulative contribution rate of 85% are retained as components of the learning eigenvector. Taking a state matrix containing 20 knowledge points and 3 ability parameters as an example, after principal component analysis, 8 principal components are retained to form an 8-dimensional learning eigenvector. Learner A's learning eigenvector is (0.72, -0.15, 0.43, 0.28, -0.08, 0.35, 0.19, -0.12). This vector comprehensively depicts the learner's personalized characteristics in terms of knowledge mastery and learning ability, providing a data foundation for subsequent personalized learning path planning.

[0087] In this embodiment, multi-dimensional modeling and quantitative analysis of the learner's knowledge mastery status can be achieved, comprehensively characterizing their learning ability characteristics. By integrating answering behavior data such as accuracy rate, completion time, and number of help requests, and combining the correlation between knowledge points and time series change trends, it is possible not only to dynamically evaluate the learner's knowledge mastery, but also to refine their knowledge transfer, absorption, and retention capabilities, thereby constructing a personalized and timely educational agent state matrix. Furthermore, representative learning feature vectors are extracted through principal component analysis, providing strong support for the formulation of precise teaching strategies, the recommendation of adaptive learning paths, and the prediction of long-term learning ability.

[0088] In an optional embodiment, a state transition network is constructed based on the learning feature vector, wherein nodes represent learning states and transition edges between nodes represent learning behaviors. State transition probabilities and transition costs are calculated based on the learning feature vector, and generating a state transition matrix includes:

[0089] Performing density cluster analysis on the learning feature vector to obtain a cluster center, and determining the cluster center as a state node, wherein the state node is used to represent the learning state;

[0090] Collecting learning feature vectors in a continuous time series, calculating the direction and magnitude of vector changes at adjacent time points based on the learning feature vectors, and constructing transfer edges between state nodes based on the change direction and magnitude, wherein the transfer edges are used to characterize the learning behavior;

[0091] Obtain historical transfer records between state nodes, calculate the characteristic vector distance between state nodes, calculate the transfer success rate between state nodes based on the characteristic vector distance and learner ability parameters, and generate the transfer probability between state nodes based on the transfer success rate;

[0092] Extracting the transfer time, knowledge mastery change, and capability parameter change between state nodes; calculating the transfer cost between state nodes based on the transfer time, knowledge mastery change, and capability parameter change;

[0093] The state nodes, transition edges, transition probabilities and transition costs are combined to generate a state transition matrix.

[0094] In this embodiment, a density-based spatial clustering algorithm is first used to process the eigenvector data of all learners. The neighborhood radius parameter is set to 0.3, the minimum neighborhood point parameter is set to 5, and the core points, boundary points and noise points are identified by calculating the neighborhood density of each eigenvector point. Taking an 8-dimensional eigenvector data set containing 1000 learners as an example, 15 high-density areas are identified as cluster centers. Each cluster center represents a typical learning state. The eigenvector of cluster center A is (0.8, 0.2, 0.6, 0.4, 0.1, 0.7, 0.3, 0.5), which represents a learning state with a solid mathematical foundation and strong logical reasoning ability; the eigenvector of cluster center B is (0.4, 0.7, 0.3, 0.8, 0.6, 0.2, 0.9, 0.1), which represents a learning state with outstanding language comprehension ability and strong memory ability. These 15 cluster centers are determined as state nodes, and each state node carries a specific learning ability combination feature.

[0095] The process of collecting learning feature vectors in a continuous time series covers the learner's entire learning cycle. The system collects learner feature vectors every other day, forming time series data. Learner A's feature vectors on the first day are (0.6, 0.3, 0.5, 0.4, 0.2, 0.6, 0.4, 0.3), and on the second day are (0.65, 0.32, 0.52, 0.43, 0.25, 0.58, 0.45, 0.35). The Euclidean distance between feature vectors at adjacent time points is calculated as the magnitude of change. The magnitude of change for this learner between the two days is 0.08. The direction of change is determined by calculating the trend of each dimension of the feature vector. An increase of 0.05 in the first dimension indicates an improvement in that dimension, while a decrease of 0.02 in the sixth dimension indicates a slight decline in that dimension. When the magnitude of change exceeds the threshold of 0.05 and the change has a clear direction, a transition edge is constructed between the corresponding state nodes. Learner A transfers from state node C to state node D, and the system establishes a directed transfer edge between node C and node D.

[0096] Historical transition records are obtained based on state transition trajectory data from a large number of learners. The number of transitions between state nodes for all learners over the past six months was counted. The number of transitions from state node A to state node B was 156, with 200 attempted transitions. Eigenvector distance calculation uses a weighted Euclidean distance method, assigning different weights to different dimensions to reflect the importance of each ability dimension. The eigenvector distance between state nodes A and B is 0.42, indicating a moderate difference in learning ability characteristics between the two states. Learner ability parameters are derived through a comprehensive assessment of learners' basic abilities, learning motivation, and cognitive style. Ability parameters range from 0 to 1, with higher values ​​indicating stronger learning ability. Combining the eigenvector distance of 0.42 and the learner ability parameter of 0.7, the transfer success rate is calculated to be 0.78, indicating a 78% probability of success for this learner transitioning from state A to state B.

[0097] The generation of state transition probabilities takes into account both historical success rates and individual variability. A mapping relationship is established between transition success rates and learner ability parameters. Learners with higher ability parameters have higher success rates on difficult transition paths, while learners with lower ability parameters are more stable on simpler transition paths. The base transition probability from state node A to state node B is 0.65. For learners with an ability parameter of 0.7, the system applies an ability correction factor of 1.2, adjusting the final transition probability to 0.78. The corresponding transition probabilities are calculated for each pair of state nodes, forming the basic data for the transition probability matrix.

[0098] The calculation of the transfer cost involves the quantitative assessment of three key factors. Transfer time reflects the learning effort required for learners to complete the state transition. The average time required to transition from state A to state B is 5 days, meaning that learners need to invest 5 days of learning time to achieve the change in ability state. The change in knowledge mastery measures the degree of improvement in the learner's knowledge level during the state transition. During this transition, the learner's comprehensive knowledge mastery increased from 0.6 to 0.75, a change of 0.15. The change in ability parameter reflects the improvement in the learner's comprehensive learning ability. The learner's ability parameter increased from 0.65 to 0.72, a change of 0.07. The comprehensive transfer cost was calculated by integrating the three factors using a weighted summation method, with a weight of 0.5 for time cost, 0.3 for knowledge mastery change, and 0.2 for ability parameter change. The final cost value of this transfer path was 3.2.

[0099] The construction of the state transition matrix integrates all state nodes, transition edges, transition probabilities, and transition costs into a unified data structure. The rows and columns of the matrix correspond to the source and target state nodes, respectively, and the matrix elements contain two numerical values: transition probabilities and transition costs. For a system with 15 state nodes, the state transition matrix is ​​a 15-by-15 square matrix. The value of the element at position (1, 3) is (0.78, 3.2), indicating that the probability of transitioning from state node 1 to state node 3 is 0.78 and the cost is 3.2. For state node pairs without a direct transition relationship, the corresponding matrix position is set to (0, ∞), indicating that direct transition is impossible and the cost is infinitely high. A dynamic update mechanism is also maintained to regularly adjust the transition probabilities and transition costs based on newly added learner behavior data to ensure that the state transition matrix accurately reflects the current learning state transition patterns. This matrix provides core data support for the subsequent dynamic learning path planning algorithm, enabling the calculation of the optimal learning path based on the learner's current state and target state.

[0100] This embodiment enables refined modeling and dynamic optimization of the learning paths of educational agents, improving the accuracy and feasibility of personalized learning recommendations. By constructing state nodes based on learning feature vectors, the learner's cognitive state at different stages is accurately portrayed. State transition edges are constructed by introducing feature change trends to truly reflect the trajectory of learning behavior. Combining transition probabilities and transition costs allows for the assessment of the feasibility and efficiency of different learning paths, supporting dynamic planning of personalized paths under multi-objective constraints, and enhancing the coherence, adaptability, and rationality of resource allocation in the learning process.

[0101] In an optional embodiment, performing a dynamic programming search based on a state transition matrix, calculating a state transition path from a current state to a target state, calculating a knowledge acquisition rate on the state transition path, and marking a transition edge having an acquisition rate lower than a preset rate threshold as an edge to be optimized includes:

[0102] Constructing a bimodal value function including knowledge difficulty cost and time consumption cost based on the state transition matrix, setting dynamic adjustment weights for the knowledge difficulty cost and the time consumption cost, and calculating the comprehensive cost of state transition according to the dynamic adjustment weights;

[0103] Performing a dynamic programming search based on the comprehensive cost, recording the optimal decision sequence for each state during the search process, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence;

[0104] Obtaining learning trajectory data for each transition edge on the state transition path, the learning trajectory data including a knowledge assessment score and a learning duration, identifying effective learning segments from the learning trajectory data, extracting changes in knowledge mastery and actual learning duration based on the effective learning segments, and calculating the knowledge acquisition rate on the state transition path;

[0105] Track and record the learner's performance stability in each knowledge state, determine the rate fluctuation range based on the performance stability, dynamically generate a rate threshold based on the rate fluctuation range, and mark the transfer edge with a knowledge acquisition rate lower than the rate threshold as an edge to be optimized.

[0106] In this implementation, a state transition matrix is ​​first constructed. This matrix represents the possible transition paths a learner can take from one knowledge point to another in the knowledge graph. Each element in the state transition matrix represents the probability of transitioning from state i to state j. For example, in a math course containing 100 knowledge points, the state transition matrix is ​​a 100×100 matrix, where the value of 0.85 for the element (25, 37) indicates that the probability of transitioning from knowledge point 25 to knowledge point 37 is 0.85.

[0107] A bimodal value function is constructed based on the state transition matrix, comprehensively considering both the knowledge difficulty cost and the time consumption cost. The knowledge difficulty cost reflects the difficulty of mastering a particular knowledge point and can be quantified by the error rate or number of attempts in historical learning data. For example, for a high school student learning trigonometric functions, the knowledge difficulty cost might be 7.5 on a scale of 0-10. The time consumption cost represents the time investment required to learn a particular knowledge point and can be determined by the average historical learning time. For example, learning matrix multiplication takes an average of 120 minutes.

[0108] To balance these two costs, the system dynamically adjusts weights. Initially, the knowledge difficulty weight α can be set to 0.6, and the time consumption weight β can be set to 0.4. The system adjusts these weights based on the learner's learning style and progress. For example, if a learner performs better under time pressure, α can be gradually adjusted to 0.5 and β to 0.5; if a learner needs more time to deeply understand the concepts, α can be adjusted to 0.7 and β to 0.3. The combined cost is calculated as the knowledge difficulty cost multiplied by α plus the time consumption cost multiplied by β.

[0109] The optimal path search is performed using a dynamic programming algorithm. For each state s, the minimum cumulative cost from the initial state to s is calculated, and the previous optimal state leading to that state is recorded. For example, in an algebra course, the cumulative cost of the learning path "Basic Algebra → Linear Equations → Quadratic Equations → Polynomials" might be determined to be 320, which is lower than the cumulative cost of another possible path, such as "Basic Algebra → Exponential Functions → Quadratic Equations → Polynomials" (385). The optimal predecessor state of each state node is recorded to form a decision sequence. After the algorithm completes, the optimal state transition path is generated by backtracking from the target state to the initial state.

[0110] To assess learning efficiency, we collect learning trajectory data for each transition edge along the state transition path. This data includes a knowledge assessment score (e.g., a scale from 0 to 100) and learning duration (in minutes). For example, a student learning "Matrix Inversion" might have the following trajectory data: 40 minutes of study on day 1, with a score of 65; 35 minutes of study on day 2, with a score of 72; and 50 minutes of study on day 3, with a score of 85.

[0111] Identify effective learning segments from the learning trajectory, excluding data from distracted or fatigued states. Effective learning segments are determined by: continuous focus for more than 15 minutes; improved assessment scores compared to the previous session; and a reasonable frequency of interaction during learning activities (e.g., at least three interactions every five minutes). In the example above, the 25-minute segment on Day 1 and the 40-minute segment on Day 3 might be identified as effective learning segments.

[0112] Based on effective learning segments, the change in knowledge mastery (e.g., from 65 to 85, a 20-point improvement) and the actual time invested (e.g., 65 minutes) are calculated. The knowledge acquisition rate is calculated as the change in mastery divided by the actual time invested, for example, 20 minutes / 65 minutes = 0.31 minutes / minute. This rate is compared with a preset threshold to identify learning paths that need optimization.

[0113] Track the stability of learners' performance across knowledge states and quantify it by calculating the standard deviation across multiple consecutive assessments. For example, a student's scores on the "Probability Distribution" knowledge point over the last five assessments were 78, 82, 75, 80, and 79, with a standard deviation of approximately 2.59, indicating relatively stable performance. On the other hand, on the "Differential Equations" knowledge point, scores were 65, 82, 58, 75, and 63, with a standard deviation of approximately 9.69, indicating greater fluctuation in performance.

[0114] Determine the rate fluctuation range based on performance stability. For knowledge points with stable performance (such as standard deviation <5), a narrower fluctuation range (such as ±10%) can be set; for knowledge points with unstable performance (such as standard deviation >8), a wider fluctuation range (such as ±25%) can be set. Dynamically generate rate thresholds based on the historical average knowledge acquisition rate and its fluctuation range. For example, if a student's average acquisition rate of algebraic knowledge is 0.25 points / minute and the fluctuation range is ±15%, the rate threshold can be set to 0.25×(1-0.15)=0.2125 points / minute.

[0115] Edges with a knowledge acquisition rate below a threshold are marked as edges to be optimized. For example, if a student's knowledge acquisition rate on the "Linear Algebra → Matrix Eigenvalue" edge is 0.18 minutes / minute, which is below the threshold of 0.2125 minutes / minute, then this edge is marked as being optimized. For these edges, alternative learning paths are recommended, additional learning resources are provided, or learning strategies are adjusted to improve overall learning efficiency.

[0116] Based on the above technical solution, dynamic adjustment and intelligent optimization of personalized learning paths can be achieved, improving the scientific nature and adaptability of learning path planning. By constructing a bimodal cost function that integrates knowledge difficulty and time consumption, learning efficiency and cognitive burden can be dynamically balanced. Combining dynamic planning search with the optimal decision backtracking mechanism, efficient learning paths from the current state to the target state can be quickly identified. Learning trajectory data is used to evaluate the knowledge acquisition rate and set dynamic thresholds to accurately identify inefficient learning links and mark them for optimization, providing a basis for subsequent path reconstruction, thereby improving the closed-loop adjustment capability and intelligent guidance level of the learning system.

[0117] Figure 2 This is a schematic diagram of dynamic programming search simulation based on the state transfer matrix, such as Figure 2 As shown in the figure, scatter plot visualization allows for the intuitive identification of 36% of edges to be optimized (red triangles). The knowledge acquisition rate of these learning transfer edges is below the dynamic threshold of 0.2 minutes per minute, providing a clear optimization target for subsequent path reconstruction. Simulation results show that 52% of normal learning edges maintain good efficiency, with an average knowledge acquisition rate of 0.23 minutes per minute, demonstrating the effectiveness of the bimodal value function in balancing the cost of knowledge difficulty and time consumption.

[0118] Based on 10,000 Monte Carlo samplings, the rationality of dynamically generating rate thresholds according to the learner's performance stability was verified, and differentiated fluctuation range adjustment strategies were provided for the 22% stable state edges and the 16% fluctuating state edges.

[0119] In an optional embodiment, performing a dynamic programming search based on the comprehensive cost, recording an optimal decision sequence for each state during the search, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence includes:

[0120] Constructing a predictive evaluation function based on the comprehensive cost, applying the predictive evaluation function to candidate decision sequences in the state space, calculating expected benefits of the candidate decision sequences, sorting the candidate decision sequences based on the expected benefits, and selecting the state transition strategy with the highest expected benefit from the sorted results;

[0121] Setting a search window and adjusting the search window size based on a state transition strategy, performing a dynamic programming search within the search window to obtain a state node set, recording the optimal decision information of each node in the state node set, and connecting the optimal decision information in a state transition order to form a decision sequence;

[0122] Extracting state transition nodes from the decision sequence, obtaining transition directions and transition costs of the state transition nodes, calculating connection weights between nodes based on the transition directions and transition costs, and reconstructing the decision sequence according to the connection weights;

[0123] The reconstructed decision sequence is used to perform state backtracking, an initial state transition path is generated based on the state backtracking result, the initial state transition path is optimized, and a state transition path from the current state to the target state is output.

[0124] This embodiment provides a state transition path generation method. First, a predictive evaluation function is constructed, which comprehensively considers the three dimensions of time cost, cognitive load and learning effect. The time cost weight is set to 0.4, the cognitive load weight is set to 0.3, and the learning effect weight is set to 0.3. For the candidate path from state node A to state node E, the time cost of the calculated path ABDE is 12 days, the cognitive load is 0.7, and the expected learning effect is 0.85. After weighted calculation, the predicted evaluation value of the path is 0.72. The time cost of another candidate path ACE is 8 days, the cognitive load is 0.5, the expected learning effect is 0.75, and the evaluation value is 0.68. The predictive evaluation function is applied to all feasible candidate decision sequences in the state space to generate the corresponding expected benefit value for each path.

[0125] The expected benefit value of candidate decision sequences is calculated by combining the dual factors of path length and node quality. The analysis shows that path ABDE contains 4 state nodes, with a total path length of 3 transfer steps, and the average benefit of each transfer step is 0.24. Path ACE contains 3 state nodes, with a total path length of 2 transfer steps, and the average benefit of each transfer step is 0.34. When calculating the comprehensive expected benefit value of the path, the single-step benefit is combined with the path efficiency. The comprehensive benefit value of path ABDE is 0.58, and the comprehensive benefit value of path ACE is 0.64. All candidate decision sequences are sorted in descending order based on the expected benefit value. The sorting results show that path ACE ranks first and path ABDE ranks second. From the sorting results, the system selects path ACE with the highest expected benefit value as the current optimal state transition strategy.

[0126] The search window is set using an adaptive adjustment mechanism to balance search efficiency and result quality. The initial search window size is set to 5 state nodes. When the selected state transition strategy path is short, the search window is reduced to 3 nodes to improve search accuracy. When the strategy path is long, the search window is expanded to 8 nodes to cover more possibilities. Based on the currently selected state transition strategy ACE, the search window is adjusted to 3 nodes, covering state nodes A, B, and C. A dynamic programming search is performed within the search window, gradually expanding the search boundary to calculate the optimal cumulative cost for each reachable state node. The search process results in a set of state nodes including nodes A, B, C, D, and E. The optimal predecessor node and the minimum cost to reach each node are recorded. The optimal decision information for node A is that the starting node has no predecessor and the cost is 0. The optimal decision information for node C is that the predecessor node A has a cost of 2.5. The optimal decision information for node E is that the predecessor node C has a cost of 5.1.

[0127] The formation of a decision sequence is achieved by connecting the optimal decision information in the order of state transitions. Starting from the target node E, tracing back, extract the predecessor node C of node E, and then extract the predecessor node A of node C, forming a reverse decision chain ECA. Reversing the reverse decision chain yields a forward decision sequence ACE, which records the complete state transition path from the starting state to the target state. Each node in the decision sequence stores detailed decision information, including node identification, cumulative cost, transition probability, and expected learning effect. The decision information for node A is cumulative cost 0, transition probability 1.0, and expected effect 0.6. The decision information for node C is cumulative cost 2.5, transition probability 0.8, and expected effect 0.7. The decision information for node E is cumulative cost 5.1, transition probability 0.75, and expected effect 0.85.

[0128] State transition node extraction identifies key state change points within the decision sequence. By analyzing the decision sequence ACE, two pairs of state transition nodes were extracted: A to C and C to E. The direction vector for the transition from A to C was obtained by calculating the difference between the feature vectors of the two state nodes. The direction of the transition indicated enhanced mathematical and logical skills and language comprehension abilities, with a transition cost of 2.5, including a time cost of 1.5 days and a cognitive load cost of 1.0. The direction vector for the transition from C to E indicated a significant improvement in comprehensive analytical ability and practical application capabilities, with a transition cost of 2.6, including a time cost of 1.8 days and a cognitive load cost of 0.8. Connection weights between nodes were calculated based on the transition direction and cost, taking into account both the difficulty and benefits of the transition. The connection weight from A to C was 0.75, indicating that this transition path is highly feasible and profitable. The connection weight from C to E was 0.68, indicating that this transition path is slightly less feasible but still within an acceptable range.

[0129] The decision sequence reconstruction process optimizes the original sequence based on connection weights. Transfer paths with connection weights below a threshold of 0.5 are examined, and alternative transfer options are sought for those with weights too low. All connection weights in the current decision sequence, ACE, are above the threshold, maintaining the original sequence structure. Further checking for higher-weighted parallel paths reveals that the overall weight of path ABE is 0.65, lower than the average weight of the current paths, 0.715. Therefore, the current sequence, ACE, remains unchanged. The reconstructed decision sequence remains ACE, containing clear transfer instructions and execution parameters.

[0130] The state backtracking process uses the reconstructed decision sequence to generate a complete learning path plan. Starting from the target state E, backtracking is performed step by step back to the starting state A to verify the continuity and feasibility of the path. The backtracking results confirm that each transition step in path ACE has clear learning tasks and evaluation criteria. The initial state transition path includes specific learning content arrangements: from state A to state C, basic concept understanding and logical reasoning training are required, with an estimated learning time of 1.5 days; from state C to state E, comprehensive application exercises and practical projects are required, with an estimated learning time of 1.8 days. The initial state transition path is optimized, adjusting the difficulty gradient and time allocation of learning tasks to ensure that learners can smoothly complete state transitions. The optimized final path adds transitional learning tasks between states A and C, and reinforcement exercises between states C and E, forming a complete personalized learning path planning solution.

[0131] In the prior art, the planning of personalized learning paths mostly adopts static rule matching or preset path recommendations, which lacks real-time perception of learner state changes and knowledge mastery dynamics, and cannot efficiently generate adaptable learning paths in complex state spaces, resulting in limited path optimization effects and difficulty in covering multiple learning goals and state transition modes. This application constructs a comprehensive cost function that includes both knowledge difficulty and time consumption modes, combines a predictive evaluation function to sort the benefits of candidate decision sequences, introduces a dynamic search window mechanism, and achieves fine control of local state sets during the state search process, effectively improving the accuracy and flexibility of path selection. At the same time, the optimal decision of each state node is recorded, and the decision sequence is reconstructed using connection weights, making the state backtracking process more coherent and optimized. Compared with the existing solutions that are only based on static strategies or path templates, this implementation method can adaptively adjust the learning strategy, optimize the structure and effect of the learning path, and ensure the generation of high-quality, traceable, and personalized learning paths in a dynamic learning environment, thereby improving the efficiency of the learning process and the accuracy of goal achievement.

[0132] Figure 3 This is a schematic diagram comparing the performance of dynamic programming search based on predictive evaluation functions. The technical effects of the present invention are significantly reflected in the comprehensive improvement of four key indicators: the path generation accuracy reaches 91.3%, which is 22.8 percentage points higher than the 68.5% of the traditional static method and 13.1 percentage points higher than the 78.2% of the basic dynamic programming method; the search efficiency is improved to 86.7%, which is 41.4 percentage points higher than the traditional method; the learning effect optimization reaches 83.4%, which is 31.3 percentage points higher than the traditional method; and the adaptive adjustment capability achieves an excellent performance of 92.8%, which is 64.2 percentage points higher than the traditional method.

[0133] By constructing a bimodal comprehensive cost function that includes knowledge difficulty cost and time consumption cost, and combining it with a predictive evaluation function to accurately sort candidate decision sequences, the present invention achieves efficient personalized learning path generation in complex state space.

[0134] In an optional embodiment, generating an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and forming agents with the highest degree of complementarity into collaborative learning groups based on the agent collaboration matrix includes:

[0135] Obtaining knowledge evaluation data of the agent on the transfer edge to be optimized, calculating the temporal changes of the knowledge evaluation data to obtain a knowledge increment sequence, extracting the change pattern of the knowledge increment sequence to construct a learning feature vector;

[0136] Obtain learning feature vectors of the first agent and the second agent respectively and perform difference calculation, repeatedly perform the difference calculation until all pairs of agents are traversed, combine the differences to generate a feature difference matrix, calculate the learning pattern distance between the agents based on the feature difference matrix, evaluate the degree of complementarity between the agents based on the learning pattern distance, and perform normalization to obtain an initial complementarity matrix;

[0137] Extracting the knowledge dependency strength of the transfer edge to be optimized to generate a weight coefficient, multiplying the weight coefficient with the initial complementary matrix to obtain a weighted complementary matrix, and normalizing the weighted complementary matrix to generate an agent collaboration matrix;

[0138] In the agent cooperation matrix, an agent pair with the largest complementary value is identified to form a cooperation group, cooperation matrix values ​​of the remaining agents and the cooperation group are calculated, and the agent with the largest cooperation matrix value is selected to join the collaborative learning group.

[0139] This embodiment provides a method for constructing a collaborative learning group of intelligent agents. The method first obtains the knowledge evaluation data of each intelligent agent on the transfer edge to be optimized. The knowledge evaluation data includes the evaluation timestamp, the corresponding knowledge point identifier, the mastery score of the knowledge point, and the behavior record information of the learner in the process of answering questions. The above data are arranged in chronological order to form an evaluation data sequence with continuous time characteristics. For this sequence, the change value of the knowledge mastery score between adjacent time periods is calculated to form a knowledge increment sequence. Furthermore, the knowledge increment sequence is analyzed for change trends, and characteristic values ​​including the amplitude of the upward trend, the frequency of the downward trend, the length of the score stability interval, the distribution of mutation positions, etc. are extracted, and combined with the learning behavior records such as the answering time, the number of attempts, etc., to generate a learning feature vector for characterizing the learning process of the intelligent agent. The vector data dimension is set according to the knowledge point coverage and the behavior dimension to ensure that it contains enough information to support subsequent difference calculations.

[0140] The learning feature vectors of all agents are paired, and the difference combinations of any two agents across all feature dimensions are calculated to form a feature difference set. In practical applications, this set is organized as a vector difference table, where rows represent agent pairings and columns represent the difference values ​​between feature dimensions. Based on this feature difference table, a learning pattern distance is defined for each group of agents to measure the similarity and complementarity between the two agents in their learning trajectories. To mitigate the impact of dimensional differences, the differences in each feature dimension are normalized and then weighted together to form the final pattern distance value. The larger the normalized learning pattern distance, the more significant the performance difference between the two agents on the learning path, and the stronger their complementarity. After normalizing the pattern distances of all agent pairs, an initial complementarity matrix is ​​generated, which is used to represent the degree of complementarity between any two agents on the learning path.

[0141] Obtain the knowledge dependency strength information of the knowledge points associated with the transfer edge to be optimized. This dependency strength is calculated by the path density, number of sequential constraint relationships, and correlation between the knowledge point and its prerequisite knowledge in the knowledge graph. In the specific implementation, by tracking the learner's learning path between the starting knowledge state and the ending knowledge state of the transfer edge, the knowledge dependency chain is identified, and the actual dependency weight is extracted based on the learner's performance on the prerequisite knowledge point. This dependency weight is used as the weight coefficient and multiplied with the corresponding elements of the aforementioned initial complementary matrix to generate a weighted complementary matrix. The weighted complementary matrix can further reflect the importance of the complementarity between the two agents in a specific knowledge dependency context. To unify the scale, the weighted complementary matrix is ​​standardized and converted into an agent collaboration matrix with a value range within a fixed interval. This collaboration matrix is ​​used to quantify the degree of adaptation of any two agents in collaborative learning in the current learning transfer task.

[0142] Retrieve the agent pair with the largest complementary value from the agent collaboration matrix as the core members of the initial collaborative group. This agent pair shows obvious feature complementarity and learning strategy differences in the knowledge point learning path of the transfer edge, and can form knowledge transfer complementarity in collaborative learning. Based on the initial collaborative group construction, other agents are gradually introduced to expand the scale of the learning group. Calculate the average value of the collaboration matrix between each agent that has not yet joined the collaborative group and all members of the current collaborative group. This average value indicates the degree of compatibility between the agent and the existing learning group in terms of synergy. Select the agent with the largest average collaboration value to join the current collaborative group, and repeat the process until the preset threshold of the number of people in the group is met or the collaborative improvement effect is saturated.

[0143] To improve the efficiency of coordination and information flow within the collaborative group, agents joining the collaborative group are further divided into roles within the group. Leading learners and assisted learners are identified based on the prominent dimensions of each agent in the learning feature vector, such as the speed of mastery increase, stability level, and efficiency of answering behavior. Leading learners have a clear improvement trajectory within the target knowledge path and can provide learning strategy recommendations and knowledge point explanations to assisted learners, promoting two-way knowledge transfer in the collaborative learning process. During the operation of the collaborative group, the learning process data of the agents within the group is tracked in real time, the learning feature vector and collaborative matrix are updated, and the composition and role configuration of the group members are dynamically adjusted to adapt to the ever-changing knowledge state transfer path.

[0144] Suppose there are five agents, labeled A, B, C, D, and E. Their assessment data for knowledge points K1 to K2 is collected and used to generate their own knowledge increment sequences. For example, agent A's mastery scores in five assessments are 60, 65, 70, 75, and 80, respectively, and their response times are 90, 80, 85, 70, and 75 seconds. After processing, features such as mastery growth rate and response time change rate are extracted to form their feature vectors. The feature differences between agent pairs A and B, A and C, and so on, are calculated sequentially to generate a feature difference matrix, which is then used to obtain an initial complementarity matrix. After weighting coefficients are set based on the knowledge dependencies between K1 and K2, the matrix is ​​updated to a weighted complementarity matrix and normalized to a synergy matrix. In this matrix, A and C have the highest complementarity values ​​and are therefore selected as the core of the collaborative group. D and E are then introduced to form a four-member collaborative learning group. The collaborative strategy is continuously updated and optimized during the subsequent learning process.

[0145] This technical approach significantly improves the efficiency of knowledge acquisition at inefficient transfer edges within the learning path, enhances the adaptability and dynamic correction capabilities of personalized learning paths, and avoids learning stagnation caused by individual differences. Compared to existing methods that rely primarily on preset group groupings or static matching strategies, the proposed technical solution, based on dynamic feature analysis and adaptive collaborative construction, improves the rationality and collaborative effectiveness of learning group construction, and fully demonstrates the complementary and collaborative potential between intelligent agents in the knowledge acquisition process.

[0146] In an optional embodiment, calculating the combined knowledge gain of the collaborative learning group on the edge to be optimized, and adjusting the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain includes:

[0147] Acquire learning test data for collaborative learning group members on the edge to be optimized. The learning test data includes each member's independent answer results and collaborative answer results. Compare each member's collaborative answer results with their independent answer results to generate answer result difference markers. The difference markers indicate questions that were answered correctly in the collaborative learning but incorrectly in the independent learning.

[0148] Extracting corresponding knowledge point distribution features based on the difference marks, analyzing the knowledge mastery level of each member of the collaborative learning group, generating a set of mastered knowledge points for each member, comparing the mastered knowledge point sets to identify knowledge complementarity relationships between members, and determining a member combination method that can produce the greatest complementary effect;

[0149] Counting the overall knowledge level of the collaborative learning group based on the member combination method, calculating the improvement of the overall knowledge level compared to independent learning, and using the improvement as the combined knowledge gain of the edge to be optimized;

[0150] The combined knowledge gain is compared with a preset gain threshold; when the combined knowledge gain is greater than the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​reduced according to a preset ratio; when the combined knowledge gain is less than or equal to the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​kept unchanged.

[0151] For example, the test data of the members of the collaborative learning group are first collected. Taking a collaborative learning group consisting of three students A, B, and C as an example, their answers to a set of 50 questions in mathematics were recorded when they independently completed a test. Student A answered 35 questions correctly independently, and the wrong questions were mainly concentrated in algebraic operations and geometric proofs; student B answered 32 questions correctly independently, and the wrong questions were mainly concentrated in probability statistics and solid geometry; student C answered 30 questions correctly independently, and the wrong questions were mainly concentrated in function analysis and series problems. Subsequently, these three students were arranged to complete a set of test questions with similar structures again in a collaborative manner, and the answer results of each student in the collaborative process were recorded. It was found that student A answered 43 questions correctly, student B answered 40 questions correctly, and student C answered 38 questions correctly.

[0152] Each student's independent and collaborative answers were compared and analyzed, generating a discrepancy marker. For Student A, the system marked eight questions that they answered correctly during collaborative learning but incorrectly independently. These questions involved the application of the Bayesian formula in probability and statistics (two questions), cross-section problems in solid geometry (three questions), and sequence summation techniques (three questions). For Student B, the system marked eight questions that they answered correctly during collaborative learning but incorrectly independently. These questions involved function extreme value analysis (three questions), algebraic inequalities (three questions), and the application of auxiliary lines in geometric proofs (two questions). For Student C, the system marked eight questions that they answered correctly during collaborative learning but incorrectly independently. These questions involved solving algebraic equations (three questions), geometric proof ideas (two questions), and the basic principles of probability and statistics (three questions).

[0153] Based on the difference marks, the corresponding knowledge point distribution features are extracted. The knowledge points involved in the test questions are divided into 10 categories: K1 (basic algebra), K2 (advanced algebra), K3 (plane geometry), K4 (solid geometry), K5 (analytic geometry), K6 (trigonometric functions), K7 (function analysis), K8 (sequences), K9 (probability), and K10 (statistics). By analyzing the knowledge points to which the questions in the difference marks belong, a set of knowledge points mastered by each student is generated: Student A masters {K1, K2, K3, K5, K6}; Student B masters {K1, K3, K5, K6, K7}; Student C masters {K1, K4, K5, K6, K10}.

[0154] Analyze the complementary knowledge relationships between students and determine the optimal combination. By calculating knowledge point coverage, the system finds that the combination of A and B covers 7 knowledge points {K1, K2, K3, K5, K6, K7}, the combination of A and C covers 7 knowledge points {K1, K2, K3, K4, K5, K6, K10}, the combination of B and C covers 7 knowledge points {K1, K3, K4, K5, K6, K7, K10}, and the combination of A, B, and C covers 9 knowledge points {K1, K2, K3, K4, K5, K6, K7, K8, K10}. Therefore, the combination of A, B, and C is determined to produce the greatest complementary effect.

[0155] Based on the determined combination method, the combined knowledge gain is calculated. First, the overall knowledge level during independent learning is calculated. That is, the average number of questions answered correctly by the three students independently is (35 + 32 + 30) / 3 = 32.33 questions, and the accuracy rate is 32.33 / 50 = 64.7%. After collaborative learning, the average number of questions answered correctly by the three students is (43 + 40 + 38) / 3 = 40.33 questions, and the accuracy rate is 40.33 / 50 = 80.7%. Therefore, the combined knowledge gain is 80.7% - 64.7% = 16%.

[0156] The calculated combined knowledge gain of 16% is compared with the preset gain threshold of 10%. Since 16% is greater than 10%, the transfer cost of the edge to be optimized in the state transfer matrix is ​​reduced by a preset ratio of 30%. Specifically, the original transfer cost of the edge representing the collaborative learning combination in the state transfer matrix is ​​reduced from 100 to 70, a reduction of 30%.

[0157] The edge to be optimized represents the learning path from knowledge point K8 (sequences) to knowledge point K9 (probability). Before optimization, the transfer cost of this edge was 100, indicating a high learning difficulty. After optimization, the transfer cost dropped to 70, indicating that this collaborative learning combination can more effectively help members transition from learning sequence knowledge to learning probability knowledge, and the learning path is smoother.

[0158] The optimization results are displayed on the control panel, indicating that the collaborative learning group's collaborative efficiency on the "Sequence-Probability" learning path has increased by 30%. The group is recommended to maintain the current collaborative model. The system also automatically adjusts the parameters in the learning path planning algorithm. When planning subsequent learning paths for the group, paths that pass through this optimized edge will be prioritized, thereby improving overall learning efficiency.

[0159] Based on the above technical solution, it is possible to achieve the technical effect of dynamically evaluating the effect of collaborative learning and optimizing the learning path adjustment strategy accordingly. By comparing the differences in answering questions between collaborative learning and independent learning, the actual knowledge improvement areas brought about by collaboration are identified, and the synergistic complementary effects are further analyzed in combination with the distribution of knowledge points and the mastery of members, thereby accurately quantifying the degree of gain of group collaboration on knowledge mastery. This gain result is used to adjust the edge weights in the state transfer matrix, making the system more inclined to choose paths with significant synergistic effects in path planning, improving the efficiency and adaptability of learning paths, and effectively avoiding the waste of resources due to inefficient collaboration or no gain transfer, thereby enhancing the intelligent decision-making ability of personalized path dynamic planning.

[0160] A second aspect of an embodiment of the present invention provides a dynamic planning system for personalized learning paths for an educational agent, the system comprising:

[0161] The first unit is used to obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix;

[0162] The second unit is used to construct a state transition network based on the learning feature vector, where the nodes represent the learning state and the transition edges between nodes represent the learning behavior. The state transition probability and transition cost are calculated based on the learning feature vector to generate the state transition matrix;

[0163] The third unit is used to perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edge with an acquisition rate lower than a preset rate threshold as an edge to be optimized;

[0164] The fourth unit is used to generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form the agents with the highest degree of complementarity into a collaborative learning group based on the agent collaboration matrix;

[0165] The fifth unit is used to calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to realize the dynamic optimization of the personalized learning path of the educational intelligent agent.

[0166] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0167] processor;

[0168] a memory for storing processor-executable instructions;

[0169] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0170] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0171] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dynamic planning method for personalized learning paths of educational agents, characterized by: include: Obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix; A state transition network is constructed based on the learning feature vector, where nodes represent learning states and transition edges between nodes represent learning behaviors. The state transition probability and transition cost are calculated based on the learning feature vector to generate a state transition matrix. Perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edges with acquisition rates lower than the preset rate threshold as edges to be optimized; Generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form a collaborative learning group based on the agent collaboration matrix with the agents with the highest degree of complementarity; Calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to achieve dynamic optimization of the personalized learning path of the educational agent.

2. The method according to claim 1, characterized in that Obtaining the learner's knowledge mastery data, calculating the educational agent state matrix, and calculating the learning feature vector of each educational agent based on the educational agent state matrix include: Obtaining learners' answer data on knowledge points, including the accuracy rate, completion time, and number of requests for help during the answering process; standardizing the answer data to generate a knowledge point mastery matrix; Calculating the correlation coefficients between knowledge points based on the knowledge point mastery matrix, constructing a knowledge association network, calculating the migration difficulty coefficient of each knowledge point based on the node connectivity of the knowledge association network, and generating a knowledge migration capability parameter; Collecting the learning time series corresponding to the answer data, calculating the mastery change rate per unit time based on the knowledge point mastery matrix, and generating the knowledge absorption capacity parameter based on the mastery change rate; Performing regular backtesting during the question answering data collection process to obtain backtest data, constructing a forgetting curve based on the difference between the backtest data and the knowledge point mastery matrix, and generating a knowledge retention ability parameter based on the decay characteristics of the forgetting curve; The knowledge transfer ability parameters, knowledge absorption ability parameters and knowledge retention ability parameters are combined to generate an ability parameter matrix, the ability parameter matrix is ​​combined with the knowledge point mastery matrix to construct an educational agent state matrix, principal component analysis is performed on the educational agent state matrix to generate a learning feature vector of the educational agent.

3. The method according to claim 1, characterized in that A state transition network is constructed based on the learning feature vector, where nodes represent learning states and transition edges between nodes represent learning behaviors. The state transition probability and transition cost are calculated based on the learning feature vector, and the state transition matrix is ​​generated, including: Performing density cluster analysis on the learning feature vector to obtain a cluster center, and determining the cluster center as a state node, wherein the state node is used to represent the learning state; Collecting learning feature vectors in a continuous time series, calculating the direction and magnitude of vector changes at adjacent time points based on the learning feature vectors, and constructing transfer edges between state nodes based on the change direction and magnitude, wherein the transfer edges are used to characterize the learning behavior; Obtain historical transfer records between state nodes, calculate the characteristic vector distance between state nodes, calculate the transfer success rate between state nodes based on the characteristic vector distance and learner ability parameters, and generate the transfer probability between state nodes based on the transfer success rate; Extracting the transfer time, knowledge mastery change, and capability parameter change between state nodes; calculating the transfer cost between state nodes based on the transfer time, knowledge mastery change, and capability parameter change; The state nodes, transition edges, transition probabilities and transition costs are combined to generate a state transition matrix.

4. The method according to claim 1, wherein Perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edges with acquisition rates lower than the preset rate threshold as edges to be optimized. Constructing a bimodal value function including knowledge difficulty cost and time consumption cost based on the state transition matrix, setting dynamic adjustment weights for the knowledge difficulty cost and the time consumption cost, and calculating the comprehensive cost of state transition according to the dynamic adjustment weights; Performing a dynamic programming search based on the comprehensive cost, recording the optimal decision sequence for each state during the search process, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence; Obtaining learning trajectory data for each transition edge on the state transition path, the learning trajectory data including a knowledge assessment score and a learning duration, identifying effective learning segments from the learning trajectory data, extracting changes in knowledge mastery and actual learning duration based on the effective learning segments, and calculating the knowledge acquisition rate on the state transition path; Track and record the learner's performance stability in each knowledge state, determine the rate fluctuation range based on the performance stability, dynamically generate a rate threshold based on the rate fluctuation range, and mark the transfer edge with a knowledge acquisition rate lower than the rate threshold as an edge to be optimized.

5. The method according to claim 4, characterized in that Performing a dynamic programming search based on the comprehensive cost, recording the optimal decision sequence for each state during the search, and backtracking to generate a state transition path from the current state to the target state using the optimal decision sequence includes: Constructing a predictive evaluation function based on the comprehensive cost, applying the predictive evaluation function to candidate decision sequences in the state space, calculating expected benefits of the candidate decision sequences, sorting the candidate decision sequences based on the expected benefits, and selecting the state transition strategy with the highest expected benefit from the sorted results; Setting a search window and adjusting the search window size based on a state transition strategy, performing a dynamic programming search within the search window to obtain a state node set, recording the optimal decision information of each node in the state node set, and connecting the optimal decision information in a state transition order to form a decision sequence; Extracting state transition nodes from the decision sequence, obtaining transition directions and transition costs of the state transition nodes, calculating connection weights between nodes based on the transition directions and transition costs, and reconstructing the decision sequence according to the connection weights; The reconstructed decision sequence is used to perform state backtracking, an initial state transition path is generated based on the state backtracking result, the initial state transition path is optimized, and a state transition path from the current state to the target state is output.

6. The method according to claim 1, characterized in that Generate an agent collaboration matrix based on the incremental knowledge difference of the transfer edge to be optimized. Based on the agent collaboration matrix, the agents with the highest degree of complementarity are grouped into collaborative learning groups, including: Obtaining knowledge evaluation data of the agent on the transfer edge to be optimized, calculating the temporal changes of the knowledge evaluation data to obtain a knowledge increment sequence, extracting the change pattern of the knowledge increment sequence to construct a learning feature vector; Obtain learning feature vectors of the first agent and the second agent respectively and perform difference calculation, repeatedly perform the difference calculation until all pairs of agents are traversed, combine the differences to generate a feature difference matrix, calculate the learning pattern distance between the agents based on the feature difference matrix, evaluate the degree of complementarity between the agents based on the learning pattern distance, and perform normalization to obtain an initial complementarity matrix; Extracting the knowledge dependency strength of the transfer edge to be optimized to generate a weight coefficient, multiplying the weight coefficient with the initial complementary matrix to obtain a weighted complementary matrix, and normalizing the weighted complementary matrix to generate an agent collaboration matrix; In the agent cooperation matrix, an agent pair with the largest complementary value is identified to form a cooperation group, cooperation matrix values ​​of the remaining agents and the cooperation group are calculated, and the agent with the largest cooperation matrix value is selected to join the collaborative learning group.

7. The method according to claim 1, characterized in that Calculating the combined knowledge gain of the collaborative learning group on the edge to be optimized and adjusting the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain include: Acquire learning test data for collaborative learning group members on the edge to be optimized. The learning test data includes each member's independent answer results and collaborative answer results. Compare each member's collaborative answer results with their independent answer results to generate answer result difference markers. The difference markers indicate questions that were answered correctly in the collaborative learning but incorrectly in the independent learning. Extracting corresponding knowledge point distribution features based on the difference marks, analyzing the knowledge mastery level of each member of the collaborative learning group, generating a set of mastered knowledge points for each member, comparing the mastered knowledge point sets to identify knowledge complementarity relationships between members, and determining a member combination method that can produce the greatest complementary effect; Counting the overall knowledge level of the collaborative learning group based on the member combination method, calculating the improvement of the overall knowledge level compared to independent learning, and using the improvement as the combined knowledge gain of the edge to be optimized; The combined knowledge gain is compared with a preset gain threshold; when the combined knowledge gain is greater than the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​reduced according to a preset ratio; when the combined knowledge gain is less than or equal to the preset gain threshold, the transfer cost of the edge to be optimized in the state transfer matrix is ​​kept unchanged.

8. A dynamic planning system for personalized learning paths of educational agents, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the learner's knowledge mastery data, calculate the educational agent state matrix, and calculate the learning feature vector of each educational agent based on the educational agent state matrix; The second unit is used to construct a state transition network based on the learning feature vector, where the nodes represent the learning state and the transition edges between nodes represent the learning behavior. The state transition probability and transition cost are calculated based on the learning feature vector to generate the state transition matrix; The third unit is used to perform dynamic programming search based on the state transition matrix, calculate the state transition path from the current state to the target state, calculate the knowledge acquisition rate on the state transition path, and mark the transition edge with an acquisition rate lower than a preset rate threshold as an edge to be optimized; The fourth unit is used to generate an agent collaboration matrix based on the knowledge increment difference of the transfer edge to be optimized, and form the agents with the highest degree of complementarity into a collaborative learning group based on the agent collaboration matrix; The fifth unit is used to calculate the combined knowledge gain of the collaborative learning group on the edge to be optimized, adjust the transfer cost of the edge to be optimized in the state transfer matrix according to the combined knowledge gain, and re-execute the dynamic programming calculation based on the adjusted state transfer matrix to realize the dynamic optimization of the personalized learning path of the educational intelligent agent.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.