Individualized learning path generation method and device based on learning behavior portrait, equipment and medium

By acquiring multimodal learning behavior data and using inverse reinforcement learning algorithms to quantify cognitive resilience, a two-dimensional learner profile is constructed and a learning path for strategic dilemma tasks is generated. This solves the problem of neglecting cognitive resilience in existing systems and improves learners' analytical and perseverance abilities in complex problems.

CN121883218APending Publication Date: 2026-04-17XIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIANGNAN UNIV
Filing Date
2025-12-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing personalized learning systems overemphasize short-term knowledge transfer efficiency, neglecting the cultivation of learners' cognitive resilience when facing moderate challenges. This results in an inability to effectively develop learners' abilities to analyze complex problems, persist in exploration, and recover from setbacks.

Method used

By acquiring multimodal learning behavior data, using inverse reinforcement learning algorithms to deduce reward functions, quantifying cognitive resilience index, constructing a two-dimensional learner profile, and generating learning paths embedded with strategic dilemma tasks through multi-objective optimization algorithms, the learning paths are monitored and adjusted in real time to cultivate cognitive resilience.

Benefits of technology

It significantly enhances learners' analytical and perseverance abilities when facing complex and ambiguous problems, enabling a shift from short-term efficiency optimization to long-term growth optimization, and systematically cultivates learners' higher-order psychological traits when facing complex problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883218A_ABST
    Figure CN121883218A_ABST
Patent Text Reader

Abstract

The invention relates to a personalized learning path generation method and device based on a learning behavior portrait, equipment and a medium. The method comprises the steps of obtaining multi-modal learning behavior data of a target learner and performing feature extraction processing to obtain a standardized behavior feature sequence; reversely deducing a reward function from expert behaviors by using a reverse reinforcement learning technology to calculate a cognitive toughness index, and constructing a two-dimensional learner portrait fusing the knowledge mastery degree and the cognitive toughness; based on the two-dimensional learner portrait, generating an initial learning path of the embedded strategic dilemma through a multi-objective optimization algorithm; and dynamically updating the learner portrait by monitoring dilemma response data in real time during path execution, and adjusting subsequent path difficulty and support according to the updated learner portrait to form an adjusted learning path. By adopting the method, beneficial dilemma and cognitive toughness concepts in educational psychology can be converted into computable and operable engineering technology practices, and the deep understanding and long-term migration ability of knowledge is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of learning path generation technology, and in particular relates to a method, apparatus, device and medium for generating personalized learning paths based on learning behavior profiles. Background Technology

[0002] With the deep integration of online education platforms and artificial intelligence technology, personalized learning systems have become a key tool for improving teaching efficiency and learning experience. These systems analyze learners' knowledge acquisition status, answer history, and other behavioral data to plan a personalized learning path best suited to each learner's current cognitive level, thus achieving an initial shift from a "one-size-fits-all" approach to "individualized instruction." The current mainstream approach involves constructing a refined profile of the learner's knowledge status and using this profile to recommend subsequent learning content and activities.

[0003] In traditional technologies, the generation of personalized learning paths primarily relies on the assessment of learners' cognitive abilities. Systems typically track explicit indicators such as learners' accuracy and reaction time across various knowledge points, and use knowledge tracking models to predict their knowledge mastery. Building upon this, path generation algorithms aim to find the "optimal" path that most efficiently fills knowledge gaps and improves test scores. Their core objective is to help learners achieve their pre-set learning goals as quickly and directly as possible, minimizing frustration and obstacles in the learning process.

[0004] However, current personalized learning methods, or traditional approaches, have a fundamental limitation: an excessive pursuit of short-term "efficiency optimization." This design philosophy, which prioritizes a smooth and easy learning path, may have advantages in immediate knowledge transfer efficiency, but from a long-term educational psychology perspective, it essentially deprives learners of the valuable opportunity to develop their ability to analyze complex problems, persevere in exploration and experimentation, and recover from setbacks when faced with appropriate challenges. This psychological trait, known as "cognitive resilience," is crucial for deep understanding of knowledge, long-term memory, and the ability to solve complex real-world problems, yet current technologies completely neglect the modeling and cultivation of such implicit psychological traits. Summary of the Invention

[0005] Therefore, it is necessary to provide a personalized learning path generation method that can go beyond simply improving the efficiency of knowledge transfer and take into account both the cultivation of learners' cognitive resilience and the development of their long-term abilities, in order to address the aforementioned technical problems.

[0006] Firstly, this application provides a method for generating personalized learning paths based on learning behavior profiles, including:

[0007] The study acquires multimodal learning behavior data of the target learners and performs feature extraction processing on the multimodal learning behavior data to obtain a standardized behavioral feature sequence. The multimodal learning behavior data includes cognitive interaction data, behavioral micro data, and emotional physiological data.

[0008] Based on standardized behavioral feature sequences, the implicit reward function is derived from the behavioral demonstrations of predefined high cognitive resilience learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequences and the reward function, the cognitive resilience index of the target learner is calculated.

[0009] Obtain the knowledge mastery level of the target learners, and construct a two-dimensional learner profile based on the cognitive resilience index and knowledge mastery level;

[0010] Based on a two-dimensional learner profile and reward function, an initial personalized learning path is generated through a multi-objective optimization algorithm; wherein, the initial personalized learning path includes at least one strategic dilemma task;

[0011] During the process of the target learner executing the initial personalized learning path, the data on the target learner's dilemma response when interacting with strategic dilemma tasks is monitored in real time.

[0012] Based on the distress response data, the dual-dimensional learner profile is updated to obtain the updated dual-dimensional learner profile. Based on the updated dual-dimensional learner profile, the difficulty level and support level of subsequent tasks in the initial personalized learning path are adjusted to obtain the adjusted learning path.

[0013] Furthermore, based on standardized behavioral feature sequences, an implicit reward function is derived from the behavioral demonstrations of predefined high cognitive resilience learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequences and the reward function, the cognitive resilience index of the target learner is calculated, including:

[0014] The behavioral trajectories of learners assessed by experts as having high cognitive resilience are obtained from a pre-set historical database, and all behavioral trajectories are aggregated to generate an expert demonstration dataset.

[0015] Based on an expert demonstration dataset, a reward function is derived by using the maximum entropy inverse reinforcement learning algorithm; the reward function is used to explain the behavioral preferences of highly resilient learners.

[0016] The cognitive resilience index of the target learner is calculated based on the reward function and standardized behavioral feature sequence.

[0017] Furthermore, based on the reward function and standardized behavioral feature sequence, the cognitive resilience index of the target learner is calculated, including:

[0018] Based on the reward function, quantitative dimensions of cognitive resilience are determined; these quantitative dimensions include challenge endurance, strategy transferability, and emotional resilience.

[0019] Based on the standardized behavioral feature sequence, calculate the raw scores for each quantitative dimension;

[0020] The cognitive resilience index of the target learner is obtained by weighted summation of the raw scores using the following formula:

[0021]

[0022] in, For cognitive resilience index, To challenge the raw score of endurance, The raw score for strategy transfer ability. The raw score for emotional resilience. To challenge the overall mean of durability, This represents the overall mean of strategy transfer capability. This represents the overall mean of emotional resilience. To challenge the overall standard deviation of persistence, The overall standard deviation of strategy transfer capability. The overall standard deviation of emotional resilience. To challenge the weighting coefficients of persistence, The weighting coefficients for policy transfer capability. This is the weighting coefficient for emotional resilience.

[0023] Furthermore, based on a two-dimensional learner profile and reward function, an initial personalized learning path is generated using a multi-objective optimization algorithm; wherein the initial personalized learning path includes at least one strategic dilemma task, including:

[0024] Construct a multi-objective path optimization function, which includes knowledge acquisition benefits, cognitive resilience development benefits, and cognitive load costs:

[0025] Based on the knowledge mastery level in the dual-dimensional learner profile, a baseline value for knowledge mastery benefits is determined.

[0026] Based on the cognitive resilience index in the two-dimensional learner profile, the difficulty range of strategic dilemma tasks is determined.

[0027] The Monte Carlo tree search algorithm is used to sample in the path space to obtain candidate paths;

[0028] Based on the baseline value of knowledge acquisition benefits and the difficulty range, three-dimensional indicators for each candidate path are calculated; among them, the three-dimensional indicators include knowledge acquisition benefits, cognitive resilience cultivation benefits, and cognitive load costs.

[0029] The three-dimensional metrics of each candidate path are substituted into the multi-objective path optimization function to generate a value score for each candidate path; and the candidate path corresponding to the maximum value score is determined as the initial personalized learning path.

[0030] Furthermore, during the initial personalized learning path execution process, the target learner's dilemma response data during interaction with strategic dilemma tasks is monitored in real time, including:

[0031] Obtain the mouse movement trajectory coordinate sequence of the target learner in the strategic dilemma task interface, and calculate the hesitation index of the mouse movement trajectory coordinate sequence;

[0032] Acquire the keyboard input event stream of the target learner during the task process, and analyze the keyboard input event stream to obtain the input rate change pattern and modification frequency;

[0033] With authorization, facial video streams of the target learner are acquired, and micro-expression change sequences are extracted from the facial video streams using a pre-trained facial expression recognition model;

[0034] Calculate the percentage of duration of negative emotion intensity in a micro-expression change sequence;

[0035] The hesitation index, input rate change pattern, and duration of negative emotion intensity are combined to form the dilemma response data.

[0036] Secondly, this application also provides a personalized learning path generation device based on learning behavior profiles, including:

[0037] The data acquisition module is used to acquire multimodal learning behavior data of the target learners and perform feature extraction processing on the multimodal learning behavior data to obtain standardized behavioral feature sequences; among which, multimodal learning behavior data includes cognitive interaction data, behavioral micro data, and emotional physiological data;

[0038] The index calculation module is used to inversely deduce the implicit reward function from the behavioral demonstrations of predefined high cognitive resilience learners based on the standardized behavioral feature sequence and the reward function, and calculate the cognitive resilience index of the target learner based on the standardized behavioral feature sequence and the reward function.

[0039] The profile building module is used to obtain the knowledge mastery level of the target learner and to build a two-dimensional learner profile based on the cognitive resilience index and knowledge mastery level.

[0040] The initial path generation module is used to generate an initial personalized learning path based on a two-dimensional learner profile and a reward function, using a multi-objective optimization algorithm; wherein the initial personalized learning path includes at least one strategic dilemma task;

[0041] The data monitoring module is used to monitor in real time the dilemma response data generated by the target learner when interacting with strategic dilemma tasks during the execution of the initial personalized learning path;

[0042] The path update module is used to update the two-dimensional learner profile based on the dilemma response data, obtain the updated two-dimensional learner profile, and adjust the difficulty level and support level of subsequent tasks in the initial personalized learning path based on the updated two-dimensional learner profile, so as to obtain the adjusted learning path.

[0043] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement any of the personalized learning path generation methods based on learning behavior profiles described in the embodiments of this application.

[0044] Fourthly, this application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the personalized learning path generation method based on learning behavior profiles as described in any of the embodiments of this application.

[0045] The aforementioned personalized learning path generation method, device, equipment, and medium based on learning behavior profiles collect multimodal behavioral data of target learners and extract features. Then, using inverse reinforcement learning technology, a reward function is derived from expert behavior to quantify cognitive resilience, constructing a dual-dimensional learner profile integrating knowledge mastery and cognitive resilience. A multi-objective optimization algorithm is used to generate an initial learning path embedded with strategic dilemmas. During path execution, the learner profile is dynamically updated by monitoring dilemma response data in real time. The difficulty and support of subsequent paths are adjusted based on the updated learner profile, forming an adjusted learning path. This approach breaks through the limitations of traditional personalized learning systems that focus solely on knowledge transfer efficiency. It transforms the concepts of "beneficial dilemmas" and "cognitive resilience" from educational psychology into calculable and operable engineering practices, significantly improving the depth of knowledge understanding and long-term transferability. It also consciously and systematically cultivates learners' higher-order psychological traits such as analysis, perseverance, and adaptation when facing complex and ambiguous problems, achieving a fundamental shift from pursuing short-term "optimal efficiency" to promoting long-term "optimal growth." Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating a personalized learning path generation method based on learning behavior profiles in one embodiment.

[0048] Figure 2 This is a flowchart illustrating the steps of calculating the cognitive resilience index of a target learner based on the standardized behavioral feature sequence, the implicit reward function derived from the behavioral demonstrations of a predefined high cognitive resilience learner using an inverse reinforcement learning algorithm, and the standardized behavioral feature sequence and reward function.

[0049] Figure 3 This is a schematic diagram of a personalized learning path generation device based on learning behavior profiles in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] In one embodiment, such as Figure 1 As shown, a method for generating personalized learning paths based on learning behavior profiles is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0052] Step S101: Obtain multimodal learning behavior data of the target learner, and perform feature extraction processing on the multimodal learning behavior data to obtain a standardized behavioral feature sequence; wherein, the multimodal learning behavior data includes cognitive interaction data, behavioral micro data and emotional physiological data.

[0053] Feature extraction processing involves filtering, transforming, and refining key features that are representative and reflect the core information of raw data such as images, text, audio, and sensor signals. Cognitive interaction data includes the learner's choices, response time series, number of attempts, and error pattern classification during the question-answering process, which directly reflects the learner's cognitive processing efficiency and knowledge application level. Behavioral micro-data records detailed micro-behaviors during human-computer interaction, such as the coordinate trajectory of mouse movement, the spatiotemporal distribution of click events, page scrolling speed, and the interval of keyboard keystrokes. These subtle behavioral patterns provide important clues for indirectly inferring the learner's attention allocation, decision hesitation, and cognitive load. Emotional and physiological data refers to facial video streams and microphone-collected voice signals collected with the learner's informed consent and authorization, which are used to identify changes in the learner's emotional state and physiological arousal level.

[0054] For example, multimodal learning behavior data such as cognitive interaction data, behavioral micro data, and emotional physiological data of target learners can be obtained. An anomaly detection algorithm based on isolated forest can be used to clean various raw multi-source heterogeneous data. Then, time series alignment of multimodal data is performed, and the Dynamic Time Warping (DTW) algorithm is used to ensure that data from different sources are synchronized on the time axis. Principal component analysis is used for feature dimensionality reduction and standardization, and finally a standardized behavioral feature sequence with a unified timestamp and standardized format is generated. Among them, the anomaly detection algorithm based on isolated forests detects anomalies by isolating anomalous samples. Its design logic differs fundamentally from traditional anomaly detection methods that rely on sample density and distance. Its core principle is that the key characteristics of anomalous samples are their small number and significant difference from normal samples. Therefore, when randomly partitioning the data space, anomalous samples are separated more quickly than normal samples (i.e., fewer partitions are required). The algorithm randomly selects a feature and a threshold for that feature, continuously bisecting the data space to construct multiple "isolation trees." Normal data, due to its dense distribution, requires more partitions to be assigned to a single leaf node, resulting in longer paths. Anomalous data (noise), due to its sparse distribution, only requires a few partitions to be isolated, resulting in shorter paths. The average path length of the multiple isolation trees is taken. If the average path length of a certain data point is much shorter than that of most data points, it is judged as an anomaly (noise). Dynamic time warping is also employed. It is a classic algorithm used to measure the similarity of two time series data of different lengths and with asynchronous time axes. It is used to solve the problem of time axis misalignment caused by differences in the acquisition speed and rhythm of two essentially similar sequences. Time series alignment of multimodal data refers to mapping time series data from multiple modalities such as text, images, audio, and sensor signals to a unified time reference to ensure that multimodal information within the same time point or time interval truly reflects the same event or state. Principal component analysis is a commonly used data dimensionality reduction and feature extraction method. It transforms high-dimensional data into low-dimensional data while preserving as much of the key information (variance) of the original data as possible, thus simplifying the complexity of analysis. Standardization refers to mapping data of different magnitudes and units to a unified range, such as [0,1] or a normal distribution with a mean of 0 and a standard deviation of 1, through mathematical transformation to eliminate the interference of magnitude differences on the analysis.

[0055] Step S102: Based on the standardized behavioral feature sequence, the implicit reward function is deduced from the predefined behavioral demonstrations of highly cognitively resilient learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequence and the reward function, the cognitive resilience index of the target learner is calculated.

[0056] Among them, Inverse Reinforcement Learning (IRL) is a machine learning algorithm that infers the underlying objective function (i.e., reward function) from the behavioral trajectory of an agent. Its core logic is the opposite of traditional Reinforcement Learning (RL). The logic of RL is "given the reward, find the optimal behavior," that is, first define "what behavior should be rewarded and what should be punished" (reward function), and then let the agent learn through trial and error how to maximize the cumulative reward, and finally output the optimal policy. The logic of IRL is "given the behavior, find the reward function," that is, observe the actual behavioral data of the agent, such as human driving trajectory and robot operation records, and infer the implicit objective (reward function) driving the behavior, which is equivalent to answering why it does this. The predefined high cognitive resilience learner refers to a learner who has been rated as having high cognitive resilience by education experts based on their long-term performance.

[0057] For example, behavioral trajectories of highly resilient learners, as evaluated by experts, are retrieved from historical databases to form an expert demonstration dataset. These behavioral trajectories are then converted into state-action sequences, where the state includes task features and the learner's real-time state. By iteratively optimizing the reward function parameters to ensure that the behavioral distribution generated based on this function aligns with the expert distribution, the reward function is derived. Subsequently, this reward function is used to calculate the current learner's cognitive resilience index.

[0058] Step S103: Obtain the knowledge mastery level of the target learner, and construct a two-dimensional learner profile based on the cognitive resilience index and knowledge mastery level.

[0059] For example, based on learners' historical interaction sequences, a Bayesian knowledge tracing model can be used to dynamically estimate and update the mastery probability of each knowledge point, with the output being a probability vector regarding the degree of knowledge mastery. Simultaneously, the obtained cognitive resilience index is used to represent the learner's psychological trait level in non-cognitive domains. A graph neural network (GNN) can be used to construct a heterogeneous information network to fuse these two heterogeneous dimensions. The node set includes both knowledge point entities and various sub-dimensions of cognitive resilience, while the edges represent the pre-learning relationships between knowledge points and the influence weight of cognitive resilience on the learning process of specific knowledge points. By utilizing the message passing and node embedding update mechanisms of the graph neural network, the complex interaction between knowledge mastery and psychological traits is dynamically represented, thus forming a comprehensive learner model that reflects both "what is known" and "how to cope with the unknown"—a two-dimensional learner profile. Among them, Bayesian Knowledge Tracing (BKT) is a commonly used probabilistic model in the field of educational data mining. It dynamically infers a student's mastery of a certain knowledge point (such as "solving a quadratic equation") by analyzing the student's answer behavior (correct / incorrect). Essentially, it uses Bayesian probability to quantify the uncertainty of whether a student has learned the knowledge. Graph Neural Networks are a type of deep learning model specifically designed to process graph-structured data. The core of this model is to enable the model to understand the relationships between nodes and edges in the graph and learn features from these relationships to complete prediction or analysis tasks. Message passing refers to enabling nodes to exchange information. The core of a graph is "nodes + edges." The purpose of message passing is to allow each node (central node) to obtain useful information from its neighboring nodes (directly connected nodes), similar to "individuals obtaining information from others through social relationships" in reality. Node embedding update mechanism refers to enabling nodes to update their own representations. Node embedding is the vectorized representation of a node (such as [0.2, 0.5, 0.1]). The core of the update mechanism is to "use its own old features + neighbor aggregated information to generate new, more globally relevant embeddings," avoiding nodes relying solely on their own isolated features.

[0060] Step S104: Based on the two-dimensional learner profile and reward function, an initial personalized learning path is generated through a multi-objective optimization algorithm; wherein, the initial personalized learning path includes at least one strategic dilemma task.

[0061] Among them, multi-objective optimization algorithms are a class of algorithms used to solve optimization problems with multiple conflicting objectives. The core is to find a set of solutions that balance the objectives and have no absolute disadvantages (i.e., Pareto optimal solutions) in scenarios where it is impossible to make all objectives optimal at the same time, rather than the unique optimal solution in single-objective optimization.

[0062] For example, a multi-objective path optimization function is constructed, which is typically a linearly weighted combination of three sub-objectives: knowledge acquisition gains, cognitive resilience development gains, and cognitive load costs. Based on this multi-objective path optimization function, a Monte Carlo tree search algorithm can be used to explore a vast space of possible paths, evaluating their long-term benefits by simulating the execution of numerous random paths. Simultaneously, a deep Q-network is used as a value function approximator to quickly evaluate the node states in the tree search, guiding the search direction and improving search efficiency. Ultimately, the algorithm outputs an optimal path that balances multiple objectives, incorporating a strategic dilemma task tailored to the learner's current level. The path space is a set of paths of a certain type, but its definition and meaning need to be clarified in conjunction with the specific scenario. For example, in topological space, a path usually refers to "a continuous mapping from the closed interval [0,1] to the target topological space X", denoted as γ:[0,1]→X. The path space is the set of all such paths that satisfy specific conditions, denoted as P(X) or Ω(X). The latter often specifically refers to the loop space, that is, the set of paths whose starting point and ending point coincide, such as γ(0)=γ(1)=x0, where x0 is a fixed point in X. By simple analogy: if X is a plane Then the path space P( A continuous curve is the sum of all curves on a plane that start from any point and end at any point, such as a straight line, a parabola, or a part of a circle, as long as it is continuous. Monte Carlo Tree Search (MCTS) is a heuristic search algorithm based on random simulation and tree structure exploration. It is used to solve decision problems with large and exhaustive state spaces, such as Go and path planning. It gradually approaches the optimal decision through finite random attempts without relying on a predefined evaluation function, which is different from traditional game tree algorithms. Deep Q-Network (DQN) is a reinforcement learning algorithm that combines deep learning with Q (Quality, action value) learning. It is used to solve the problem that traditional Q-learning cannot directly look up tables when the state space is too large, allowing the agent to learn the optimal decision strategy autonomously in a high-dimensional, continuous environment. The value function approximator is the core tool to solve the problem that "the traditional table-based method fails due to the large state / action space". Its core logic is to use a parameterized function to replace the exhaustive table to estimate the value of a state or state-action pair.

[0063] Step S105: During the process of the target learner executing the initial personalized learning path, monitor in real time the data on the dilemma response generated by the target learner when interacting with the strategic dilemma task.

[0064] For example, during the initial personalized learning path execution by the target learner, their interaction with strategic dilemma tasks is monitored frequently in real time to capture dilemma response data. This data includes mouse movement trajectories, keyboard input event streams, and continuously collected learner facial video streams with explicit authorization. All these multimodal data streams undergo rigorous time synchronization and feature-level fusion to form a multidimensional dilemma response data vector. This vector is used to objectively characterize the learner's psychological and behavioral response patterns when facing challenges in real time. Time synchronization refers to using a unified clock source (e.g., assigning the same time reference to all devices) and calibrating for delay biases. For example, if the camera is measured to be 20ms slower than the microphone, the image data is "delayed by 20ms" before matching to ensure that multimodal data at the same point in time correspond to the same event state. Feature-level fusion refers to first extracting key features from the raw data of each modality—for example, extracting speech features such as tone and voiceprint vectors corresponding to keywords from the user's audio—and then integrating these features, rather than directly merging the raw data.

[0065] Step S106: Based on the distress response data, update the two-dimensional learner profile to obtain the updated two-dimensional learner profile. Based on the updated two-dimensional learner profile, adjust the difficulty level and support level of subsequent tasks in the initial personalized learning path to obtain the adjusted learning path.

[0066] For example, incremental updates are performed on the two-dimensional learner profile: The cognitive resilience index is updated by using adversity response data as new features, inputting them into the corresponding quantitative model. For instance, an online learning algorithm is used to fine-tune the weights of the reward function, reflecting the learner's latest resilience level changes. The knowledge point mastery update is based on the learner's final response in the strategic dilemma task, applying Bayesian knowledge tracking update rules to adjust the mastery probability of the corresponding knowledge points. Then, based on the updated learner profile, the difficulty level and support level of subsequent learning tasks not yet executed are adjusted in real time. Difficulty adjustment typically uses the updated cognitive resilience index and knowledge point mastery as input variables, inferring the specific amount of difficulty adjustment through a pre-defined fuzzy rule base. Support level adjustment is reflected in dynamically adding appropriate scaffolding support to the task, such as providing multi-step hints, displaying solution examples, or adjusting the level of detail in feedback. Finally, an adjusted learning path is generated to ensure the learning experience remains within a personalized optimal challenge range. Incremental updates refer to updating only the changed parts, rather than replacing the entire model, aiming to improve efficiency and reduce resource consumption. Online learning algorithms are learning-as-you-go algorithms that do not require collecting all data before training. Instead, they can receive new data in real time and dynamically update model parameters, unlike offline learning which requires training before use and cannot be adjusted in real time. Bayesian knowledge tracking dynamically judges a learner's mastery of a knowledge point through probability updates, usually simplified to two implicit states: mastered or not mastered. Its update rules revolve around the logic of "prior probability → likelihood probability → posterior probability". Essentially, it uses new evidence (answer results) to continuously revise the belief in the student's knowledge state: after each answer, the "posterior probability of the previous round" is transformed into the "prior probability of the next round", and this process is iterated to ultimately achieve dynamic tracking of the student's mastery state. The pre-set fuzzy rule base refers to the set of fuzzy logic rules in the form of "condition-conclusion" that transforms human experience judgments in specific scenarios into a set of rules. For example, "IF (if) cognitive resilience is high and the mastery of knowledge points is moderate, THEN (then) appropriately increase the difficulty."

[0067] In this embodiment, multimodal behavioral data of the target learners is collected and features are extracted. Then, inverse reinforcement learning technology is used to deduce the reward function from expert behavior to quantify cognitive resilience, constructing a dual-dimensional learner profile that integrates knowledge mastery and cognitive resilience. A multi-objective optimization algorithm is used to generate an initial learning path embedded with strategic dilemmas. During path execution, the learner profile is dynamically updated by monitoring dilemma response data in real time. The difficulty and support of subsequent paths are adjusted based on the updated learner profile, forming an adjusted learning path. This approach breaks through the limitations of traditional personalized learning systems that focus solely on knowledge transfer efficiency. It transforms the concepts of "beneficial dilemmas" and "cognitive resilience" from educational psychology into calculable and operable engineering practices, significantly improving the depth of knowledge understanding and long-term transferability. It also consciously and systematically cultivates learners' higher-order psychological traits such as analysis, perseverance, and adaptation when facing complex and ambiguous problems, achieving a fundamental shift from pursuing short-term "optimal efficiency" to promoting long-term "optimal growth."

[0068] In one embodiment, such as Figure 2 As shown, based on standardized behavioral feature sequences, an implicit reward function is derived from the behavioral demonstrations of predefined high cognitive resilience learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequences and the reward function, the cognitive resilience index of the target learner is calculated, including:

[0069] Step S201: Obtain the behavioral trajectories of learners assessed by experts as having high cognitive resilience from the preset historical database, and summarize all behavioral trajectories to generate an expert demonstration dataset.

[0070] The pre-defined historical database refers to a learner behavior data storage system that is pre-built and maintained over a long period of time. It includes a large number of learners' multimodal learning behavior records, corresponding expert evaluation tags, and related metadata. The predefined high cognitive resilience learner evaluation criteria is a quantitative evaluation system developed by education experts based on cognitive psychology theory and teaching practice, from multiple dimensions such as learning persistence, effectiveness of adversity coping strategies, emotion regulation ability, and problem-solving innovation. The criteria have been pre-fixed in the tag rules of the database.

[0071] For example, all learner behavior trajectories labeled "high cognitive resilience" by expert assessments are selected from a pre-defined historical database. These trajectories encompass multimodal information such as cognitive interaction data sequences, micro-operation records, and emotional and physiological response time-series data of these learners in various learning tasks. This data has already undergone preliminary cleaning and standardization upon entry into the database. The retrieved trajectories are then deduplicated and filtered for outliers, such as removing incomplete trajectories caused by network latency. After time alignment, they are concatenated to form a standardized set containing multiple expert behavior sequences.

[0072] Step S202: Based on the expert demonstration dataset, the reward function is derived by using the maximum entropy inverse reinforcement learning algorithm; whereby the reward function is used to explain the behavioral preferences of highly resilient learners.

[0073] Among them, the maximum entropy inverse reinforcement learning algorithm is an inverse reinforcement learning algorithm based on the principle of maximum entropy. While explaining expert behavior, it maximizes the entropy value of the behavior trajectory and improves the generalization ability of the reward function. The technical principle of this algorithm is to overcome the limitation of traditional inverse reinforcement learning algorithms that are prone to getting trapped in local optima. By introducing maximum entropy constraints, the inverse reward function can not only explain the observed expert behavior, but also remain open to reasonable behaviors that have not been observed, thus avoiding overfitting to a single behavior pattern.

[0074] For example, using a constructed expert demonstration dataset as input, a maximum entropy inverse reinforcement learning algorithm is employed to inversely deduce the reward function. First, the behavioral trajectories in the expert demonstration dataset are converted into state-action sequences recognizable by the algorithm. The states contain information such as learning task features and the learner's real-time state, while the actions correspond to the learner's specific learning behavior choices. Through iterative optimization, the weight parameters of the reward function are continuously adjusted so that the distribution of behavioral trajectories generated by the optimal strategy based on this reward function tends to be consistent with the distribution of behavioral trajectories in the expert demonstration dataset. The resulting reward function can quantitatively characterize the behavioral preferences of learners with high cognitive resilience, clarifying which learning behaviors, strategy choices, and emotional states are more consistent with high cognitive resilience traits.

[0075] Step S203: Based on the reward function and the standardized behavioral feature sequence, the cognitive resilience index of the target learner is calculated.

[0076] For example, the reward function indirectly defines the core quantitative dimensions of cognitive resilience—challenge persistence, strategy transferability, and emotional resilience—by quantifying the behavioral value preferences of learners with high cognitive resilience. The standardized behavioral feature sequences of the target learners are matched with the corresponding behavioral representation indicators of the three quantitative dimensions. Raw scores for each quantitative dimension are obtained through methods such as behavioral pattern similarity comparison and feature contribution calculation. Subsequently, the raw scores are standardized by converting the raw data of different dimensions into standard normal distribution data with a mean of 0 and a standard deviation of 1 using the overall mean and standard deviation of the data. The standardized scores are then weighted and summed to obtain a cognitive resilience index that comprehensively reflects the cognitive resilience level of the target learners. Among them, behavioral pattern similarity comparison is a technical method used to quantify the degree of matching between the behavioral characteristics of a target object and a reference object (such as a learner with high cognitive resilience). It involves converting two types of behavioral data into structured feature vectors, such as the target learner's vectors for "duration on challenge tasks" and "frequency of strategy adjustments," and comparing them with the corresponding feature vectors of learners with high cognitive resilience. Then, a pre-set similarity measurement algorithm (such as cosine similarity) is used to calculate the degree of matching between the vectors. For example, in cognitive resilience assessment, behavioral patterns such as the target learner's "sequence of changes in hesitation index" and "timeline of recovery from negative emotions" when facing difficult tasks are compared with typical behavioral pattern vectors of learners with high cognitive resilience in expert demonstration datasets. The higher the similarity value, the better the target learner's performance in that dimension. The stronger the fit with high cognitive resilience traits, the better. Feature contribution calculation is used to measure the influence weight of a single behavioral feature on the overall evaluation index (such as the cognitive resilience index). Through statistical or algorithmic analysis, the importance of different behavioral features in representing target attributes (such as cognitive resilience) is identified. For example, in the calculation of the cognitive resilience index, it is necessary to analyze the contribution ratio of each sub-feature under dimensions such as challenge persistence, strategy transfer ability, and emotional recovery ability, such as "task persistence duration", "cross-task strategy reuse rate" and "negative emotion fading speed" to the final index. Correlation analysis (such as Pearson correlation coefficient) can be used to determine the correlation strength between features and the index, ensuring that the cognitive resilience index obtained by subsequent weighted summation can accurately reflect the role of core features and avoid secondary features interfering with the evaluation results.

[0077] In this embodiment, by filtering and integrating the behavioral trajectories of highly cognitively resilient learners evaluated by experts from a pre-set historical database, a widely representative expert demonstration dataset is generated. The maximum entropy inverse reinforcement learning algorithm is used to deduce a reward function that accurately explains the behavioral preferences of highly cognitively resilient learners. Based on this reward function, the quantitative dimensions of cognitive resilience are defined, and the cognitive resilience index of the target learner is calculated. This effectively solves the technical problem that traditional methods cannot quantify cognitive resilience, an implicit psychological trait. The calculated cognitive resilience index can comprehensively and accurately represent the learner's core psychological traits.

[0078] In one embodiment, the cognitive resilience index of the target learner is calculated based on the reward function and the standardized behavioral feature sequence, including:

[0079] Step S301: Based on the reward function, determine the quantitative dimensions of cognitive resilience; among which, the quantitative dimensions include challenge persistence, strategy transferability, and emotional resilience.

[0080] For example, the reward function quantifies the behavioral value preferences of learners with high cognitive resilience, clearly defining the core components supporting cognitive resilience. Based on the function's selection and weighting of high-value behavioral characteristics, three key quantitative dimensions are extracted: challenge persistence, strategy transferability, and emotional resilience. Challenge persistence refers to the learner's willingness and duration of focus and commitment when facing difficult tasks beyond their current capabilities; strategy transferability refers to the learner's ability to flexibly apply learned problem-solving methods and thinking patterns to new, similar, or related learning scenarios; and emotional resilience refers to the learner's ability to quickly adjust their mental state and re-engage in learning after experiencing learning setbacks and negative emotions. For example, these three quantitative dimensions are determined through logical deduction and feature matching, combining the core definition of cognitive resilience in educational psychology with the behavioral preferences reflected by the reward function, comprehensively covering the core connotations of cognitive resilience. Among them, cognitive resilience is defined as the psychological trait and ability of an individual to proactively maintain cognitive engagement, flexibly adjust cognitive strategies, and continuously pursue cognitive goals when facing difficulties and challenges such as complex learning tasks, knowledge comprehension obstacles, and thinking bottlenecks in cognitive activities. It includes three key dimensions: cognitive persistence, strategic flexibility, and goal orientation. Logical deduction refers to the process of drawing conclusions through step-by-step reasoning based on clear rules, axioms, or causal relationships. Feature matching is the process of judging whether two things are similar or related by comparing them with known samples or templates based on the key attributes (features) of things.

[0081] Step S302: Calculate the raw scores for each quantitative dimension based on the standardized behavioral feature sequence.

[0082] For example, based on standardized behavioral feature sequences, a set of feature-to-score mapping rules or calculation models are preset for each dimension: For challenge persistence, its raw score can be obtained by calculating the learner's total duration and number of effective attempts (excluding invalid very short attempts) in a single challenging task, and then weighted and summed in combination with the basic difficulty coefficient of the task; For strategy transfer ability, different problem-solving strategies are identified from the behavioral feature sequence, for example, by matching using predefined strategy templates, and then the total number of strategies used, the differences between strategies (such as the cosine distance based on the strategy feature vector), and the frequency of strategy switching are calculated, and finally these indicators are combined into a strategy effectiveness score; For emotional resilience, its raw score calculation relies on the temporal analysis of emotion-related behavioral features. By identifying the time point when an error occurs or negative feedback is received, the changes in behavioral features over a period of time afterward are analyzed, such as calculating whether the response time of subsequent attempts is shortened or the accuracy rate is improved, and combining the decay rate of emotional physiological data (such as the negative emotional intensity of facial expressions), a regression model is used to estimate the speed of emotional recovery. Among them, the predefined strategy template refers to defining key behavioral characteristics for each known strategy in advance to form a template. For example, the template characteristics of the "trial and error method" can be set as: ① no explicit formula / logical derivation, directly substitute numerical values; ② modify the substituted values ​​multiple times (≥2 times); ③ finally determine the answer by "eliminating incorrect values"; the regression model is a core tool for analyzing the relationship between variables and predicting continuous results. It finds the mathematical law between the independent variable (influencing factors) and the dependent variable (result to be predicted / analyzed) through known data, and then uses this law to explain the existing data or predict the results of new data.

[0083] Step S303: The original scores are weighted and summed using the following formula to obtain the cognitive resilience index of the target learner:

[0084]

[0085] in, For cognitive resilience index, To challenge the raw score of endurance, The raw score for strategy transfer ability. The raw score for emotional resilience. To challenge the overall mean of durability, This represents the overall mean of strategy transfer capability. This represents the overall mean of emotional resilience. To challenge the overall standard deviation of persistence, The overall standard deviation of strategy transfer capability. The overall standard deviation of emotional resilience. To challenge the weighting coefficients of persistence, The weighting coefficients for policy transfer capability. This is the weighting coefficient for emotional resilience.

[0086] The weighting coefficients α, β, and γ are obtained by analyzing expert demonstration datasets of learners with high cognitive resilience: on this dataset, the covariance matrix of the standardized scores of each dimension is calculated or regression analysis is performed to determine the contribution of each dimension to the final result of being identified as "highly resilient" by experts, thereby deriving the optimal weight combination and ensuring that the sum of the weights is 1.

[0087] For example, since the raw scores for different dimensions may have different dimensions and distribution ranges, direct summation can lead to distortion of the weights for some dimensions. Therefore, it is necessary to standardize each raw score, that is, to obtain the population mean (μ) and population standard deviation (σ) of the raw scores for each dimension from historical data, and then to standardize the current learner's raw score (μ). , , The standardized score is obtained by subtracting the mean of the corresponding dimension and then dividing by the corresponding standard deviation. This transforms the scores of different dimensions into comparable values ​​with a mean of 0 and a standard deviation of 1. Subsequently, the standardized scores are weighted and summed according to pre-set weighting coefficients to obtain the cognitive resilience index.

[0088] In this embodiment, three core quantitative dimensions of cognitive resilience are scientifically defined based on a reward function. Standardized behavioral feature sequences are used to calculate the raw scores for each dimension, ensuring a strong correlation between the scores and the corresponding traits. Standardization eliminates dimensional differences, and a weighted sum is performed using appropriate weights to obtain a comprehensive cognitive resilience index. This effectively addresses the technical bottleneck of traditional methods in quantifying cognitive resilience, an implicit psychological trait.

[0089] In one embodiment, an initial personalized learning path is generated based on a two-dimensional learner profile and a reward function using a multi-objective optimization algorithm; wherein the initial personalized learning path includes at least one strategic dilemma task, including:

[0090] Step S401: Construct a multi-objective path optimization function, wherein the multi-objective path optimization function includes knowledge acquisition benefits, cognitive resilience cultivation benefits, and cognitive load costs.

[0091] For example, a multi-objective path optimization function is constructed as the core decision mechanism for path generation. This function uses a linear weighted sum method to integrate multiple optimization objectives into a quantifiable scalar function, the mathematical expression of which is:

[0092]

[0093] in, It represents the benefit of knowledge acquisition, which is quantified by calculating the sum of the increased probability of mastering all knowledge points in the path; The benefit of cognitive resilience training is calculated based on the matching degree between the number and difficulty of strategic dilemma tasks in the path and the learner's current level of cognitive resilience. Representing cognitive load cost, cognitive burden is assessed through indicators such as path length, task switching frequency, and the magnitude of difficulty variation. Weighting coefficients. , and The Analytic Hierarchy Process (AHP) combined with scores from educational experts can be used to determine the relative importance of each objective, ensuring that the relative importance of each objective is reasonably reflected. AHP is a multi-criteria decision analysis method that breaks down complex decision-making problems such as optimal solution selection and determination of indicator weights into a clear hierarchical structure. By combining subjective judgment with objective calculation, it arrives at a scientific decision result.

[0094] Step S402: Determine the baseline value of knowledge acquisition benefits based on the knowledge mastery level in the dual-dimensional learner profile.

[0095] For example, the mastery probability vector of each knowledge point is extracted from the two-dimensional learner profile. ,in ∈[0,1] represents the mastery level of the i-th knowledge point. This is based on the knowledge point importance weight vector defined in the domain knowledge graph. ,in This represents the importance weight of the i-th knowledge point, used to calculate the basic value of the current knowledge state. The baseline value for knowledge acquisition gains can be set from the current state to the state of complete mastery (all...). The theoretical maximum return value of (=1) The baseline value is determined based on the knowledge space theory in educational measurement, which models the learning process as a transfer of knowledge states. This baseline value reflects the gap between the learner's current knowledge level and the ideal state. Knowledge Space Theory (KST) is a crucial theoretical framework in educational measurement used to quantitatively analyze learners' knowledge states and optimize assessment and teaching decisions. Its core is to transform knowledge into a computable and mappable spatial structure, thereby accurately locating the learner's knowledge mastery. Domain knowledge graphs are structured collections of knowledge focused on specific industries or professional fields. Their core is to organize key information such as concepts, entities, attributes, and relationships within that field into a "entity-relationship-entity" graph format, forming a machine-understandable knowledge network, distinct from general knowledge graphs covering multiple fields (such as Baidu Knowledge Graph).

[0096] Step S403: Based on the cognitive resilience index in the two-dimensional learner profile, determine the difficulty range of the strategic dilemma task.

[0097] For example, the cognitive resilience index in the two-dimensional learner profile is converted into a difficulty level through a linear mapping function. The mapping function is calibrated based on statistics of the difficulty learners with different resilience levels successfully coped with difficulties in historical data. Then, using... Expand upwards from the center Determine the upper limit of difficulty and expand downwards. Determine the lower limit of difficulty and form a range. Expansion range and This can be achieved using Vygotsky's Zone of Proximal Development (ZPD) theory, ensuring that the difficulty of the challenge is both higher than the learner's current ability to solve problems independently, and within their potential range with appropriate support. Vygotsky's ZPD theory is a core concept in his sociocultural theory, revealing the potential space for individual learning and development, rather than focusing solely on existing abilities. Simply put, it refers to the gap between two levels: first, the actual developmental level, which is the individual's current level of ability to complete tasks and solve problems independently without assistance, such as a child being able to independently calculate addition and subtraction within 10; second, the potential developmental level, which is the level of task an individual cannot yet complete independently, but can complete with guidance, demonstration, or collaboration from others (such as teachers), such as a child initially unable to calculate subtraction within 20 with carrying and borrowing, but gradually able to do so after parental prompting using the "making ten" method. This gap represents the key space for individual cognitive growth.

[0098] Step S404: The Monte Carlo tree search algorithm is used to sample in the path space to obtain candidate paths.

[0099] The path space is defined as the set of all possible sequences of learning activities, where each node represents a learning task and the edges represent the transition relationships between tasks.

[0100] For example, the Monte Carlo Tree Search (MCTS) algorithm is used to sample the path space. The MCTS algorithm executes in a loop through four phases: Selection, Expansion, Simulation, and Backpropagation. In the Selection phase, starting from the root node, child nodes are selected based on the Upper Confidence Bound (UCB) formula. In the Expansion phase, new branches are expanded when incompletely explored nodes are encountered. In the Simulation phase, the path value is quickly evaluated using a randomized strategy. In the Backpropagation phase, the simulation results are backpropagated to update node statistics. After multiple iterations, the algorithm outputs a set of candidate paths with high potential. The UCB formula is a core algorithmic tool in the multi-armed slot machine problem used to balance exploration and exploitation. Its core logic is: for each option (such as the recommended strategy), not only is its historical average reward (exploitation) considered, but an uncertainty penalty term (exploration) is also added. Finally, the option with the highest "average reward + upper uncertainty limit" is selected to avoid missing potential optimal solutions due to insufficient information. The core formula is: ; The upper confidence interval value of the i-th option is the basis for the final decision; the higher the value, the higher the priority of selection. is the historical average return of the i-th option, the core of "utilization", used to reflect the known good or bad of this option; t is the current total number of trials, that is, the total number of times all options have been selected, which increases as the trials progress; The key to "exploration" is the historical number of times the i-th option has been selected. The smaller the value, the higher the uncertainty of the option, and the larger the penalty. As a penalty for uncertainty, the quantification of "exploration" is: the more total trials t, the more likely this option is to be selected. The less information a given item has, the larger this item becomes, forcing the algorithm to prioritize options with less information. Random strategies are strategies that randomly select action plans rather than relying on fixed logic, preferences, or predictions during decision-making or game theory. The core idea is to use randomness to break determinism in order to cope with scenarios with insufficient information, opponent predictions, or complex uncertainties.

[0101] Step S405: Based on the baseline value of knowledge acquisition benefits and the difficulty range, calculate the three-dimensional indicators for each candidate path; wherein, the three-dimensional indicators include knowledge acquisition benefits, cognitive resilience cultivation benefits, and cognitive load costs.

[0102] For example, for each candidate path, the knowledge acquisition benefit is obtained by weighted summing of the expected improvement in knowledge point mastery for each task in the path, and then normalizing by dividing by the baseline value; the cognitive resilience development benefit is calculated by statistically analyzing the number of challenging tasks falling into the difficulty range in the path, and weighting the matching degree between their difficulty coefficients and the learner's resilience level; the cognitive load cost is estimated by linear combination of features such as path length and task complexity switching frequency. During the calculation, a prediction model based on Item Response Theory (IRT) can be used to predict the learner's probability of success on new tasks based on their historical performance, ensuring the accuracy of the indicator calculation. Specifically, the prediction model based on Item Response Theory (IRT) infers the level of potential traits such as ability, attitude, and knowledge acquisition from the subject's responses to specific items (such as exam questions), and further predicts their performance in similar items or scenarios based on this trait level.

[0103] Step S406: Substitute the three-dimensional indicators of each candidate path into the multi-objective path optimization function to generate a value score for each candidate path; and determine the candidate path corresponding to the maximum value score as the initial personalized learning path.

[0104] For example, the three-dimensional index vector of each candidate path is standardized to eliminate dimensional differences, and then input into the multi-objective path optimization function to calculate the comprehensive score. The top k paths with the highest value scores are selected to form a candidate set. Finally, the path with the highest score and best stability is selected as the initial personalized learning path according to the variance reduction criterion. The variance reduction criterion is a core optimization idea in statistics and numerical computation used to reduce the variance of estimates and improve the accuracy of results. The core logic is to reduce the fluctuation of random errors by designing more efficient calculation methods or data utilization methods without changing the "unbiasedness" of the estimate or controlling the bias within an acceptable range, so that the estimation results are more stable and closer to the true value.

[0105] In this embodiment, a multi-objective path optimization function encompassing knowledge, resilience, and workload is constructed. Based on a two-dimensional learner profile, a baseline value for knowledge acquisition gains and a difficulty range for strategic dilemma tasks are determined. A Monte Carlo tree search algorithm is used to efficiently sample candidate paths from a vast path space. The three-dimensional indicators of each candidate path are calculated and substituted into the optimization function to obtain a value score, ultimately selecting the optimal initial personalized learning path. This effectively solves the technical problems of traditional personalized learning paths that focus only on knowledge transfer, neglect the cultivation of cognitive resilience, and have low path generation efficiency. The generated initial path not only matches the learner's current knowledge and psychological characteristics but also promotes their all-round development.

[0106] In one embodiment, during the initial personalized learning path executed by the target learner, real-time monitoring of the learner's dilemma response data when interacting with strategic dilemma tasks includes:

[0107] Step S501: Obtain the mouse movement trajectory coordinate sequence of the target learner in the strategic dilemma task interface, and calculate the hesitation index of the mouse movement trajectory coordinate sequence.

[0108] For example, a sequence of mouse movement coordinates within the interface of a strategic dilemma task is obtained. This sequence is a continuous set of coordinate points including timestamps, recording the mouse's movement path on the screen. A sliding window smoothing algorithm can be used to filter the original trajectory coordinate sequence, removing noise points caused by device jitter or unintentional operation, and interpolation methods can be used to compensate for missing data points due to sampling frequency fluctuations. Subsequently, multiple kinematic features are extracted from the cleaned trajectory data, such as the coefficient of variation of movement speed, the number of abrupt changes in movement direction, the hovering duration on specific interface elements (such as submit buttons or tooltip icons), and the tortuosity of the trajectory (i.e., the ratio of the actual path length to the straight-line distance between the start and end points). Based on these features, a hesitation index is calculated using a pre-trained regression model. Among them, the pre-trained regression model refers to the model trained on a large amount of labeled data, such as the degree of hesitation assessed by experts based on behavioral videos. It can map multi-dimensional motion features to a scalar value between 0 and 1. The higher the value, the stronger the uncertainty and hesitation of the learner in the decision-making process. The sliding window smoothing algorithm is a basic data processing method for denoising and smoothing fluctuations in time series data. By setting a fixed-size window, the window moves gradually from the beginning to the end of the time series. For all data points in each window, a local statistic is calculated and used to replace the original data value at the center of the window or a certain position in the window, thereby weakening the impact of random noise or short-term abnormal fluctuations and preserving the overall trend of the data. The interpolation method refers to the estimation of function values ​​at any unknown position between several known discrete data points (called interpolation nodes) by constructing a continuous function (called the interpolation function). It requires that the interpolation function must pass through all known nodes, that is, satisfy the interpolation condition: at each node, the interpolation function value is equal to the known data value.

[0109] Step S502: Obtain the keyboard input event stream of the target learner during the task process, and analyze the keyboard input event stream to obtain the input rate change pattern and modification frequency.

[0110] For example, the keyboard input event stream of the target learner during the task is acquired. This event stream includes the key value of each keystroke, the timestamps of pressing and releasing, and the event type, such as keydown. The raw event stream is analyzed to organize discrete keystroke events into semantically meaningful input units, such as a sequence of inputs of a complete word or a mathematical formula. Then, temporal pattern analysis is performed to calculate key indicators characterizing input fluency, including the average keystroke interval (IKI), the standard deviation of the keystroke interval reflecting the stability of the input rhythm, and the frequency and distribution patterns of specific correction behaviors (such as the delete key, Backspace). For example, identifying long pauses followed by rapid input segments, or frequent "input-delete-re-input" loop patterns. The input rate variation pattern is quantified into a time-series feature vector, while the correction frequency is calculated as the number of correction operations per unit time. Temporal pattern analysis is the core analytical method for time-series data, aiming to extract patterns, trends, or anomalies from the temporal correlation of the data to provide a basis for prediction and decision-making.

[0111] Step S503: With authorization, acquire the facial video stream of the target learner and extract the micro-expression change sequence from the facial video stream using a pre-trained facial expression recognition model.

[0112] For example, with the learner's explicit informed consent and authorization, facial video streams of the target learner are collected as input data. The collection process must ensure stable lighting conditions and clear visibility of the facial area. A pre-trained Facial Expression Recognition (FER) model is then used to process the video stream. By performing face detection and alignment on each frame, key facial points, such as the corners of the eyes and mouth, are located. The aligned facial areas are then input into a CNN, outputting probability distribution vectors corresponding to basic emotions such as happiness, sadness, and anger. Particular attention is paid to micro-expressions related to frustration and confusion—that is, brief, involuntary facial expressions—capturing these instantaneous reactions by analyzing rapid and subtle changes in expression probabilities between consecutive frames. The final output is a sequence of micro-expression changes, which is a data structure corresponding to timestamps and expression category probability distributions. Pre-trained facial expression recognition models are typically based on a deep convolutional neural network (CNN) architecture and are initially trained on a large-scale general facial expression dataset (such as FER2013). After acquiring basic expression recognition capabilities, they can be directly used for inference or fine-tuned to adapt to specific scenarios. FER2013 (Facial Expression Recognition 2013) is a classic public dataset for facial expression recognition, mainly used for training and evaluating algorithms related to facial expression classification in the field of computer vision. The architecture of deep convolutional neural networks is a deep learning architecture inspired by the structure of the biological visual cortex. Its core advantage lies in automatically extracting spatial features, such as image edges, textures, shapes, and even complex semantics, and significantly reducing the amount of computation through parameter sharing. Essentially, it "uses convolutional kernels for local feature extraction, pooling for feature compression, and layer stacking for feature abstraction," optimizing parameter utilization efficiency and feature representation capabilities.

[0113] Step S504: Calculate the percentage of duration of negative emotion intensity in the micro-expression change sequence.

[0114] For example, a set of facial expression categories associated with negative emotions is defined, typically including anger, sadness, fear, and disgust. For each frame in the micro-expression change sequence, its negative emotion intensity value is calculated. This can be obtained by weighted summation of the probabilities of each expression category within the set, with the weights set according to the strength of the association between different expressions and learning frustration. Next, a continuous emotional state determination is performed. When the negative emotion intensity of several consecutive frames exceeds a preset threshold, the learner is determined to be in a negative emotional state. The entire sequence is iterated, and the duration of all determined negative emotional states is accumulated, calculating its percentage of the total task interaction time, i.e., the percentage of negative emotion intensity duration. The preset threshold refers to a pre-defined critical standard or numerical limit. When a certain data point reaches, exceeds, or falls below this limit, a preset response is triggered, providing a clear triggering condition for decision-making or behavior.

[0115] Step S505: Summarize the hesitation index, input rate change pattern, and duration of negative emotion intensity to form dilemma response data.

[0116] For example, hesitation index, input rate change patterns (including their feature vectors and statistics), and the proportion of negative emotion intensity duration are summarized and fused. Since these data originate from different modalities and have different dimensions, standardization is first performed before feature-level fusion to unify the indicators to the same numerical range. Then, a feature-level fusion strategy can be used to concatenate the standardized multi-dimensional features into a high-dimensional feature vector. This vector comprehensively characterizes learners' behavioral hesitation, cognitive fluency, and emotional state in challenging tasks. The resulting challenging response data is a structured, multi-dimensional data object. The feature-level fusion strategy is one of the core technologies in multi-source information fusion. It first integrates and optimizes the basic features extracted from various data sources, and then conducts subsequent tasks based on the enhanced fused features, rather than directly fusing the original data or the final decision results.

[0117] In this embodiment, hesitation index is extracted and quantified from mouse trajectories, input patterns are analyzed from keyboard event streams to assess cognitive fluency, and facial expression recognition technology, under authorized conditions, is used to capture micro-expression changes and calculate the duration of negative emotions to quantify emotional responses. These three heterogeneous multimodal behavioral indicators are effectively integrated to form comprehensive distress response data. This enables non-invasive, objective, and high-granular measurement of learners' implicit psychological states such as frustration and perseverance, completely overcoming the shortcomings of traditional self-report methods, such as strong subjectivity and significant lag.

[0118] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0119] Based on the same inventive concept, this application also provides a personalized learning path generation device based on learning behavior profiles for implementing the personalized learning path generation method based on learning behavior profiles described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the personalized learning path generation device based on learning behavior profiles provided below can be found in the limitations of the personalized learning path generation method based on learning behavior profiles described above, and will not be repeated here.

[0120] In one exemplary embodiment, such as Figure 3 As shown, a personalized learning path generation device 300 based on learning behavior profiles is provided, including:

[0121] The data acquisition module 301 is used to acquire multimodal learning behavior data of the target learner and perform feature extraction processing on the multimodal learning behavior data to obtain a standardized behavioral feature sequence; wherein, the multimodal learning behavior data includes cognitive interaction data, behavioral micro data and emotional physiological data;

[0122] The index calculation module 302 is used to inversely deduce the implicit reward function from the behavioral demonstrations of predefined high cognitive resilience learners through an inverse reinforcement learning algorithm, and calculate the cognitive resilience index of the target learner based on the standardized behavioral feature sequence and the reward function.

[0123] The profile building module 303 is used to obtain the knowledge mastery of the target learner and to build a two-dimensional learner profile based on the cognitive resilience index and knowledge mastery.

[0124] The initial path generation module 304 is used to generate an initial personalized learning path based on a two-dimensional learner profile and a reward function, using a multi-objective optimization algorithm; wherein the initial personalized learning path includes at least one strategic dilemma task.

[0125] The data monitoring module 305 is used to monitor in real time the dilemma response data generated by the target learner when interacting with strategic dilemma tasks during the execution of the initial personalized learning path.

[0126] The path update module 306 is used to update the two-dimensional learner profile based on the dilemma response data, obtain the updated two-dimensional learner profile, and adjust the difficulty level and support level of subsequent tasks in the initial personalized learning path based on the updated two-dimensional learner profile, so as to obtain the adjusted learning path.

[0127] In one embodiment, the exponent calculation module 302 is further configured to:

[0128] The behavioral trajectories of learners assessed by experts as having high cognitive resilience are obtained from a pre-set historical database, and all behavioral trajectories are aggregated to generate an expert demonstration dataset.

[0129] Based on an expert demonstration dataset, a reward function is derived by using the maximum entropy inverse reinforcement learning algorithm; the reward function is used to explain the behavioral preferences of highly resilient learners.

[0130] The cognitive resilience index of the target learner is calculated based on the reward function and standardized behavioral feature sequence.

[0131] In one embodiment, the exponent calculation module 302 is further configured to:

[0132] Based on the reward function, quantitative dimensions of cognitive resilience are determined; these quantitative dimensions include challenge endurance, strategy transferability, and emotional resilience.

[0133] Based on the standardized behavioral feature sequence, calculate the raw scores for each quantitative dimension;

[0134] The cognitive resilience index of the target learner is obtained by weighted summation of the raw scores using the following formula:

[0135]

[0136] in, For cognitive resilience index, To challenge the raw score of endurance, The raw score for strategy transfer ability. The raw score for emotional resilience. To challenge the overall mean of durability, This represents the overall mean of strategy transfer capability. This represents the overall mean of emotional resilience. To challenge the overall standard deviation of persistence, The overall standard deviation of strategy transfer capability. The overall standard deviation of emotional resilience. To challenge the weighting coefficients of persistence, The weighting coefficients for policy transfer capability. This is the weighting coefficient for emotional resilience.

[0137] In one embodiment, the initial path generation module 304 is further configured to:

[0138] Construct a multi-objective path optimization function, which includes knowledge acquisition benefits, cognitive resilience development benefits, and cognitive load costs:

[0139] Based on the knowledge mastery level in the dual-dimensional learner profile, a baseline value for knowledge mastery benefits is determined.

[0140] Based on the cognitive resilience index in the two-dimensional learner profile, the difficulty range of strategic dilemma tasks is determined.

[0141] The Monte Carlo tree search algorithm is used to sample in the path space to obtain candidate paths;

[0142] Based on the baseline value of knowledge acquisition benefits and the difficulty range, three-dimensional indicators for each candidate path are calculated; among them, the three-dimensional indicators include knowledge acquisition benefits, cognitive resilience cultivation benefits, and cognitive load costs.

[0143] The three-dimensional metrics of each candidate path are substituted into the multi-objective path optimization function to generate a value score for each candidate path; and the candidate path corresponding to the maximum value score is determined as the initial personalized learning path.

[0144] In one embodiment, the data monitoring module 305 is further configured to:

[0145] Obtain the mouse movement trajectory coordinate sequence of the target learner in the strategic dilemma task interface, and calculate the hesitation index of the mouse movement trajectory coordinate sequence;

[0146] Acquire the keyboard input event stream of the target learner during the task process, and analyze the keyboard input event stream to obtain the input rate change pattern and modification frequency;

[0147] With authorization, facial video streams of the target learner are acquired, and micro-expression change sequences are extracted from the facial video streams using a pre-trained facial expression recognition model;

[0148] Calculate the percentage of duration of negative emotion intensity in a micro-expression change sequence;

[0149] The hesitation index, input rate change pattern, and duration of negative emotion intensity are combined to form the dilemma response data.

[0150] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the personalized learning path generation method based on learning behavior profiles as described above.

[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0152] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0153] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for generating personalized learning paths based on learning behavior profiles, characterized in that, The method includes: The multimodal learning behavior data of the target learner is acquired, and feature extraction processing is performed on the multimodal learning behavior data to obtain a standardized behavioral feature sequence; wherein, the multimodal learning behavior data includes cognitive interaction data, behavioral micro data, and emotional physiological data; The implicit reward function is derived from the behavioral demonstrations of predefined high cognitive resilience learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequence and the reward function, the cognitive resilience index of the target learner is calculated. The knowledge mastery level of the target learner is obtained, and a two-dimensional learner profile is constructed based on the cognitive resilience index and the knowledge mastery level. Based on the dual-dimensional learner profile and the reward function, an initial personalized learning path is generated through a multi-objective optimization algorithm; wherein, the initial personalized learning path includes at least one strategic dilemma task; During the process of the target learner executing the initial personalized learning path, the data on the target learner's dilemma response when interacting with strategic dilemma tasks are monitored in real time; Based on the aforementioned distress response data, the dual-dimensional learner profile is updated to obtain an updated dual-dimensional learner profile. Based on the updated dual-dimensional learner profile, the difficulty level and support level of subsequent tasks in the initial personalized learning path are adjusted to obtain an adjusted learning path.

2. The method according to claim 1, characterized in that, The implicit reward function is derived from the predefined behavioral demonstrations of highly cognitively resilient learners using an inverse reinforcement learning algorithm. Based on the standardized behavioral feature sequence and the reward function, the cognitive resilience index of the target learner is calculated, including: The behavioral trajectories of learners assessed by experts as having high cognitive resilience are obtained from a pre-defined historical database, and all such behavioral trajectories are aggregated to generate an expert demonstration dataset. Based on the expert demonstration dataset, a reward function is derived using the maximum entropy inverse reinforcement learning algorithm; wherein, the reward function is used to explain the behavioral preferences of highly resilient learners; Based on the reward function and the standardized behavioral feature sequence, the cognitive resilience index of the target learner is calculated.

3. The method according to claim 2, characterized in that, The calculation of the cognitive resilience index of the target learner based on the reward function and the standardized behavioral feature sequence includes: Based on the reward function, quantitative dimensions of cognitive resilience are determined; wherein, the quantitative dimensions include challenge endurance, strategy transferability, and emotional resilience. Based on the standardized behavioral feature sequence, calculate the original score for each of the quantification dimensions; The cognitive resilience index of the target learner is obtained by weighted summation of the original scores using the following formula: in, For cognitive resilience index, To challenge the raw score of endurance, The raw score for strategy transfer ability. The raw score for emotional resilience. To challenge the overall mean of durability, This represents the overall mean of strategy transfer capability. This represents the overall mean of emotional resilience. To challenge the overall standard deviation of persistence, The overall standard deviation of strategy transfer capability. The overall standard deviation of emotional resilience. To challenge the weighting coefficients of persistence, The weighting coefficients for policy transfer capability. This is the weighting coefficient for emotional resilience.

4. The method according to claim 1, characterized in that, Based on the dual-dimensional learner profile and the reward function, an initial personalized learning path is generated using a multi-objective optimization algorithm; wherein, the initial personalized learning path includes at least one strategic dilemma task, including: A multi-objective path optimization function is constructed, wherein the multi-objective path optimization function includes knowledge acquisition benefits, cognitive resilience cultivation benefits, and cognitive load costs: Based on the knowledge point mastery level in the dual-dimensional learner profile, a baseline value for the knowledge mastery benefit is determined. Based on the cognitive resilience index in the dual-dimensional learner profile, the difficulty range of the strategic dilemma task is determined; The Monte Carlo tree search algorithm is used to sample in the path space to obtain candidate paths; Based on the baseline value of the knowledge acquisition benefit and the difficulty range, a three-dimensional indicator is calculated for each candidate path; wherein, the three-dimensional indicator includes the knowledge acquisition benefit, the cognitive resilience cultivation benefit, and the cognitive load cost; The three-dimensional indicators of each candidate path are substituted into the multi-objective path optimization function to generate a value score for each candidate path; and the candidate path corresponding to the maximum value score is determined as the initial personalized learning path.

5. The method according to claim 1, characterized in that, The process of the target learner executing the initial personalized learning path includes real-time monitoring of the target learner's dilemma response data when interacting with strategic dilemma tasks, including: Obtain the mouse movement trajectory coordinate sequence of the target learner within the strategic dilemma task interface, and calculate the hesitation index of the mouse movement trajectory coordinate sequence; The keyboard input event stream of the target learner during the task is obtained, and the keyboard input event stream is analyzed to obtain the input rate change pattern and modification frequency. With authorization, the facial video stream of the target learner is acquired, and a pre-trained facial expression recognition model is used to extract micro-expression change sequences from the facial video stream; Calculate the percentage of duration of negative emotion intensity in the micro-expression change sequence; The hesitation index, the input rate change pattern, and the percentage of duration of negative emotion intensity are summarized to form the dilemma response data.

6. A personalized learning path generation device based on learning behavior profiles, characterized in that, The device includes: The data acquisition module is used to acquire multimodal learning behavior data of the target learner and perform feature extraction processing on the multimodal learning behavior data to obtain a standardized behavioral feature sequence; wherein, the multimodal learning behavior data includes cognitive interaction data, behavioral micro data and emotional physiological data; The index calculation module is used to deduce the implicit reward function from the predefined behavioral demonstrations of highly cognitively resilient learners based on the standardized behavioral feature sequence using an inverse reinforcement learning algorithm, and to calculate the cognitive resilience index of the target learner based on the standardized behavioral feature sequence and the reward function. The profile building module is used to obtain the knowledge mastery level of the target learner and to build a two-dimensional learner profile based on the cognitive resilience index and the knowledge mastery level. An initial path generation module is used to generate an initial personalized learning path based on the dual-dimensional learner profile and the reward function using a multi-objective optimization algorithm; wherein the initial personalized learning path includes at least one strategic dilemma task; The data monitoring module is used to monitor in real time the dilemma response data generated by the target learner when interacting with the strategic dilemma task during the execution of the initial personalized learning path by the target learner; The path update module is used to update the two-dimensional learner profile based on the dilemma response data to obtain the updated two-dimensional learner profile, and to adjust the difficulty level and support level of subsequent tasks in the initial personalized learning path based on the updated two-dimensional learner profile to obtain the adjusted learning path.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.