Generative and multi-modal sensing integrated agent learning system
Through the integrated generative and multimodal perception agent learning system, the dynamic knowledge graph is established using multimodal data and the teaching strategies are optimized in real time, the problems of inefficient learning efficiency and insufficient privacy protection in the online education system are solved, and more efficient learning paths and teaching strategies are achieved.
Patent Information
- Application Number
- CN202510775321.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing online education system has problems such as low learning efficiency, insufficient relevance of suggestions, failure to integrate teaching feedback mechanisms, inability to reflect real engineering capabilities, and lack of privacy protection mechanisms.
Adopt a fusion generative and multimodal perception of the agent learning system, including learner modeling module, teaching strategy planning module, multi-scale evaluation module, strategy optimization module and feedback coordination module, and establish a dynamic knowledge graph by collecting learners' multimodal data, evaluate and optimize teaching strategies in real time, and realize multimodal data fusion and intelligent collaboration.
It improves the adaptability of learning paths, reduces the response time of path reconstruction, improves the accuracy of learning behavior characteristics capture, enhances the adaptability and privacy protection of teaching strategies, and achieves the improvement of learning efficiency and effectiveness.
Smart Images

Figure CN120277628A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent learning systems, and in particular, relates to an intelligent agent learning system that integrates generative and multimodal perception. Background Art
[0002] With the digital transformation of software engineering education, existing technologies have made progress in the fields of online learning platforms, intelligent programming tools, and automated evaluation systems. However, since traditional online education systems (such as Moodle and Coursera) mostly adopt a linear course structure, they cannot dynamically adjust the teaching content according to the cognitive characteristics of learners (such as knowledge blind spots and learning rhythm). Studies have shown that about 68% of learners have low learning efficiency due to fixed teaching paths. Although mainstream IDE plug-ins (such as IntelliSense) can achieve basic code completion, they lack understanding of the project context, resulting in insufficient relevance of suggestions. For example, GitHub Copilot has an error suggestion rate of up to 42% in complex architecture projects (ICSE 2023), and does not integrate a teaching feedback mechanism. Existing automated evaluation tools (such as JUnit and SonarQube) focus on code correctness and static quality, ignoring development process behaviors (such as debugging strategies and document review modes) and soft skills (such as architecture design and technical debt management). ACM survey shows that 83% of software engineers believe that the existing evaluation system cannot reflect real engineering capabilities. Learner models, teaching engines, and programming tools usually run independently, forming "data islands". For example, the learning records of MOOCs platforms cannot guide programming behaviors in IDEs in real time, resulting in a disconnect between teaching strategies and practical training. Existing systems generally lack privacy protection mechanisms (such as differential privacy) and have weak cross-platform support. The 2022 NIST report pointed out that 61% of educational technology systems are at risk of sensitive data leakage, and only 23% support seamless mobile connections.
[0003] Therefore, the technical problems existing in the online education system in the existing technology, such as low learning efficiency, insufficient suggestion relevance, failure to integrate teaching feedback mechanism, inability to reflect real engineering capabilities, and lack of privacy protection mechanism, are technical problems that need to be solved urgently. Summary of the invention
[0004] The purpose of the present invention is to provide an intelligent agent learning system that integrates generative and multimodal perception, so as to solve the technical problems existing in the online education system in the prior art, such as low learning efficiency, insufficient suggestion relevance, no integrated teaching feedback mechanism, inability to reflect real engineering capabilities, and lack of privacy protection mechanism.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows: An intelligent agent learning system integrating generative and multimodal perception, including a learner modeling module, a teaching strategy planning module, a multi-scale evaluation module, a strategy optimization module, and a feedback coordination module. The learner modeling module is connected to the teaching strategy planning module, the teaching strategy planning module is connected to the multi-scale evaluation module, and the feedback coordination module is connected to each module respectively; The learner modeling module establishes a dynamic knowledge graph of the learner by collecting various daily learning data, physiological indicators, and basic data of the learner; The teaching strategy planning module creates a teaching strategy plan for the learner based on the dynamic knowledge graph of the learner; The multi-scale evaluation module evaluates the specified indicators of the learner and the specified indicators of the system after the learner executes the specified time period based on the teaching strategy plan; The strategy optimization module optimizes the teaching strategy planning module based on the evaluation results of the specified indicators of the multi-scale evaluation module, and then optimizes the teaching strategy plan of the learner; The feedback coordination module coordinates and optimizes the learner modeling module, the teaching strategy planning module, the multi-scale evaluation module, and the strategy optimization module based on the evaluation results of the specified indicators of the system.
[0006] Preferably, various daily learning data of the learner are collected through a multimodal behavior capture module. The multimodal behavior capture module includes a keyboard dynamics sensor, an eye tracker, and a speech emotion analyzer. The learner modeling module constructs a dynamic knowledge graph of the learner based on the relevant data obtained by the keyboard dynamics sensor, the eye tracker, and the speech emotion analyzer. The dynamic knowledge graph of the learner includes a knowledge topology structure, a metacognitive strategy library, and a neural representation encoding.
[0007] Preferably, the teaching strategy planning module includes a learning path generation module, a resource recommendation module, and a learning difficulty adaptive adjustment module. The learning path generation module constructs a dynamic behavior feature vector of the learner based on the learning interaction data of the learner including answering time, step jumpiness, and resource stay duration, and physiological data including pupil diameter change rate and brain wave signal. It applies a quantization clustering algorithm to decompose the subject knowledge system to form a dynamically combinable knowledge unit network, establishes a transfer relationship between knowledge points based on a semantic reasoning engine, locates knowledge blind spots through an ability diagnostic test, generates a basic learning path in combination with historical learning data, and deploys a reinforcement learning model to evaluate the learning effect in real time. When a thinking path deviation is detected, it triggers path reconstruction. For example, when a step jump error occurs in the learning of quadratic functions, the combination of number and shape training is enhanced. Based on a forgetting threshold prediction model, including combining the Ebbinghaus curve with real-time memory intensity monitoring, it automatically arranges the rhythm of knowledge reproduction.
[0008] Preferably, the teaching strategy planning module also generates a three-dimensional learning network by integrating cognitive levels, including knowledge mastery, learning styles, including visual / auditory preferences, development goals, including academic orientation and career orientation, simulates the long-term effects of different learning strategies through adversarial generative networks, selects the optimal path branch, and maps the progress of knowledge point mastery through a learning trajectory visualization system. A heat map is used to display the priority of ability enhancement, and a dynamic update mechanism is used to perform incremental training every hour, and the difficulty gradient of subsequent paths is adjusted according to the latest learning performance.
[0009] Preferably, the multi-scale assessment module includes a cognitive feature extraction module, a behavior pattern creation module and an assessment module; The cognitive feature extraction module is based on the synchronous analysis of pupil diameter change rate and brain wave signal, quantifies the intensity of attention allocation in the debugging process in real time, constructs a Markov decision process model, analyzes the correlation between code submission frequency and debugging path selection, and obtains learner cognitive features; The behavior pattern creation module adopts a federated contrastive learning framework, aggregates learner behavior data across platforms, including shortcut key usage preferences and breakpoint setting rules, deploys neural radiation field technology, and maps keyboard dynamics data (pressure sensitivity 0.1N) into a three-dimensional cognitive load heat map; The evaluation module evaluates the learner's designated indicators and the system's designated indicators based on the learner's cognitive characteristics and the three-dimensional cognitive load heat map.
[0010] Preferably, the strategy optimization module includes an evaluation index mapping module and a data fusion module; The evaluation indicator mapping module establishes a spatiotemporal mapping relationship between micro-behavior analysis, including the duration of stay on knowledge points, attribution patterns of wrong questions, and macro-ability prediction (including technical growth trajectory), constructs a learner ability development matrix, and analyzes teaching design evaluation data (including classroom interaction intimacy, teaching resource adaptability, etc.) through a cross-modal attention mechanism to generate a teaching strategy improvement heat map; The data fusion module applies quantized tensor decomposition technology to fuse three types of heterogeneous data: cognitive load characteristics (including eye tracking data), emotional state (including voice emotion recognition) and knowledge mastery (including test scores). Based on a dynamic weight allocation model, the contribution rate of each evaluation dimension is automatically adjusted according to the teaching stage, including focusing on cognitive load in the new course stage and focusing on knowledge transfer ability in the review stage.
[0011] Preferably, the strategy optimization module constructs a teaching strategy decision tree based on the learner ability development matrix and the teaching strategy improvement heat map, automatically switches the teaching mode according to the learner's real-time cognitive state, including working memory capacity and attention fluctuation, deploys a course difficulty elastic regulator, dynamically expands or contracts the teaching content boundary based on the knowledge graph mastery (node completion rate ≥ 85%), and identifies the group characteristics of learners, including visual / audio / kinesthetic types, through cluster analysis, generates personalized teaching resource packages, including 3D models / voice explanations / virtual reality, and triggers the reinforcement training module at the knowledge decay critical point (e.g., memory strength ≤ 60%) based on the forgetting curve prediction model.
[0012] Preferably, the feedback coordination module integrates the learner's learning behavior trajectories (including answering time and step correction rate), cognitive data (including the intensity of the β wave in the brain wave signal), and the knowledge graph node mastery, constructs a three-dimensional learner state matrix, deploys an LSTM-GRU hybrid network to monitor the changes in the learning environment in real time, identifies the fluctuations in the teaching resource adaptability, and can be set to trigger the strategy reconstruction at a ±20% threshold. Based on the temporal graph convolutional network, it predicts the potential obstacle nodes in the knowledge transfer path, defines the collaborative action space of the teaching agent (including the knowledge point explanation module), the training agent (including the exercise generation module), and the evaluation agent (including the ability diagnosis module), adopts the CPQL algorithm (a reinforcement learning strategy based on the consistency model) to achieve real-time strategy update, constructs a multi-scale reward function: micro reward: the knowledge point mastery rate (including the specified percentage increase in the correct rate per unit time), medium reward: the effectiveness of skill transfer (including the success rate of the transfer from trigonometric functions to vector operations), macro reward: the long-term ability development index (including the quarterly growth of the logical thinking index), and applies the MAPPO algorithm to achieve the collaborative optimization of multi-agent strategies, and balances the individual and global goals through sharing the critic network.
[0013] Preferably, it further includes a knowledge reinforcement module and a skill improvement path planning module. The knowledge reinforcement module identifies the weak nodes (mastery < 70%) in the knowledge graph based on the separated multi-modal attention mechanism, generates a three-dimensional visual reinforcement path, deploys a generative adversarial network (GAN) to simulate the effects of different training schemes, and selects the optimal reinforcement combination (including the reorganization of wrong questions + micro-lesson explanations); the skill improvement path planning module constructs a dynamic course learning model, automatically inserts interval review nodes according to the forgetting curve prediction (including the memory decay critical point ±5 minutes), and generates an elastic learning plan with time window constraints by integrating the teaching expert experience through the neuro-symbolic system.
[0014] The beneficial effects of the present invention include: The intelligent agent learning system integrating generative and multimodal perception provided by the present invention establishes a dynamic knowledge graph of learners by collecting various daily learning data, physiological indicators, and basic data of learners, creates a teaching strategy plan for learners based on the dynamic knowledge graph of learners. After the learners execute the specified time period based on the teaching strategy plan, the specified indicators of the learners and the specified indicators of the system are evaluated, and the teaching strategy planning module is optimized based on the evaluation results of the specified indicators of the multi-scale evaluation module, thereby optimizing the teaching strategy plan of the learners. The various modules are coordinately optimized based on the evaluation results of the specified indicators of the system. It solves the technical problems of low learning efficiency, insufficient relevance of suggestions, lack of integrated teaching feedback mechanism, inability to reflect real engineering capabilities, and lack of privacy protection mechanism existing in the existing online education system.
[0015] First, the data fusion module realizes multimodal data fusion. Through the collaborative perception of the keyboard dynamics sensor and the eye tracker, the accuracy of capturing learning behavior characteristics is effectively improved, and the Pearson correlation coefficient between brain wave signals and knowledge mastery is effectively improved, realizing the construction of a cognitive quantification model for neural representation encoding.
[0016] Second, by decomposing the subject knowledge system, a knowledge unit network that can be dynamically combined is generated, improving the adaptability of the learning path. The reinforcement learning model detects the deviation of the thinking path in real time, greatly reducing the response time for triggering path reconstruction and improving the path correction accuracy.
[0017] Third, based on the completion rate of knowledge graph nodes, the content boundary is dynamically expanded, effectively improving the coverage rate of high-order thinking training. The three-dimensional learning network maps the progress of knowledge point mastery, and the heat map shows the priority of ability strengthening, effectively improving the training intensity of key knowledge.
[0018] Third, through cross-modal evaluation of the federated contrast learning framework, the recognition value of behavior patterns is effectively improved, the misjudgment rate is reduced, and the multimodal fusion of emotional state and knowledge mastery significantly reduces the prediction error of teaching strategy adaptability.
[0019] Finally, the learner modeling module, teaching strategy planning module, multi-scale evaluation module, and strategy optimization module are optimized through the feedback coordination module, realizing intelligent collaboration among the modules. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the architecture of the intelligent agent learning system integrating generative and multimodal perception of the present invention.
[0021] Figure 2 It is a schematic diagram of the architecture of a traditional online learning system. DETAILED DESCRIPTION OF THE INVENTION
[0022] The following will further elaborate on the present invention in conjunction with the accompanying Figures 1 - 2 drawings: Embodiment 1 Referring to the accompanying Figure 1 drawings, an intelligent agent learning system integrating generative and multimodal perception includes a learner modeling module, a teaching strategy planning module, a multi-scale evaluation module, a strategy optimization module, and a feedback coordination module. The learner modeling module is connected to the teaching strategy planning module, the teaching strategy planning module is connected to the multi-scale evaluation module, and the feedback coordination module is connected to each module respectively. The learner modeling module creates a dynamic knowledge graph of the learner by collecting various daily learning data, physiological indicators, and basic data of the learner. The teaching strategy planning module creates a teaching strategy plan for the learner based on the dynamic knowledge graph of the learner. The multi-scale evaluation module evaluates the specified indicators of the learner and the specified indicators of the system after the learner executes the specified time period based on the teaching strategy plan. The strategy optimization module optimizes the teaching strategy planning module based on the evaluation results of the specified indicators of the multi-scale evaluation module, and then optimizes the teaching strategy plan of the learner. The feedback coordination module coordinately optimizes the learner modeling module, the teaching strategy planning module, the multi-scale evaluation module, and the strategy optimization module based on the evaluation results of the specified indicators of the system.
[0023] Referring to Figure 2 , in a traditional online learning system, learners register through a user registration module and conduct corresponding learning in a learning interaction module. A series of relevant learning data will be generated in the learning interaction module. The data analysis module analyzes the learning situation of the learners based on the learning data, and the system management module manages and adjusts the learning interaction module based on the data analysis results. Therefore, the traditional online learning system has technical problems such as low learning efficiency, insufficient relevance of suggestions, lack of integration of teaching feedback mechanisms, inability to reflect real engineering capabilities, and lack of privacy protection mechanisms. In this embodiment, the learner modeling module creates a dynamic knowledge graph of the learner by collecting various daily learning data, physiological indicators, and basic data of the learner. The daily learning data includes exam data, homework completion, classroom quizzes, and feedback data in the current learning stage and historical learning stages. The exam data includes exam time, exam scores, exam question types, exam difficulty, etc. The feedback data includes evaluations of courses and exams, evaluations of teachers, evaluations of the system, etc., reflecting the learning experience and learning effects of students. The physiological indicators include capturing keystroke pressure and answering interval time through the deployment of keyboard dynamics sensors to construct learning behavior time series data, using an eye tracker to record the pupil diameter change rate and fixation hotspot distribution, and integrating an Emotiv Epoc+ EEG headset to collectβ Wave intensity, quantifying the cognitive load level. The sampling rate of the keyboard dynamics sensor is 1000 Hz, and the resolution for capturing keystroke pressure is 0.1 N. The sampling rate of the eye tracker is 1200 Hz, and the β frequency range of the wave is 13 - 30 Hz.
[0024] In another implementation of this embodiment: The learner modeling module fuses multi-source data through a federated contrastive learning framework to construct map nodes containing the following dimensions: The first dimension is the knowledge topological structure, including decomposing the mathematics subject into several knowledge points to form a probabilistic graph model, and establishing connection edges with a conditional probability ≥ 0.85. The second dimension is the metacognitive strategy library, recording cross-disciplinary transfer paths, such as the success rate of transferring geometric proofs to physical mechanics. The third dimension is the neural representation encoding, which β establishes a mapping relationship between wave intensity and the degree of knowledge point mastery, with a Pearson coefficient r = 0.82. A dynamic knowledge graph of the learner is established based on the first dimension, the second dimension, and the third dimension.
[0025] The teaching strategy planning module decomposes junior high school mathematics into several knowledge units based on a quantization clustering algorithm, constructs a reconfigurable network, and evaluates the learning effect in real time through a reinforcement learning model.
[0026] Path reconstruction is triggered when the following situations are detected: The same knowledge point is wrongly repeated 3 times, the confidence level ≥ 90%, and the brain wave signal shows cognitive overload. The judgment criterion for the cognitive overload is β The wave intensity is greater than 45 μV and lasts for 5 minutes.
[0027] An adaptive adjustment mechanism is set: The course difficulty elastic regulator dynamically adjusts according to the completion rate of the map nodes of the knowledge graph: When the mastery level is lower than 60%, basic micro-lessons are inserted, and the duration of the basic micro-lessons does not exceed 8 minutes. When the mastery level is greater than 85%, advanced thinking training is extended. The visual interface maps the progress of knowledge points through a three-dimensional learning network, and a red hot zone is established to represent the knowledge nodes that need to be strengthened in training.
[0028] The multi-scale evaluation module extracts features from keyboard dynamics data through a ConvLSTM network to identify ineffective debugging modes (such as frequently using print statements), constructs a Markov decision process model from eye movement data to calculate the attention allocation efficiency, and determines that it meets the standard when the proportion of effective fixation is greater than 75%. The neural radiance field technology converts the electroencephalogram signal into a three-dimensional cognitive load heat map. The federated learning framework aggregates data across platforms to calculate the effectiveness index of skill transfer, and a transfer success rate greater than or equal to 68% is considered qualified. A causal inference model is deployed to verify the correlation between teaching strategies and performance improvement, with a Granger causal value greater than 0.6. The quarterly ability development index calculates the standardized growth in dimensions such as logical thinking and spatial imagination.
[0029] The strategy optimization module performs dynamic parameter adjustment of the system, fuses three types of heterogeneous data including quantization tensor decomposition technology, cognitive load characteristics, emotional state, and knowledge mastery. The weight of the cognitive load characteristics is set to 0.4, the emotional state includes the confidence of speech emotion recognition, and its weight is set to 0.3, and the weight of knowledge mastery is set to 0.3. The strategy reconstruction engine includes establishing a teaching strategy decision tree with 32 branch nodes, automatically switching modes according to the real-time state. When β the wave intensity suddenly increases by 20%, it switches to the video explanation mode. When the correct rate of two consecutive practices is greater than 90%, the difficulty level is increased. The forgetting curve intervention is to predict the knowledge decay point based on the Ebbinghaus model. When the memory retention rate is less than 50%, a review is triggered, and intensive training questions are pushed 5 minutes before the critical point, and the question type similarity is greater than 80%.
[0030] Example 2 Based on Example 1, various daily learning data of the learner are collected through a multi-modal behavior capture module. The multi-modal behavior capture module includes a keyboard dynamics sensor, an eye tracker, and a speech emotion analyzer. The learner modeling module constructs a dynamic knowledge graph of the learner based on the relevant data obtained by the keyboard dynamics sensor, the eye tracker, and the speech emotion analyzer. The dynamic knowledge graph of the learner includes a knowledge topology structure, a metacognitive strategy library, and a neural representation encoding.
[0031] The teaching strategy planning module includes a learning path generation module, a resource recommendation module, and a learning difficulty adaptive adjustment module. The learning path generation module constructs a dynamic behavior feature vector of the learner based on the learning interaction data of the learner including the answering time, step skipping, and resource staying duration, and the physiological data including the pupil diameter change rate and brain wave signals. It applies the quantization clustering algorithm to decompose the subject knowledge system to form a dynamically combinable knowledge unit network, establishes the migration relationship between knowledge points based on the semantic reasoning engine, locates the knowledge blind area through the ability diagnosis test, generates a basic learning path in combination with historical learning data, and deploys a reinforcement learning model to evaluate the learning effect in real time. When a thinking path deviation is detected, a path reconstruction is triggered. For example, when a step skipping error occurs in the learning of quadratic functions, the combination of number and shape training is enhanced. Based on the forgetting threshold prediction model, including combining the Ebbinghaus curve with real-time memory intensity monitoring, the knowledge reproduction rhythm is automatically arranged.
[0032] This embodiment takes the special training of quadratic functions in mathematics as an example: First, construct a dynamic behavior feature vector, collect learning interaction data, including answering time, step skipping index, and resource staying duration, and integrate physiological data, including pupil diameter change rate and prefrontal β wave intensity.
[0033] Then, perform quantization clustering on the knowledge system, decompose the quadratic function into several knowledge units, and establish migration relationships through a semantic reasoning engine.
[0034] Next, design a path dynamic reconstruction mechanism, deploy a PPO reinforcement learning model, and trigger reconstruction when the following situations are detected: step jump errors, such as directly applying formulas and skipping image analysis; a sudden drop in β-wave intensity greater than 25% and lasting for 30 seconds is determined as cognitive fatigue. Example of a reconstruction strategy: insert 3 special training questions on the combination of numbers and shapes, including dynamic geometry demonstrations.
[0035] The resource recommendation module obtains cross-platform (PC / tablet) aggregated behavior data through a federated contrast learning framework: high-frequency misoperation patterns, video viewing completion rate, and uses a differential privacy mechanism (ε = 1.5, δ = 1e-5) to ensure data security. Construct resource feature vectors through a multi-modal feature matching algorithm.
[0036] The learning difficulty adaptive adjustment module monitors real-time cognitive load, with a pressure perception threshold model: maintain the current difficulty in the green interval (β-wave intensity 15 - 35 μV); a red warning (β-wave > 45 μV for 2 minutes), automatically downgrade to basic question types. And set an elastic adjustment mechanism to calculate the difficulty coefficient D, with the specific expression as follows: D = 0.6×(historical correct rate / benchmark value) + 0.4×(real-time reaction speed / threshold); Among them, the historical correct rate represents the proportion of correct answers in specified tests and answering activities carried out in the past, the benchmark value is a preset fixed value as a reference standard for measuring the historical correct rate, the real-time reaction speed represents the time taken by the test taker from the presentation of the question to giving a response during the current answering or task execution, and the threshold is a preset fixed value as a reference limit for measuring the real-time reaction speed. The dynamic adjustment range is 0.5 - 1.8, with 0.5 being the basic difficulty and 1.8 being the extended difficulty.
[0037] The teaching strategy planning module also generates a three-dimensional learning network by integrating cognitive levels, including knowledge mastery, learning style, and development goals. The learning style includes visual / auditory preferences, and the development goals include entrance examination orientation and career orientation. Simulate the long-term effects of different learning strategies through a generative adversarial network, select the optimal path branch, a learning trajectory visualization system maps the progress of knowledge point mastery, shows the priority of ability enhancement through a heat map, and a dynamic update mechanism performs incremental training every hour and adjusts the subsequent path difficulty gradient according to the latest learning performance.
[0038] The teaching strategy planning module also calculates the knowledge mastery degree when integrating cognitive levels, and the calculation formula is as follows: KML = w 1×(historical correct rate / benchmark value) +w 2 × (Average score of the most recently specified tests / Full score); where, KML is the knowledge mastery level, w 1 and w 2 are the weights of historical knowledge mastery level and most recent knowledge mastery level respectively. The historical correct rate represents the proportion of correctly answered questions in the specified tests or learning evaluations conducted in the past. The benchmark value is a preset fixed value, serving as a reference standard for measuring the historical correct rate. The average score of the most recently specified tests represents the average value of each test score in the most recently specified number of tests.
[0039] Learning style recognition is measured by the visual preference index and the auditory preference index. The visual preference index is calculated by the proportion of video viewing duration × pupil focus stability. The auditory preference index is calculated by the audio resource replay rate + the frequency of using voice notes.
[0040] Development goal embedding includes college entrance examination orientation and career orientation. For the college entrance examination orientation, the weight of training with college entrance examination real questions is increased by 30%. For the career orientation, the priority of the engineering case library is raised.
[0041] Generate a three-dimensional feature vector based on the learner's knowledge mastery level, learning style, and development goal. Construct a three-dimensional vector space based on the three-dimensional feature vector, and use the t-SNE algorithm to perform dimensionality reduction visualization on the three-dimensional vector space. Create an adversarial generative network, input the three-dimensional feature vector and time series learning data into the adversarial generative network, and the adversarial generative network outputs the learning path branches within a specified future time period.
[0042] The adversarial generative network is optimized through a discriminator, and a dynamic knowledge gain evaluation function is set: ; where, R is the knowledge gain evaluation value, t is the time step, T is the termination time step, γ is the time discount factor, which is a constant between 0 and 1. As the time step t increases, the value of γ t−1 will gradually decrease, is the knowledge gain coefficient, is the cognitive load coefficient. The knowledge gain t represents the knowledge gain obtained at the time step t , which is a numerical value quantifying the degree of knowledge acquisition or improvement at this time point. The cognitive load t represents the cognitive load borne at the time step t , which is a numerical value quantifying the degree of cognitive resources or stress consumed by the brain when processing information at this time point.
[0043] Perform dynamic optimization of the learning path based on the values of the dynamic knowledge gain evaluation function. Execute incremental Monte Carlo tree search at specified intervals, and establish trigger conditions for switching learning path branches. The switching trigger conditions can be set to the daily knowledge gain being lower than a preset expected value threshold, or the cognitive load being greater than a preset threshold for a continuous specified period.
[0044] The multi-scale evaluation module includes a cognitive feature extraction module, a behavior pattern creation module, and an evaluation module. The cognitive feature extraction module quantifies the attention allocation intensity during the debugging process in real time based on the synchronous analysis of the pupil diameter change rate and brain wave signals, constructs a Markov decision process model, analyzes the correlation between code submission frequency and debugging path selection, and obtains the learner's cognitive features. The behavior pattern creation module uses a federated contrast learning framework to aggregate learner behavior data across platforms, including shortcut key usage preferences and breakpoint setting rules, and deploys neural radiance field technology to map keyboard dynamics data into a three-dimensional cognitive load heat map. The evaluation module evaluates the specified indicators of the learner and the system based on the learner's cognitive features and the three-dimensional cognitive load heat map.
[0045] When the multi-scale evaluation module quantifies the attention allocation intensity during the debugging process in real time, it constructs a spatio-temporal alignment model to calculate the attention intensity index from the acquired multi-modal data, constructs the state transition matrix of the Markov decision process model, and analyzes the optimal debugging path through the Viterbi algorithm to identify inefficient patterns, such as frequent ineffective breakpoints.
[0046] The multi-scale evaluation module creates a differential privacy mechanism based on the federated contrast learning framework: ; where σ is the noise standard deviation, here σ = 0.3, is the privacy budget, which is the maximum logarithmic difference controlling the probability ratio of the algorithm output. The smaller it is, the higher the privacy protection intensity. is the relaxation parameter, which represents the upper limit of the acceptable privacy leakage probability. It allows the algorithm to break the strict privacy protection limit with a certain probability, usually taking a very small positive number, and N is the number of data samples participating in federated learning.
[0047] The three-dimensional cognitive load heat map is generated based on the pressure-time mapping algorithm and outputs heat map levels based on spatio-temporal convolutional kernels. In the 10-school joint experiment, federated learning improved the accuracy of the model in identifying the following patterns. The detection accuracy of formula derivation bottlenecks increased from 78% to 92%, and the identification of misused geometric auxiliary lines increased from 65% to 84%.
[0048] The evaluation metrics of the multi-scale evaluation module include the thinking coherence index, the timeliness of system recommendations, and technical effect data. The calculation formula of the thinking coherence index (TCI) is as follows: ; Wherein, i is the status serial number, n is the total number of states, and the effective state duration i represents the duration of the thinking in the effective state in the i th state. The total problem-solving time represents the total duration from the start to the end of problem-solving, and the number of abnormal fluctuations represents the number of times of abnormal fluctuations in thinking during the entire problem-solving process. TCI Students with >0.7 have about a 23% improvement in final exam scores.
[0049] The timeliness of system recommendations is achieved by defining the golden intervention window, and a push prompt is sent within 30 seconds after detecting a 20% decrease in TCI.
[0050] The strategy optimization module includes an evaluation index mapping module and a data fusion module; The evaluation index mapping module establishes a spatio-temporal mapping relationship between micro-behavior analysis, including the knowledge point residence duration and the wrong-question attribution pattern, and macro-capability prediction, constructs a learner ability development matrix, and analyzes the instructional design evaluation data through a cross-modal attention mechanism to generate a heat map for improving teaching strategies. Macro-capability prediction includes the technical growth trajectory, and the instructional design evaluation data includes classroom interaction intimacy, teaching resource suitability, etc.
[0051] The data fusion module applies quantization tensor decomposition technology to fuse three types of heterogeneous data: cognitive load characteristics, emotional state, and knowledge mastery. Based on a dynamic weight allocation model, it automatically adjusts the contribution rate of each evaluation dimension according to the teaching stage. For example, in the new lesson stage, it focuses on cognitive load, and in the review stage, it focuses on knowledge transfer ability. Cognitive load characteristics include eye movement tracking data, emotional state includes speech emotion recognition, and knowledge mastery includes test scores.
[0052] The strategy optimization module constructs a teaching strategy decision tree based on the learner ability development matrix and the heat map for improving teaching strategies, automatically switches the teaching modality according to the learner's real-time cognitive state, including working memory capacity and attention fluctuations, deploys a course difficulty elastic regulator, dynamically expands or contracts the teaching content boundary based on the knowledge graph mastery, and identifies learner group characteristics, including visual / aural / kinesthetic types, through cluster analysis to generate personalized teaching resource packages, including 3D models / speech explanations / virtual reality. Based on the forgetting curve prediction model, a reinforcement training module is triggered at the knowledge decay critical point, and the knowledge decay critical point is set to a memory strength lower than 60%.
[0053] The feedback coordination module integrates the learning behavior trajectory, cognitive data, and the mastery degree of knowledge graph nodes of learners to construct a three-dimensional learner state matrix. The learning behavior trajectory includes the answering time, step correction rate, and the cognitive data includes brain wave signal β wave intensity. Deploy an LSTM-GRU hybrid network to monitor the changes in the learning environment in real time, identify fluctuations in the adaptability of teaching resources, and a threshold of ±20% can be set to trigger policy reconstruction. Based on the temporal graph convolutional network, predict potential obstacle nodes in the knowledge transfer path. Define the collaborative action space of teaching agents (including knowledge point explanation modules), training agents, and evaluation agents, and use the CPQL algorithm to achieve real-time policy updates. Construct a multi-scale reward function: micro reward: the knowledge point mastery rate, medium reward: the effectiveness of skill transfer, macro reward: long-term ability development indicators, and apply the MAPPO algorithm to achieve multi-agent policy collaborative optimization, and balance individual and global goals through sharing the critic network.
[0054] The teaching agent includes a knowledge point explanation module, the training agent includes an exercise generation module, the evaluation agent includes an ability diagnosis module, the CPQL algorithm is a reinforcement learning strategy based on a consistency model, the knowledge point mastery rate includes an increase in the correct rate by a specified proportion per unit time, the effectiveness of skill transfer includes the success rate of transferring trigonometric functions to vector operations, and the long-term ability development indicator includes the quarterly growth of the logical thinking index.
[0055] It also includes a knowledge reinforcement module and a skill improvement path planning module. The knowledge reinforcement module, based on a separate multi-modal attention mechanism, identifies weak nodes in the knowledge graph, generates a three-dimensional visual reinforcement path, deploys an adversarial generation network to simulate the effects of different training schemes, and selects the optimal reinforcement combination. The optimal reinforcement combination includes wrong question reorganization + micro-lesson explanation. The skill improvement path planning module constructs a dynamic course learning model, predicts and automatically inserts spaced review nodes according to the forgetting curve, and generates an elastic learning plan with time window constraints by integrating the experience of teaching experts through a neuro-symbolic system.
[0056] The separate multi-modal attention mechanism of the knowledge reinforcement module performs data fusion on the input data to create the characteristics of the learner's knowledge graph nodes. The input data includes the knowledge mastery degree, the attention weights of related knowledge points, learning behavior data, and physiological signals. The learning behavior data includes the wrong question repetition rate, video pause frequency, and interaction response time. The calculation formula for the attention weight is as follows: w i = softmax ( Q · k i T / d k ); Among them, Q is the query vector of the current knowledge node, k i is the multi-modal feature key vector, with a dimension d k of 128, and the output weight is used to identify weak nodes.
[0057] The three-dimensional visualization reinforcement path is realized based on the constructed learning reinforcement path topological structure. The visualization rule is to use specified color spheres to represent unmastered knowledge nodes, and green connecting lines to represent the optimal learning reinforcement path.
[0058] The skill improvement path planning module is realized by creating a dynamic course learning model. The dynamic course learning model sets a forgetting curve prediction model. The expression of the forgetting curve prediction value of the forgetting curve prediction model is as follows: R ( t ) = e -λt + β· EEG memory marking intensity; Among them, λ is the knowledge decay coefficient set to 0.03, β is the physiological signal correction weight, set to 0.2, t is the time step. The EEG memory marking intensity represents the intensity or activity of neural activities related to specific memories in the brain. Through EEG signal monitoring technology, EEG activity characteristics related to memories are obtained and quantified into a numerical value.
[0059] When the forgetting curve prediction model predicts that the memory retention rate R ( t ) drops to 58% ± 3%, an interval review node is automatically inserted. When the review node conflicts with the original plan, start the priority sorting: Priority = 0.6 × Forgetting Urgency + 0.4 × Knowledge System Centrality Priority = 0.6 × Forgetting Urgency + 0.4 × Knowledge System Centrality.
[0060] In summary, the intelligent agent learning system that integrates generative and multimodal perception provided by the present invention establishes a dynamic knowledge graph of learners through various daily learning data, physiological indicators, and basic data of learners, and then creates a teaching strategy plan for learners. After the learners execute the specified time period based on the teaching strategy plan, the specified indicators of the learners are evaluated, and the teaching strategy plan module is optimized based on the evaluation results, and then the teaching strategy plan for learners is optimized. Each module is coordinately optimized based on the evaluation results of the specified indicators. This solves the technical problems existing in the online education system in the prior art, such as low learning efficiency, insufficient relevance of suggestions, lack of integration of teaching feedback mechanisms, inability to reflect real engineering capabilities, and lack of privacy protection mechanisms.
[0061] Multimodal data fusion is achieved through the data fusion module, and the collaborative perception of keyboard dynamics sensor data and eye movement tracker data is fused, effectively improving the accuracy of capturing learning behavior characteristics and realizing the construction of a cognitive quantification model for neural representation coding. By decomposing the subject knowledge system, a knowledge unit network that can be dynamically combined is generated, improving the adaptability of the learning path. The reinforcement learning model detects the deviation of the thinking path in real time, greatly reducing the response time for triggering path reconstruction and improving the accuracy of path correction.
[0062] Based on the completion rate of knowledge graph nodes, the content boundary is dynamically expanded, effectively improving the coverage rate of higher-order thinking training. The three-dimensional learning network maps the progress of knowledge point mastery, and the heat map shows the priority of ability enhancement, effectively improving the training intensity of key knowledge. Through the cross-modal evaluation of the federated contrast learning framework, the recognition value of behavior patterns is effectively improved, the misjudgment rate is reduced, and the multimodal fusion of emotional state and knowledge mastery significantly reduces the prediction error of teaching strategy adaptability. Through the feedback coordination module, the learner modeling module, the teaching strategy planning module, the multi-scale evaluation module, and the strategy optimization module are optimized, realizing the intelligent collaboration between modules.
Claims
1. An intelligent agent learning system that integrates generative and multi-modal perception, characterized in that, It includes a learner modeling module, a teaching strategy planning module, a multi-scale evaluation module, a strategy optimization module, and a feedback coordination module. The learner modeling module is connected to the teaching strategy planning module, the teaching strategy planning module is connected to the multi-scale evaluation module, and the feedback coordination module is connected to each module respectively; The learner modeling module creates a dynamic knowledge graph of the learner by collecting various daily learning data, physiological indicators, and basic data of the learner; The teaching strategy planning module creates a teaching strategy plan for the learner based on the dynamic knowledge graph of the learner; The multi-scale evaluation module evaluates the specified indicators of the learner and the specified indicators of the system after the learner executes the specified time period based on the teaching strategy plan; The strategy optimization module optimizes the teaching strategy planning module based on the evaluation results of the specified indicators of the multi-scale evaluation module, and further optimizes the teaching strategy plan of the learner; The feedback coordination module coordinates and optimizes the learner modeling module, the teaching strategy planning module, the multi-scale evaluation module, and the strategy optimization module based on the evaluation results of the specified indicators of the system.
2. The intelligent agent learning system integrating generative and multimodal perception according to claim 1, characterized in that, The various daily learning data of the learner are collected through a multi-modal behavior capture module. The multi-modal behavior capture module includes a keyboard dynamics sensor, an eye tracker, and a speech emotion analyzer. The learner modeling module constructs a dynamic knowledge graph of the learner based on the relevant data obtained by the keyboard dynamics sensor, the eye tracker, and the speech emotion analyzer. The dynamic knowledge graph of the learner includes a knowledge topology structure, a metacognitive strategy library, and a neural representation encoding.
3. An intelligent agent learning system integrating generative and multimodal perception according to claim 1, characterized in that, The teaching strategy planning module includes a learning path generation module, a resource recommendation module, and a learning difficulty adaptive adjustment module. The learning path generation module constructs a dynamic behavior feature vector of the learner based on the learning interaction data of the learner including answer time consumption, step jumpiness, and resource stay duration, and physiological data including pupil diameter change rate and brain wave signals, applies a quantization clustering algorithm to decompose the subject knowledge system to form a dynamically combinable knowledge unit network, establishes a migration relationship between knowledge points based on a semantic reasoning engine, locates knowledge blind spots through an ability diagnostic test, generates a basic learning path in combination with historical learning data, and deploys a reinforcement learning model to evaluate the learning effect in real time. When a thinking path deviation is detected, it triggers path reconstruction, and automatically arranges the knowledge reproduction rhythm based on a forgetting threshold prediction model.
4. An intelligent agent learning system integrating generative and multimodal perception according to claim 3, characterized in that, The teaching strategy planning module also generates a three-dimensional learning network by integrating cognitive level, learning style, and development goal, simulates the long-term effects of different learning strategies through an adversarial generation network, selects the optimal path branch, a learning trajectory visualization system maps the progress of knowledge point mastery, displays the priority of ability enhancement through a heat map, and a dynamic update mechanism performs incremental training every hour and adjusts the subsequent path difficulty gradient according to the latest learning performance.
5. An intelligent agent learning system integrating generative and multimodal perception according to claim 3, characterized in that, The multi-scale evaluation module includes a cognitive feature extraction module, a behavior pattern creation module, and an evaluation module; The cognitive feature extraction module quantifies the attention allocation intensity during the debugging process in real time based on the synchronous analysis of the pupil diameter change rate and brain wave signals, constructs a Markov decision process model, analyzes the correlation between code submission frequency and debugging path selection, and obtains the learner's cognitive features; The behavior pattern creation module adopts a federated contrast learning framework, aggregates learner behavior data across platforms and deploys neural radiance field technology to map keyboard dynamics data into a three-dimensional cognitive load heat map; The evaluation module evaluates the specified indicators of the learner and the specified indicators of the system based on the learner's cognitive features and the three-dimensional cognitive load heat map.
6. An intelligent agent learning system integrating generative and multimodal perception according to claim 1, characterized in that, The strategy optimization module includes an evaluation index mapping module and a data fusion module; The evaluation index mapping module establishes a spatio-temporal mapping relationship between micro-behavior analysis and macro-capability prediction, constructs a learner ability development matrix, and analyzes the instructional design evaluation data through a cross-modal attention mechanism to generate a heat map for improving teaching strategies; The data fusion module applies quantization tensor decomposition technology to fuse three types of heterogeneous data: cognitive load features, emotional states, and knowledge mastery levels. Based on a dynamic weight allocation model, it automatically adjusts the contribution rate of each evaluation dimension according to the teaching stage.
7. An intelligent agent learning system integrating generative and multimodal perception according to claim 6, characterized in that, The strategy optimization module constructs a teaching strategy decision tree based on the learner ability development matrix and the heat map for improving teaching strategies, automatically switches the teaching modality according to the learner's real-time cognitive state, deploys a course difficulty elastic regulator, dynamically expands or contracts the teaching content boundary based on the knowledge graph mastery level, and identifies the group characteristics of learners through cluster analysis to generate a personalized teaching resource package. Based on the forgetting curve prediction model, it triggers the reinforcement training module at the critical point of knowledge decay.
8. An intelligent agent learning system integrating generative and multimodal perception according to claim 6, characterized in that, The feedback coordination module integrates the learner's learning behavior trajectory, cognitive data, and knowledge graph node mastery level, constructs a three-dimensional learner state matrix, and deploys an LSTM-GRU hybrid network to monitor changes in the learning environment in real time, identify fluctuations in the adaptability of teaching resources, predict potential obstacle nodes in the knowledge transfer path based on a temporal graph convolutional network, define the collaborative action space of teaching agents, training agents, and evaluation agents, adopt the CPQL algorithm to achieve real-time policy updates, construct a multi-scale reward function: micro reward: knowledge point mastery rate, meso reward: effectiveness of skill transfer, macro reward: long-term ability development indicator, and apply the MAPPO algorithm to achieve multi-agent policy collaborative optimization, and balance individual and global goals through a shared critic network.
9. The intelligent agent learning system integrating generative and multi-modal perception according to claim 1, characterized in that, It also includes a knowledge reinforcement module and a skill improvement path planning module. The knowledge reinforcement module, based on a separate multi-modal attention mechanism, identifies weak nodes in the knowledge graph, generates a three-dimensional visual reinforcement path, deploys an adversarial generation network to simulate the effects of different training schemes, and selects the optimal reinforcement combination; the skill improvement path planning module constructs a dynamic course learning model, automatically inserts spaced review nodes according to the forgetting curve prediction, and generates an elastic learning plan with time window constraints by integrating the experience of teaching experts through a neuro-symbolic system.
Citation Information
Cited By
AI-enabled intelligent teaching personalized service method and system
CN120542747A
An AI-enabled intelligent teaching personalized service method and system
CN120542747B
Endocrine nursing teaching decision-making system based on deep learning
CN120598751A
Personalized scientific education system based on AI agent driving
CN120725253A
An AI agent-driven personalized science education system
CN120725253B