A Big Data-Based Information Technology Teaching Optimization Method
By collecting students' multimodal data, using deep learning and multi-head attention mechanisms to fusion data, building a dynamic knowledge graph, and using graph convolution networks and reinforcement learning to optimize learning paths, the problem of existing systems being difficult to dynamically adjust learning paths is solved, and the scientific optimization of personalized learning and teaching strategies is achieved, which improves learning efficiency and teaching effectiveness.
Patent Information
- Application Number
- CN202510259032.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing information-based teaching system is difficult to dynamically adjust students' personalized learning paths and cannot update students' knowledge points in real time, resulting in insufficient optimization of teaching strategies, making it difficult to diagnose learning bottlenecks and analyze the causal relationship between teaching behavior and learning results.
By collecting students' behavior, emotions and physiological data, using deep learning and multi-head attention mechanisms to fusion data, building a dynamic knowledge graph, updating students' knowledge mastery status in real time, and optimizing learning paths using graph convolutional networks and reinforcement learning. At the same time, based on causal inference, the impact of teaching behavior on learning effects is analyzed and the teaching strategy is adjusted.
It realizes dynamic adjustment of personalized learning paths and scientific optimization of teaching strategies, improves learning efficiency and teaching effect, and can more accurately diagnose learning bottlenecks and analyze the causal relationship between teaching behavior and learning results.
Smart Images

Figure CN119741175B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information-based teaching, and more specifically, the present invention relates to an information-based teaching optimization method based on big data. Background Art
[0002] The optimization of information-based teaching is a method based on digitization and intelligence that optimizes the teaching process through big data technology, artificial intelligence, and data-driven decision-making. By collecting students' learning data, teachers' teaching data, and the usage data of teaching resources, it can accurately analyze the problems existing in the teaching process and provide a scientific basis for personalized learning path recommendation, resource allocation optimization, and teaching strategy adjustment.
[0003] Deficiencies of the prior art: The dynamic adjustment ability of students' personalized learning paths is limited. Most teaching systems lack in-depth mining and precise adaptation to students' real-time learning behaviors, and it is difficult to dynamically adjust learning content and difficulty according to students' immediate performance. The construction of knowledge graphs is mostly a static model, which cannot update students' knowledge mastery in real time and be associated with the dynamics of the subject, resulting in inaccurate diagnosis of learning bottlenecks. The analysis of teaching behaviors and learning outcomes mostly stays at the correlation level, and causal inference technology is not used to deeply explore the direct impact of different teaching strategies on learning effects, making it difficult to provide a scientific basis for optimizing teaching strategies and causing the lag of teaching strategies. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the following solutions are provided to solve the problem of poor teaching management in the information-based teaching process in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] An information-based teaching optimization method based on big data, comprising the following steps:
[0007] Collect students' behavior, emotion, and physiological data, perform synchronous alignment and normalization processing, extract features and reduce dimensions according to deep learning, and fuse various modal features through a multi-head attention mechanism;
[0008] Construct a knowledge point and dependency relationship graph, update students' knowledge mastery status in real time, and based on students' performance, propagate knowledge features according to the graph convolutional network to adjust the learning path and knowledge mastery status;
[0009] Define the state space and action space, optimize the personalized learning path according to immediate feedback and reinforcement learning, and adjust the task in real time;
[0010] Analyze the impact of teaching behaviors on learning effects based on causal inference, establish a causal relationship model to quantify the intervention effect, and adjust teaching strategies.
[0011] In a preferred embodiment, student behavior, emotion, and physiological data are collected, synchronized, aligned, and normalized. Features are extracted and dimensionally reduced based on deep learning. The specific steps are as follows:
[0012] The collected behavior data includes clickstream logs, task completion duration, answer accuracy rate, and learning frequency;
[0013] Emotion data is collected by capturing facial expressions and postures through a camera and collecting speech through a microphone;
[0014] Physiological data is collected by obtaining students' physiological signals through wearable devices. The physiological signals include heart rate changes and skin conductance;
[0015] For the collected data at different time granularities, the time window slicing method is used for alignment, and the logarithmic scaling function is used for normalization;
[0016] Behavior data feature extraction is performed to extract complex patterns in the time series, including the correlation between task completion time and click behavior;
[0017] Emotion data feature extraction is performed to extract emotional features from facial expressions and speech signals through a convolutional neural network. The emotional features include attention and stress changes;
[0018] Physiological data feature extraction is performed to extract the feature vectors of physiological signals using a variational autoencoder;
[0019] For the extracted features, a non-linear dimensionality reduction method is used for dimensionality reduction processing.
[0020] In a preferred embodiment, multi-modal features are fused through a multi-head attention mechanism. The specific steps include:
[0021] The collected data features are constructed into a physiological data feature matrix, an emotional data feature matrix, and a behavior data feature matrix respectively, and the feature matrices are fused through a multi-head attention mechanism;
[0022] The fused feature matrix is input into a graph embedding model to generate a global embedding vector of the student's learning state.
[0023] In a preferred embodiment, a knowledge point and dependency relationship graph is constructed to update the student's knowledge mastery status in real time. The specific steps are as follows:
[0024] Knowledge point and association relationship initialization is performed to extract knowledge points and the dependency relationships between knowledge points from teaching resources;
[0025] Node definition is performed. The knowledge point is the basic unit of the graph, and each node represents a specific teaching knowledge point;
[0026] Define the edges and calculate the weights. The edges represent the association relationships between knowledge points, and the co-occurrence probability is used to mine the forward and backward dependence relationships between knowledge points from teaching resources. The formula is: , where is the knowledge point and are the number of times they appear simultaneously in the same question. The co-occurrence probability represents the association strength between knowledge points and serves as the weight of the edge;
[0027] Initialize the status of knowledge points. The initial status Based on the matching degree calculation between the student characteristics embedding and the knowledge points, the dot product method is used to measure the matching degree between the student embedding vector and the knowledge point embedding vector : , where is the mapping matrix for learning, is the bias term, T represents the transpose operation, is the Sigmoid function;
[0028] Update the student performance feedback. Dynamically update the status of knowledge points according to the student's performance in tests and assignments : , where represents the test score of the student on the knowledge point , represents the dependence strength between the knowledge points , , is the learning rate.
[0029] In a preferred embodiment, and based on the student performance, propagate the knowledge features according to the graph convolutional network, and adjust the learning path and the knowledge mastery status, including the following steps:
[0030] Use the graph attention network to propagate the knowledge point features and update the feature representation of each knowledge point: , where the attention weight is calculated through the association strength between the knowledge points and and the current feature representation: , where and are the student's learning status vector and the knowledge point's feature vector, sim is the similarity measurement function, and N is the total number of knowledge points;
[0031] Based on the student's latest learning performance and the propagated features, adjust the node attributes and edge weights of the knowledge graph in real time.
[0032] In a preferred embodiment, a state space and an action space are defined, and the personalized learning path is optimized according to immediate feedback and reinforcement learning, and the tasks are adjusted in real time, including the following steps:
[0033] Define the state space, where the states in the state space include the current learning state of the student and the node features of the dynamic knowledge graph. The learning state of the student is represented by a feature vector, and the node features of the dynamic knowledge graph are represented by the updated knowledge point features;
[0034] Define the action space, where the actions in the action space represent the recommended learning tasks, and the action set is dynamically generated based on the node relationships in the dynamic knowledge graph and the student's knowledge mastery status;
[0035] Define the immediate reward as the trade-off between the improvement of knowledge mastery and the time cost after completing the learning task as the reward function;
[0036] Capture the student's current learning situation through the state, provide task selection through the action, and evaluate the learning task through the reward function;
[0037] Input the state into the policy function, output the action probability distribution, and use the PPO algorithm to optimize the policy function;
[0038] Based on the optimized policy, execute the recommended learning tasks at each time step, collect the student's learning feedback, and adjust the tasks according to the learning feedback.
[0039] In a preferred embodiment, analyze the impact of teaching behaviors on learning effects based on causal inference, establish a causal relationship model to quantify the intervention effect, and adjust the teaching strategy, including the following steps:
[0040] Define the causal variables. The independent variables represent the intervention variables of teaching behaviors, including the frequency of classroom interaction, the task assignment strategy, and the complexity of teaching content;
[0041] The dependent variables represent the student learning outcomes, including the knowledge point mastery degree and the answer correctness rate;
[0042] The confounding variables represent the factors that simultaneously affect the independent variables and the dependent variables;
[0043] Use a Bayesian network to automatically learn the causal graph structure from historical teaching data. The nodes in the causal graph represent variables, and the edges represent causal relationships. The learning process determines the network structure according to the method of maximizing the log-likelihood estimation;
[0044] Input the constructed causal graph into the causal inference stage, analyze the intervention effect of the teaching strategy, and define the causal effect;
[0045] According to the Do-Calculus rules, isolate the estimation path of the intervention effect from the causal graph;
[0046] Optimize the teaching strategy and dynamically adjust the teaching behavior according to the analysis results of the intervention effect.
[0047] The technical effects and advantages of an information-based teaching optimization method based on big data according to the present invention:
[0048] Through multi-modal data collection and deep learning technology, the present invention accurately captures students' behaviors, emotions, and physiological states, realizes data fusion through the multi-head attention mechanism, generates accurate learning state features, and based on the constructed dynamic knowledge graph and graph convolutional network, updates the students' knowledge mastery status in real time and adjusts the learning path to ensure personalized learning support. Then, reinforcement learning is used to optimize the personalized learning path, and the learning tasks are adjusted in real time in combination with instant feedback to improve learning efficiency and experience. At the same time, through causal inference, the impact of teaching behaviors on learning effects is analyzed, the intervention effect is quantified, and the teaching strategy is scientifically adjusted to improve the teaching effect. This solution combines multi-modal fusion, graph convolutional network, and reinforcement learning to optimize the learning path and teaching strategy, and has significant advantages in personalized learning support, teaching effect improvement, and learning efficiency optimization. Brief Description of the Drawings
[0049] Figure 1 It is a schematic flowchart of an information-based teaching optimization method based on big data according to the present invention. Detailed Embodiments
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] To achieve the above object, Figure 1 A schematic structural diagram of an information-based teaching optimization method based on big data according to the present invention is given, which specifically includes the following steps;
[0052] Collect students' behavior, emotion, and physiological data, perform synchronous alignment and normalization processing, extract features and reduce dimensions according to deep learning, and fuse features of each modality through the multi-head attention mechanism;
[0053] Construct a knowledge point and dependency relationship graph, update the students' knowledge mastery status in real time, and based on the students' performance, propagate knowledge features according to the graph convolutional network, and adjust the learning path and knowledge mastery status;
[0054] Define the state space and action space, optimize the personalized learning path according to immediate feedback and reinforcement learning, and adjust tasks in real time;
[0055] Analyze the impact of teaching behaviors on learning effects based on causal inference, establish a causal relationship model to quantify the intervention effect, and adjust teaching strategies.
[0056] Step 1: Perform multimodal data integration and feature modeling. Obtain multimodal data (behavior, emotion, physiology, etc.), and transform these heterogeneous data into available information through a unified feature modeling and fusion process. The specific steps are as follows:
[0057] Collect behavioral data , record the operation behaviors of students in the information-based teaching system, including click stream logs, task completion duration, answer accuracy rate, learning frequency, etc.;
[0058] Collect emotional data , use sentiment analysis technology to capture facial expressions and postures through a camera, collect voice emotional features through a microphone, and extract the learning emotions of students;
[0059] Collect physiological data , obtain the physiological signals of students through wearable devices (such as smart bracelets), including heart rate changes, skin conductance, etc., for judging the stress and concentration levels of the learning state;
[0060] Behavioral data is the core data directly reflecting students' learning behaviors and efficiency; emotional data is used to capture the emotional fluctuations of students during the learning process, such as fatigue, anxiety, etc., by analyzing the emotional features of students' expressions and voices; physiological data provides the physiological state of students, such as heart rate changes and stress indices, through wearable devices, indirectly reflecting the concentration level and learning load of students;
[0061] The learning state of students is multi-dimensional, and single-modal data may be one-sided or misleading. For example, although behavioral data can record the number of tasks completed by students, it cannot judge their true learning engagement; while emotional and physiological data can make up for this defect. Therefore, through multimodal data collection, the learning state of students can be comprehensively described, providing an accurate information basis for teaching optimization;
[0062] For data with different time granularities, the time window slicing method is used for alignment. Assume that the sampling frequency of behavioral data is , the sampling frequency of emotional data is , the sampling frequency of physiological data is , the aligned time window can be expressed as: , where lcm represents the least common multiple operation, mapping all data to the time window , and fill in the missing values to unify the time dimension;
[0063] Normalization uses a logarithmic scaling function to avoid the sensitivity of traditional linear normalization to extreme values. The normalization formula is as follows: , where is the original data sample, is the result after normalization, max(X) is the maximum value of the feature dimension, X is the data set, expressed as .
[0064] Data at different time scales may lead to data loss or redundancy. For example, the sampling frequency of emotion data is relatively high, while the sampling of behavior data is relatively low. If not aligned, it will cause the behavior data in the time window to be too sparse, affecting the timeliness of analysis. Similarly, different data modalities (such as heart rate and answering time) have different dimensions, and direct analysis may produce biases.
[0065] Use self-supervised learning techniques to extract high-dimensional features from different data modalities, and use different advanced algorithms to extract features from the collected multi-modal data respectively:
[0066] For the extraction of behavior data features, use TransformerEncoder to extract complex patterns in time series, such as the correlation between task completion time and click behavior;
[0067] For the extraction of emotion data features, extract emotional features such as attention and stress changes from facial expressions and speech signals through a convolutional neural network (CNN);
[0068] For the extraction of physiological data features, use a variational autoencoder (VAE) to extract the latent feature vectors of physiological signals and capture the focused and fatigued states during the learning process;
[0069] For the extracted high-dimensional features, use a non-linear dimensionality reduction method (such as t-SNE or UMAP) for dimensionality reduction, retain the local manifold structure of the features, and at the same time retain the local manifold structure of the data to generate a dimensionality-reduced feature matrix: , where, is the dimensionality-reduced feature dimension, L is the number of input samples of the data set, and R represents the dimensionality of the dimensionality-reduced feature matrix;
[0070] High-dimensional features may have collinearity and noise. Direct modeling will affect the computational efficiency and prediction performance of the model. Dimensionality reduction can remove irrelevant information and reduce data complexity. For example, some features in behavior data (such as irrelevant click behaviors) have limited predictive significance for learning status. The dimensionality reduction method retains the core information and removes redundant features through non-linear mapping.
[0071] Feature fusion and embedding modeling are carried out. Since the features of different modalities are complementary, multi-head attention mechanism (MHA) is needed for feature fusion. MHA calculates the correlation weights between modalities to achieve dynamic interaction and integration of cross-modal features. The specific steps are as follows:
[0072] Construct the extracted features into feature matrices, namely physiological data feature matrix, emotional data feature matrix, and behavioral data feature matrix. Given the feature matrices of three modalities 、 and , which respectively represent the learning features of students in three different modalities, such as physiological data, emotional data, and behavioral data. The fusion process applies the MHA mechanism multiple times, processes the interaction between modalities in each operation, and retains the important information of each modality: , where the fusion of the first pair of modality features represents the interaction features extracted from physiological data ( ) and emotional data ( ); is the fusion of the second pair of modality features, representing the interaction features extracted from emotional data ( ) and behavioral data ( ); is the fusion of the third pair of modality features, representing the interaction features extracted from behavioral data ( ) and physiological data ( );
[0073] The fusion of modalities can effectively capture the similarities and differences between modalities through the interaction of queries, keys, and values;
[0074] Input the fused feature matrix into the graph embedding model (GraphEmbeddingModel) to generate the global embedding vector E of the student's learning state, thereby representing the comprehensive learning state of the student based on multi-modal features throughout the learning process: , where is the feature matrix after MHA fusion, containing information from multiple modalities; G is the knowledge graph structure, representing the relationship between different knowledge points in the learning task, which can be a graph structure based on subject knowledge, such as the hierarchical relationship between mathematical knowledge points; N is the number of students, that is, the number of students in the dataset; is the dimension of the embedding vector, that is, the representation of each student's learning state in the low-dimensional space;
[0075] The goal of the graph embedding model is to embed the multi-modal features of students into a low-dimensional space so that the embedded features can effectively represent the students' learning state and knowledge mastery. The calculation process of the embedding is as follows:
[0076] Initialize the feature matrix. The initial features of each student are composed of the fused modal features, representing the learning status of the student in different dimensions;
[0077] Construct the adjacency matrix. The adjacency matrix A of the knowledge graph is used to represent the relationships between knowledge points or the relationships between students;
[0078] Perform feature propagation to update the student features through graph convolutional operations (GCN) or other graph neural network techniques: , where represents the feature matrix of the k-th layer, W is the learned weight matrix, is the activation function;
[0079] After multiple layers of graph convolution, the final embedding matrix E is the comprehensive learning status of each student.
[0080] Graph embedding can utilize the relationships between nodes in the knowledge graph to capture the mutual influence of students in the process of knowledge mastery and learning, thereby generating a low-dimensional embedding vector for each student. This vector contains the overall learning status of the student, not only reflecting their performance in a single modality but also integrating the interaction information between different modalities and knowledge points.
[0081] Through data collection and preprocessing, feature extraction and dimensionality reduction, feature fusion and embedding modeling, a complete process from raw data to the representation of students' learning status is constructed to solve the heterogeneous problems of multimodal data in terms of time, dimension, and information interaction.
[0082] Step 2: Perform dynamic knowledge graph construction and real-time update. The dynamic knowledge graph construction and real-time update are to dynamically capture the knowledge point mastery status of students through their learning data and model the relationships between knowledge points as a graph structure. The specific steps are as follows:
[0083] Initialize the knowledge points and their associated relationships, and extract the knowledge points and the dependency relationships between them from the teaching resources;
[0084] Teaching resources refer to various materials and tools that support teaching activities, such as textbooks, courseware, exercise question banks, teaching syllabuses, etc.; knowledge points refer to the specific content or concepts that need to be learned in a course, which can be a single concept, principle, theorem, etc. For example, "algebraic equations" in mathematics;
[0085] Define the nodes. The knowledge point is the basic unit of the graph, and each node represents a specific teaching knowledge point M. For example, "solving quadratic equations of one variable" is a mathematics knowledge point, and the acquisition of the node set comes from the structured analysis of course textbooks, question banks, and teaching syllabuses;
[0086] Define the edges and calculate their weights. The edges represent the association relationships between knowledge points, and the co-occurrence probability is used to mine the forward and backward dependence relationships between knowledge points from the question bank or textbooks. The formula is: , where is the knowledge point and are the number of times they appear simultaneously in the same question or chapter. The co-occurrence probability represents the association strength between knowledge points and serves as the weight of the edge. is the number of times the knowledge point appears, and is the number of times the knowledge point
[0087] The definition of nodes reflects the integrity of the knowledge point set, while the weights of the edges characterize the association strength and the forward and backward order of knowledge points.
[0088] Model the student's knowledge mastery status. Calculate the mastery status of each knowledge point based on the student's learning behavior data , which is used to initialize and update the node attributes in the knowledge graph in real time;
[0089] Initialize the knowledge point status. The initial status Based on the matching degree calculation between the student characteristics embedding and the knowledge points, use the dot product method to measure the matching degree between the student embedding vector and the knowledge point embedding vector : , where is the learnable mapping matrix, is the bias term, T represents the transpose operation, is the Sigmoid function;
[0090] Update the student performance feedback. Dynamically update the knowledge point status according to the student's performance in tests and assignments: , where represents the test score of the student on the knowledge point , represents the dependence strength between knowledge points, is the learning rate;
[0091] Through the matching calculation between the knowledge point embedding and the student status embedding, initialize the student's knowledge mastery level and dynamically adjust the mastery status to reflect the real-time learning results. For example, a high score of the student on "function graph" will improve their mastery status, and vice versa.
[0092] It should be noted that the degree of students' mastery of each knowledge point is dynamic and needs to be modeled by combining historical data and real-time learning behaviors. Through the matching calculation of the student's embedding vector and the knowledge point embedding vector, the current knowledge state of the student can be quantified. In the test feedback, the performance of the student on tasks related to the knowledge point (such as the answering accuracy rate and error distribution) is used to update the mastery state. In the update formula, the dependency between knowledge points is considered: if is a prerequisite knowledge point of and the mastery state of is relatively poor,
[0093] the update amplitude of
[0094] will be restricted. The matching degree of the embedding vector provides the initial state according to the mastery state of the student's knowledge point and changes dynamically with the completion of the learning task, while the dynamic feedback update makes the state better reflect the actual mastery of the student. , where the attention weight is calculated through the correlation strength between the knowledge point and and the current feature representation: , where and can be the student's learning state vector, the feature vector of the knowledge point, or the embedding vector generated by the neural network. Sim is the similarity metric function used to calculate the similarity between the input feature representation (such as the student state vector) and the knowledge point representation, such as dot product, cosine similarity, etc. N is the total number of knowledge points;
[0095] The graph attention network dynamically adjusts the weight of feature propagation by calculating the contribution of neighbor nodes to the current node. For example, if the student has a good mastery of "function extreme value" and "derivative application" is its directly related knowledge point, then the state of "function extreme value" will have a positive impact on the state of "derivative application";
[0096] In the information-based teaching scenario, the feature propagation of knowledge points simulates the "knowledge flow" of teaching content. For example, the process by which students gradually master a group of interrelated knowledge points during the learning process is modeled through feature propagation, enhancing the dynamic adaptability of the knowledge graph;
[0097] Based on the latest learning performance and dissemination characteristics of students, the node attributes and edge weights of the knowledge graph are adjusted in real time. By combining the real-time performance of students with the dissemination characteristics of knowledge points, a comprehensive evaluation of the mastery level of knowledge points is achieved. The dynamic adjustment of edge weights reflects the learning trajectory of students on different knowledge points. For example, if a student repeatedly encounters "function images" and "derivative formulas" in certain questions, the system will automatically increase the edge weight between these two knowledge points, indicating that the system needs to pay more attention to this type of knowledge point combination in the future. The finally updated knowledge graph will be used for subsequent learning path optimization.
[0098] In the construction and update of the dynamic knowledge graph, from the initialization of knowledge points to feature dissemination and then to real-time update, it clearly covers the structured management of teaching content and the dynamic adjustment of students' learning status. This process transforms the static teaching knowledge structure into a dynamic graph that can reflect learning performance in real time, providing support for the optimization of information-based teaching.
[0099] Step 3: Optimize the personalized learning path driven by reinforcement learning. By modeling the association between the dynamic learning status of students and knowledge points, design an adaptive learning path planning mechanism. Using the dynamic knowledge graph as input and combining the real-time learning feedback of students, generate a personalized learning path that can be continuously optimized. The specific steps are as follows:
[0100] The design of the reinforcement learning model requires clarifying the components, including the state space, action space, reward function, and policy update method. Each component is closely related to the scenario. The specific steps are as follows:
[0101] State space: State includes the current learning status of the student and the node features of the dynamic knowledge graph. The learning status of the student is represented by its feature vector The node features of the dynamic knowledge graph are represented by the updated knowledge point features in Step 2 Therefore, the state space is defined as: , where K represents the set of knowledge points on the current learning path;
[0102] Action space: Action represents the recommended learning tasks. The tasks can be knowledge points , related exercise question sets, or a teaching video. The action set is dynamically generated based on the node relationships in the dynamic knowledge graph and the student's knowledge mastery status. For example, if the mastery level of a certain knowledge point is relatively low and there is a strong dependence relationship with the knowledge point , then the knowledge point is preferentially recommended as a learning task;
[0103] The reward function is used to measure the effect of learning tasks. Define the immediate reward To balance the improvement of knowledge mastery and time cost after completing learning tasks: , where and respectively represent the knowledge points after and before the task 's mastery status, T represents the time required to complete the task, and are weight coefficients used to balance learning efficiency and time cost;
[0104] The reinforcement learning model captures the student's current learning situation through the state, provides flexible task selection through the action, and evaluates the effectiveness of the learning task through the reward function, forming a dynamic closed-loop. The state and action are input into the reinforcement learning framework by the dynamic knowledge graph and the student's learning performance.
[0105] For the learning path optimization strategy, in reinforcement learning, the policy defines the probability distribution of selecting action a in state s. The optimization goal of the policy is to maximize the cumulative return to improve the effect of the learning path;
[0106] The policy function is represented by a neural network, with the input being the state , and the output being the action probability distribution . The parameters of the policy network are initialized using the Xavier method to avoid initial gradient explosion or disappearance;
[0107] Use the PPO algorithm (Proximal Policy Optimization) to optimize the policy network. PPO maintains the stability and sample efficiency of policy updates by restricting the magnitude of policy updates. The objective function is: , where represents the expected value of the samples obtained by sampling at different time steps t, is the probability ratio between the current policy and the old policy, is the advantage function, indicating the relative goodness or badness of taking action in state compared to the average behavior. ϵ is a hyperparameter used to limit the magnitude of policy updates;
[0108] Through backpropagation, the PPO algorithm updates the parameters of the policy network to maximize the objective function. The optimization process includes:
[0109] Sample the actions taken by the student in each state from the environment, along with the corresponding rewards and state transition information, calculate the advantage function, and capture the relative advantages and disadvantages of the actions;
[0110] By maximizing the objective function , update the parameters of the policy network and use the Adam optimizer or other optimizers for parameter update.
[0111] Policy optimization maximizes the cumulative reward of the learning effect while maintaining the stability and efficiency of model training. For example, in a teaching scenario, the system dynamically adjusts the recommended learning tasks to maximize the efficiency of students' knowledge acquisition. The optimized policy provides a recommended task sequence for personalized learning paths.
[0112] Based on the optimized policy , execute the recommended learning tasks at each time step and collect students' learning feedback;
[0113] The recommended tasks include task content (such as question sets, teaching videos) and task objectives (such as improving knowledge mastery). After students complete the tasks, the task completion time T, the answering accuracy rate R, and the status changes of relevant knowledge points are recorded in real time ;
[0114] Perform feedback recording and use the students' task completion data as the basis for state update in the next time step. For example, if the mastery of a knowledge point significantly improves after a task, the recommended tasks in the next time step will prioritize advanced knowledge points related to the knowledge point ; the task completion feedback is used to update the knowledge graph and the reinforcement learning state . After each policy execution, the state and the node states of the knowledge graph can also be recalculated. Adjust the priority of the recommended tasks according to real-time data. For example, for nodes that students have not mastered but are highly relevant to the current knowledge point, the system will increase their ranking in the recommended task sequence;
[0115] If a student's performance in a certain task is significantly lower than expected (e.g., too long completion time or too high error rate), the system will trigger a jump mechanism and directly recommend review tasks for relevant basic knowledge points to fill in the knowledge gaps, and use the adjusted learning path sequence as the final output to provide to students and the teaching system.
[0116] By designing the reinforcement learning model, path optimization strategy, task execution and feedback, and path dynamic adjustment, a personalized learning path system that can be optimized in real time is constructed. Combining the students' dynamic learning status and the structured information of the knowledge graph, through the policy update and feedback mechanism, the accuracy and efficiency of teaching task recommendation are achieved.
[0117]
[0118] Step 4: Optimize teaching strategies driven by causal inference. By constructing a causal relationship model, analyze the causal effects between teaching behaviors (such as teaching methods, task allocation) and learning outcomes (such as knowledge point mastery, learning interest), so as to optimize teaching strategies. The specific steps are as follows:
[0119] The construction of the causal relationship model requires clarifying the causal path between teaching behaviors and learning outcomes, and modeling the relationships between these variables through a structured method. The causal variables are defined as follows:
[0120] The independent variable X represents the intervention variable of teaching behavior, such as the frequency of classroom interaction, task allocation strategy, and complexity of teaching content;
[0121] The dependent variable Y represents the learning outcomes of students, such as the degree of knowledge point mastery , the correct answer rate R, and the learning emotion index E;
[0122] The confounding variable Z represents factors that may affect both X and Y simultaneously, such as students' learning ability and the quality of teaching resources;
[0123] Use the Bayesian Network to automatically learn the causal graph structure G from historical teaching data, where nodes represent variables and edges represent causal relationships. The learning process determines the network structure using the method of maximizing the log-likelihood estimate: , where, represents the set of parent nodes of node , and N is the total number of nodes;
[0124] The causal relationship model models teaching behaviors, learning outcomes, and potential confounding factors as a directed acyclic graph, clarifying the potential impact paths of different teaching strategies on learning effects. For example, the frequency of classroom interaction (X) indirectly affects the degree of knowledge point mastery by enhancing students' participation (Z);
[0125] The constructed causal graph G is input into the causal inference stage to quantitatively analyze the intervention effects of teaching strategies.
[0126] Conduct quantitative analysis of intervention effects. Through causal inference techniques, quantitatively analyze the intervention effects of teaching behaviors, so as to evaluate the contributions of different strategies to learning outcomes;
[0127] Define the causal effect. Let represent the implementation of a certain teaching strategy, represent the non-implementation of this strategy. The intervention effect is defined as: , where, represents the external forced intervention on , rather than the observed natural state;
[0128] For example, assume that there are two settings for the classroom question frequency (X): a lower frequency ( ), and a higher frequency ( ). By inferring and calculating the change in the correct answer rate (Y) of students after the increase in the question frequency through intervention, the specific impact of questioning on students' learning effects can be obtained;
[0129] Using the Do-Calculus rule, isolate the estimation path of the intervention effect from the causal graph G. For example, if the path in the causal graph is from X to Z and then to Y, the intervention effect can be decomposed into: , where P(Y∣Z = z, X = x) represents the probability of Y under the conditions of Z = z and X = x;
[0130] Decompose the intervention effect into a direct effect and an indirect effect. The direct effect is from X to Y, such as classroom questioning directly stimulating students' thinking ability;
[0131] The indirect effect is from X to Z and then to Y, such as indirectly improving the learning effect (Y) by increasing students' concentration (Z);
[0132] The quantitative analysis of the intervention effect clarifies the advantages and disadvantages of different teaching strategies through mathematical forms. For example, if a certain strategy does not significantly improve the learning effect, it should be optimized or replaced.
[0133] According to the analysis results of the intervention effect, optimize the teaching strategy and dynamically adjust the teaching behavior. The strategy optimization rule is:
[0134] If (effect threshold), it indicates a significant improvement in the effect, retain and promote this strategy. If , it indicates that the effect is not significant or negative, replace or adjust the strategy. If the analysis finds that a high interaction frequency significantly improves the mastery level of the knowledge point "judgment of function monotonicity" ( ), but has no obvious effect on "drawing of function images" ( ), then reduce the interaction intensity in the teaching of the latter knowledge point;
[0135] During the teaching process, update the causal model according to real-time data. For example, for students with better foundations, reduce the proportion of low-order tasks; for students with weaknesses, increase the review and practice of basic knowledge points.
[0136] Strategy optimization and dynamic adjustment ensure the flexibility and adaptability of teaching behavior, implement differentiated strategies for different groups of students, make the allocation of teaching resources more targeted, and the optimized strategy is used to generate a new teaching plan and dynamically update it into the teaching system.
[0137] It should be noted that the threshold information related in this embodiment is set in advance by professionals and will not be explained in detail here. In the embodiment, there are cases where some parameter English letters are the same, but different meanings are explained during use, and they will not be explained one by one here.
[0138] Through multi-modal data collection and deep learning technology, the present invention accurately captures students' behaviors, emotions and physiological states, realizes data fusion through the multi-head attention mechanism, generates accurate learning state features, and based on the constructed dynamic knowledge graph and graph convolutional network, updates the students' knowledge mastery status in real time and adjusts the learning path to ensure personalized learning support. Then, reinforcement learning is used to optimize the personalized learning path, and the learning task is adjusted in real time in combination with instant feedback to improve learning efficiency and experience. At the same time, through causal inference analysis of the impact of teaching behaviors on learning effects, the intervention effect is quantified, and teaching strategies are scientifically adjusted to improve teaching effects. This solution combines multi-modal fusion, graph convolutional network and reinforcement learning to optimize the learning path and teaching strategies, and has significant advantages in personalized learning support, teaching effect improvement and learning efficiency optimization.
[0139] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0140] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.
[0141] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0142] In addition, the functional modules in each embodiment of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0143] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0144] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A teaching optimization method based on big data informationization, characterized by: The steps include: Collect students' behavior, emotions and physiological data, perform synchronous alignment and normalization, extract features and reduce dimensions based on deep learning, and fuse features of each modality through a multi-head attention mechanism; Construct a graph of knowledge points and dependencies, update students’ knowledge mastery status in real time, and adjust learning paths and knowledge mastery status based on student performance and knowledge propagation characteristics of graph convolutional networks; Define state space and action space, optimize personalized learning paths based on immediate feedback and reinforcement learning, and adjust tasks in real time; Analyze the impact of teaching behavior on learning outcomes based on causal inference, establish a causal relationship model to quantify the intervention effect, and adjust the teaching strategy; Construct a knowledge point and dependency graph to update students’ knowledge mastery status in real time. The specific steps are as follows: Initialize knowledge points and associations, extract knowledge points and dependencies between knowledge points from teaching resources; Define nodes, knowledge points It is the basic unit of the graph, each node Indicates specific teaching knowledge points; Define the edge and calculate the weight. The edge represents the relationship between knowledge points. Use the co-occurrence probability to mine the previous and next dependencies between knowledge points from teaching resources. The formula is: ,in, It is a knowledge point and The number of times it appears in the same question at the same time, It is a knowledge point The number of occurrences, It is a knowledge point Number of occurrences, co-occurrence probability Indicates the strength of association between knowledge points and serves as the weight of the edge; Initialize the knowledge point state, initial state Based on the matching calculation between student feature embedding and knowledge points, the dot product method is used to measure the student embedding vector and knowledge point embedding vector The matching degree: ,in, is the mapping matrix used for learning, is the bias term, T represents the transposition operation, is the Sigmoid function; Provide student performance feedback updates and dynamically update knowledge point status based on student performance in tests and assignments : ,in, Indicates that students are at the knowledge point The test scores of Representing knowledge points , The strength of the dependence between is the learning rate.
2. The method for optimizing teaching based on big data informationization according to claim 1, characterized in that: Collect student behavior, emotion, and physiological data, perform synchronous alignment and normalization, extract features and reduce dimensions based on deep learning. The specific steps are as follows: The collected behavioral data include clickstream logs, task completion time, answer accuracy, and learning frequency; Collect emotional data, capture facial expressions and gestures through cameras, and collect speech through microphones; Collect physiological data and obtain students' physiological signals through wearable devices. Physiological signals include heart rate changes and skin conductance. For the collected data of different time granularities, the time window slicing method is used for alignment and the logarithmic scaling function is used for normalization; Perform behavioral data feature extraction to extract complex patterns in time series, including the correlation between task completion time and click behavior; Extract emotional data features and extract emotional features from facial expressions and voice signals through convolutional neural networks. Emotional features include changes in attention and stress. Perform physiological data feature extraction and use variational autoencoder to extract feature vectors of physiological signals; For the extracted features, nonlinear dimensionality reduction method is used for dimensionality reduction.
3. The method for optimizing teaching based on big data informationization according to claim 2 is characterized in that: The multi-head attention mechanism is used to fuse the features of each modality. The specific steps include: The collected data features are constructed into physiological data feature matrix, emotional data feature matrix, and behavioral data feature matrix, and the feature matrices are fused through a multi-head attention mechanism; The fused feature matrix is input into the graph embedding model to generate a global embedding vector of the student’s learning status.
4. The method for optimizing teaching based on big data informationization according to claim 3 is characterized in that: Based on the student performance and the knowledge propagation characteristics of the graph convolutional network, the learning path and knowledge mastery status are adjusted, including the following steps: Use the graph attention network to propagate the knowledge point features and update the feature representation of each knowledge point: , where the attention weight Through knowledge points and The association strength and current feature representation calculation: ,in, and is the student's learning state vector and the feature vector of the knowledge point, sim is the similarity measurement function, and N is the total number of knowledge points; Represents the feature vector of neighbor node j in the tth layer; is the learning weight matrix in the t-th layer graph attention network, which is used to calculate the feature vector of neighbor node j Perform linear transformation; Represents the feature vector obtained by the knowledge point at the t+1th layer; Based on students' latest learning performance and communication characteristics, the node attributes and edge weights of the knowledge graph are adjusted in real time.
5. The method for optimizing teaching based on big data informationization according to claim 4 is characterized in that: Define the state space and action space, optimize the personalized learning path based on immediate feedback and reinforcement learning, and adjust the task in real time, including the following steps: Define the state space. The state of the state space includes the student's current learning state and the node features of the dynamic knowledge graph. The student's learning state is represented by a feature vector, and the node features of the dynamic knowledge graph are represented by the updated knowledge point features. Define an action space, where the actions in the action space represent recommended learning tasks. The action set is dynamically generated based on the node relationships in the dynamic knowledge graph and the student's knowledge mastery status. Define the immediate reward as the trade-off between the improvement in knowledge mastery and the time cost after completing the learning task as the reward function; The state captures the student's current learning situation, the action provides task selection, and the reward function evaluates the learning task; Input the state into the policy function, output the action probability distribution, and use the PPO algorithm to optimize the policy function; Based on the optimized strategy, the recommended learning task is executed at each time step, and the students' learning feedback is collected and the task is adjusted according to the learning feedback.
6. The method for optimizing teaching based on big data informationization according to claim 5 is characterized by: Based on causal inference, we analyze the impact of teaching behavior on learning outcomes, establish a causal relationship model to quantify the intervention effect, and adjust the teaching strategy, including the following steps: The causal variables were defined, and the independent variables represented the intervening variables of teaching behaviors, including the frequency of classroom interactions, task allocation strategies, and the complexity of teaching content; The dependent variables represent the students’ learning outcomes, including the mastery of knowledge points and the accuracy of answering questions; Confounding variables represent factors that affect both the independent and dependent variables; Use Bayesian networks to automatically learn causal graph structures from historical teaching data. Causal graph nodes represent variables, and edges represent causal relationships. The learning process determines the network structure based on the method of maximizing log-likelihood estimation. The constructed causal diagram is input into the causal inference stage to analyze the intervention effect of the teaching strategy and define the causal effect; According to the Do-Calculus rule, the estimated path of the intervention effect is separated from the causal diagram; Based on the analysis results of the intervention effects, optimize teaching strategies and dynamically adjust teaching behaviors.
Citation Information
Patent Citations
Accurate teaching management method and system based on adaptive learning analysis
CN118396804A
Personalized training scheme generation method based on big data analysis
CN118761874A