Online teaching optimization method and system based on emotion recognition
Through multimodal sensors and advanced algorithms, learners' emotional data are collected in real time, and online teaching content is dynamically adjusted, which solves the problem of insufficient personalization in traditional teaching and improves learning effect and experience.
Patent Information
- Application Number
- CN202510416360.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional online teaching model lacks personalization and cannot understand learners' emotions and cognitive status in real time, resulting in disconnection between teaching content and learners' needs and poor learning results and experience.
Learner emotional data is collected through multimodal sensors, graph convolutional neural network is used to perform feature fusion, combine meta-reinforcement learning model to optimize teaching strategies, dynamically generate course content, and personalized optimization is achieved through hierarchical teaching control models.
It realizes dynamic adjustment of teaching content based on learners' real-time emotions and cognitive status, improves learning efficiency and enthusiasm, and improves learning experience and evaluation accuracy.
Smart Images

Figure CN120355539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of online teaching management, and particularly to an online teaching optimization method and system based on emotion recognition. Background Art
[0002] In today's digital age, online teaching has become an important part of the education field by virtue of its advantages such as flexibility, convenience, and resource sharing. It breaks the limitations of time and space, enabling learners to access rich learning resources anytime and anywhere, providing new opportunities for the popularization and personalized development of education. However, current online teaching still faces many challenges, severely restricting the improvement of teaching effects and learning experiences.
[0003] Most traditional online teaching models adopt a "one-size-fits-all" approach, providing the same teaching content and teaching progress for all learners, ignoring individual differences among learners in terms of cognitive level, learning style, emotional state, etc. Each learner has a different speed of accepting knowledge and understanding ability. The unified teaching rhythm makes it difficult for some learners to keep up with the teaching progress, resulting in frustrated learning enthusiasm; for learners with stronger learning abilities, the teaching content may lack challenges and cannot fully stimulate their learning potential. For example, on some online course platforms, regardless of the basic knowledge and learning ability of learners, teaching is carried out in accordance with a fixed course sequence and teaching duration, making it difficult for learners to obtain a learning experience that meets their own needs. In face-to-face teaching, teachers can understand the learning status of learners in real time, such as whether they understand the teaching content and whether they are concentrated, by observing the facial expressions, body languages, classroom responses, etc. of learners, and adjust teaching strategies in a timely manner. However, in the online teaching environment, there is a time and space separation between teachers and learners, making it difficult to directly observe the real-time status of learners. Although some online teaching platforms provide simple interaction functions, such as chat windows, question buttons, etc., the information obtained through these methods is limited and cannot comprehensively reflect the emotional and cognitive states of learners. This makes it difficult for teachers to timely discover the problems encountered by learners during the learning process and unable to adjust teaching methods and progress in a timely manner, affecting the pertinence and effectiveness of teaching.
[0004] Existing online teaching interactions mainly focus on simple text interaction methods such as Q&A and discussion forum exchanges, lacking diverse interaction means. This single interaction mode cannot meet the diverse learning needs of learners and is also difficult to create a real and vivid learning atmosphere. For example, for some teaching contents that require intuitive demonstrations and practical operations, it is difficult for learners to deeply understand and master them only through text interaction. In addition, different learners have different preferences for interaction methods. The single interaction mode cannot adapt to the characteristics of each learner, reducing the participation and learning interest of learners. The content of online teaching courses is usually determined before class, lacking a mechanism for dynamic adjustment according to the real-time learning situation of learners. As the learning process progresses, the knowledge mastery level, learning interest, and emotional state of learners will change, but the course content cannot be changed accordingly in a timely manner. This results in the disconnection between the teaching content and the actual needs of learners, unable to achieve the best teaching effect. For example, when learners have difficulty understanding a certain knowledge point and their learning enthusiasm decreases, the course content cannot automatically adjust the difficulty or provide additional auxiliary learning resources, making it possible for learners to encounter more difficulties in subsequent learning. The current evaluation of the learning effect of online teaching mainly relies on methods such as homework and exams. This evaluation method focuses on the examination of knowledge mastery and ignores the impact of factors such as the learning process, learning attitude, and emotional changes of learners on the learning effect. Moreover, these evaluation methods are often phased and cannot reflect the learning status and progress of learners in real time. For example, a learner may have low learning efficiency due to emotional fluctuations during the learning process, but may achieve good results in the exam due to temporary review, which cannot truly reflect the problems in their learning process. This incomplete and inaccurate evaluation method is not conducive to teachers discovering problems in the teaching process in a timely manner, nor is it conducive to learners understanding their own learning status and adjusting their learning strategies.
[0005] To solve the above problems, researchers have tried to apply various technologies to online teaching. For example, data analysis technology is used to analyze the learning behavior data of learners to understand their learning habits and needs. However, this method can only obtain the external behavior data of learners and cannot deeply understand their internal emotional and cognitive states. The development of artificial intelligence technology has brought new hope to online teaching. Some intelligent teaching systems have begun to try to use machine learning algorithms to achieve the recommendation of teaching content and the optimization of teaching strategies. However, most of these systems do not fully consider the emotional factors of learners, and the accuracy and adaptability of the algorithms still need to be improved. Summary of the Invention
[0006] The purpose of the present invention is to provide an online teaching optimization method and system based on emotion recognition to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions: An online teaching optimization method based on emotion recognition, the method comprising:
[0008] Collecting real-time emotion data of learners through multi-modal sensors, the multi-modal sensors including a facial expression collection module, a speech emotion analysis module, an eye movement tracking module, and a physiological signal detection module; performing multi-modal feature fusion on the real-time emotion data based on a graph convolutional neural network to generate an emotion state feature vector; inputting the emotion state feature vector into a meta-reinforcement learning model, the meta-reinforcement learning model adopting a double-layer policy network structure, and optimizing teaching strategies across tasks based on a dynamic reward mechanism to generate teaching strategy parameters;
[0009] Constructing a dynamic course generation model according to the teaching strategy parameters, the dynamic course generation model taking maximizing knowledge coverage and balancing cognitive load as optimization objectives, and using an improved genetic algorithm to optimize the sequence of teaching content, wherein the improved genetic algorithm introduces an adaptive crossover probability and a mutation probability, and constrains the solution space through a knowledge graph; outputting optimal teaching sequence data based on the dynamic course generation model;
[0010] Constructing a hierarchical teaching control model according to the optimal teaching sequence data, the hierarchical teaching control model including a content planning layer, an interaction feedback layer, and an evaluation and adjustment layer, wherein the content planning layer performs global knowledge path planning based on the teaching strategy parameters, the interaction feedback layer adjusts the teaching interaction mode in real time based on the optimal teaching sequence data, and the evaluation and adjustment layer dynamically evaluates the learning state based on a hidden Markov model; generating teaching control instructions through the hierarchical teaching control model to achieve personalized online teaching optimization.
[0011] Preferably, the performing multi-modal feature fusion on the real-time emotion data based on a graph convolutional neural network to generate an emotion state feature vector includes:
[0012] Constructing a multi-modal emotion graph structure, taking facial expression key points, speech spectrum features, eye movement trajectory coordinates, and physiological signal waveforms as graph nodes, and constructing an edge weight matrix based on the inter-modal correlation; using a graph attention mechanism to dynamically aggregate node features, and calculating the weight distribution between nodes through a multi-head attention layer;
[0013] Designing a graph convolutional residual module, fusing shallow local features and deep global features through skip connections, and outputting a multi-modal joint embedding vector; performing temporal alignment processing on the joint embedding vector, and using a long short-term memory network to capture the time dependence of the emotion state to generate a temporal emotion feature sequence;
[0014] Introducing an adversarial training mechanism, distinguishing real emotion labels from generated feature distributions through a discriminator, and optimizing the generalization ability of the graph convolutional neural network.
[0015] Preferably, the meta-reinforcement learning model adopts a two-layer policy network structure, including:
[0016] Construct a meta-policy network and a base policy network, where the meta-policy network outputs meta-parameters shared by tasks, and the base policy network generates specific teaching actions based on the meta-parameters; design a dynamic reward function, which includes a knowledge mastery reward term, an emotional stability reward term, and an interaction activity reward term. Among them, the knowledge mastery reward term is calculated through a knowledge point association matrix and the answer correct rate, the emotional stability reward term is measured by an emotional entropy value, and the interaction activity reward term is generated based on a composite function of the question asking frequency and the response delay;
[0017] Adopt a curriculum progressive training method to update network parameters in stages from simple teaching scenarios to complex scenarios, and dynamically adjust the training task distribution through a task sampler; construct a policy distillation loss function to perform KL divergence constraint on the parameter distribution output by the meta-policy network and the action distribution of the base policy network to achieve policy consistency optimization.
[0018] Preferably, the improved genetic algorithm for sequence optimization of teaching content includes:
[0019] Based on the knowledge graph, construct a chromosome coding rule to map the teaching unit sequence to a gene chain. Each gene contains a knowledge point identifier, a difficulty coefficient, and an association weight; design a multi-objective fitness function, which includes a knowledge coherence measure, a cognitive load balance degree, and an emotional matching score. Among them, the emotional matching score is calculated through the cosine similarity between the emotional state feature vector and the knowledge point emotional label;
[0020] Introduce an adaptive genetic operator to dynamically adjust the crossover probability and mutation probability according to the population diversity index. Among them, the crossover probability decays exponentially as the population similarity increases, and the mutation probability increases in a piecewise linear manner as the number of iterations increases; adopt an elite retention strategy to directly retain the optimal solution of each generation to the next generation, and constrain the mutation operation through the knowledge graph to ensure that the new solution conforms to the prior knowledge dependency relationship.
[0021] Preferably, the content planning layer performs global knowledge path planning, including:
[0022] Construct a knowledge topology graph, use knowledge points as nodes and dependency relationships as directed edges, and calculate the knowledge point weights through the PageRank algorithm; design a path generation model, represent the learning path as an implicit Markov chain, and solve the maximum probability path based on the Viterbi algorithm;
[0023] Introduce a dynamic weight adjustment mechanism to update the knowledge point priorities according to the real-time emotional state feature vector, and balance the contribution degrees of historical learning data and the current emotional state through a weight decay factor; use Monte Carlo tree search to expand and simulate candidate paths, and select the path with the highest cumulative reward as the global planning result.
[0024] Preferably, the real-time adjustment of the teaching interaction mode by the interaction feedback layer includes:
[0025] Construct a multi-channel interaction decision-making model, which receives the emotional state feature vector, the attributes of the current teaching unit, and the environmental context data, and outputs the interaction mode encoding; the interaction mode encoding includes the voice broadcast rate, the interface element layout strategy, and the Q&A difficulty gradient; use a fuzzy logic controller to finely adjust the interaction parameters, define the membership function of the input variables and the fuzzy rule base, and obtain the precise control quantity through the centroid method for defuzzification; design a real-time feedback loop to dynamically update the fuzzy rule weights based on the learner's operation delay and the attention focus area.
[0026] Preferably, the dynamic evaluation of the learning state by the evaluation and adjustment layer based on the hidden Markov model includes:
[0027] Define a set of hidden states, including the cognitive level state, the emotional fluctuation state, and the metacognitive ability state; construct a set of observation states, including the answer correctness rate, the interaction response time, and the physiological signal statistics.
[0028] Preferably, the specific implementation of the knowledge graph to constrain the solution space includes:
[0029] Construct a logical constraint expression to encode the prerequisite conditions, exclusion relationships, and optimal learning intervals between knowledge points into first-order predicate logic; use a constraint satisfaction problem solver to verify the feasibility of the candidate solutions generated by the genetic algorithm, and eliminate the individuals that violate the logical constraints; design a repair operator to locally adjust the chromosomes that partially violate the constraints, and select the gene loci with the least conflicts through the greedy algorithm for replacement to ensure that the new solutions meet the topological constraints of the knowledge graph.
[0030] Preferably, the specific implementation of the multi-channel interaction decision-making model includes:
[0031] Adopt a multi-task learning framework to share the underlying feature extraction network and parallelly output voice, visual, and text interaction strategies; design a collaborative loss function between channels to constrain the consistency of the decision results of different channels through contrastive learning;
[0032] Introduce a memory enhancement module to store the historical interaction modes and their effect evaluations, and retrieve the best strategies for similar scenarios through the attention mechanism to achieve experience transfer optimization.
[0033] Preferably, the present invention further includes an online teaching optimization system based on emotion recognition to implement the above-mentioned online teaching optimization method. The system includes:
[0034] A data acquisition module for obtaining the facial expressions, voice signals, eye movement trajectories, and physiological signals of learners through multimodal sensors;
[0035] A feature fusion module for jointly representing and learning multimodal data based on a graph convolutional neural network to generate an emotion state feature vector;
[0036] A strategy generation module for outputting cross-task teaching strategy parameters using a meta-reinforcement learning model and driving a dynamic curriculum generation model to optimize the teaching content sequence;
[0037] A hierarchical control module including a content planning unit, an interaction feedback unit, and an evaluation and adjustment unit, which respectively implement knowledge path planning, real-time interaction adjustment, and learning state evaluation;
[0038] An execution and output module for converting the optimized teaching instructions into multimodal outputs such as voice, interface elements, and Q&A content.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] The online teaching optimization method and system based on emotion recognition proposed by the present invention effectively solve the pain points of traditional online teaching and produce significant positive effects in many aspects.
[0041] By collecting the real-time emotion data of learners through multimodal sensors, fusing and processing it through a graph convolutional neural network, a feature vector that can accurately reflect their emotion state is generated. The meta-reinforcement learning model optimizes the teaching strategy based on this vector, and the dynamic curriculum generation model adjusts the teaching content sequence accordingly. The hierarchical teaching control model plans personalized knowledge paths. In this way, learners with different learning abilities and emotion states can obtain suitable learning content and progress, and the learning efficiency and enthusiasm are significantly improved. With the help of multimodal data collection and advanced algorithms, the system can deeply and real-time understand the emotion and cognitive states of learners. The evaluation and adjustment layer combines various observation data and uses a hidden Markov model to dynamically evaluate the learner state, which can timely detect problems in the learning process and provide a reliable basis for teaching strategy adjustment.
[0042] The multi-channel interaction decision-making model of the interaction feedback layer synthesizes multi-faceted data to output interaction mode coding, and the fuzzy logic controller and real-time feedback loop further optimize the interaction parameters. This enables teaching interaction to be flexibly adjusted according to the real-time state of learners, enhancing learner engagement and learning experience, and breaking away from the limitations of traditional single interaction modes. The dynamic curriculum generation model is based on a knowledge graph and an improved genetic algorithm, and optimizes the teaching content sequence with the goal of knowledge coverage and cognitive load balance. The knowledge graph restricts the solution space, ensuring the reasonable coherence of teaching content, and enabling the curriculum content to be adjusted in real time according to the learning situation and emotional state of learners, enhancing the pertinence of teaching.
[0043] The evaluation and adjustment layer changes the traditional single evaluation method and evaluates the cognitive, emotional, and metacognitive abilities of learners by synthesizing multiple factors. This comprehensive evaluation can more accurately reflect the actual learning situation, provide a precise direction for subsequent tutoring and training, and help learners improve their comprehensive learning abilities. Brief Description of the Drawings
[0044] Figure 1 is the working principle diagram of the online teaching optimization method based on emotion recognition described in the present invention;
[0045] Figure 2 is the working principle diagram of the meta-reinforcement learning model adopting a double-layer policy network structure;
[0046] Figure 3 is the working flowchart of the solution space constraint of the genetic algorithm. Detailed Embodiments
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Please refer to Figures 1 - 3 , the present invention provides a technical solution: an online teaching optimization method based on emotion recognition, and the method includes:
[0049] Collect real-time emotion data of learners using multi-modal sensors. Among them, the facial expression acquisition module can use a camera to capture the facial expressions of learners and obtain information on facial expression key points; the speech emotion analysis module processes the speech signals of learners and extracts speech spectrum features; the eye movement tracking module records the eye movement trajectory coordinates with an eye movement tracking device; the physiological signal detection module collects physiological signal waveforms such as heart rate and skin conductance response through wearable devices. These data in different modalities reflect the emotional state of learners from multiple dimensions.
[0050] Based on the graph convolutional neural network, multi-modal feature fusion is performed on the collected real-time emotion data. Facial expression key points, speech spectrum features, eye movement trajectory coordinates, and physiological signal waveforms are used as graph nodes, and an edge weight matrix is constructed according to the inter-modal correlation. The graph attention mechanism is used to dynamically aggregate node features, and the weight distribution between nodes is calculated through the multi-head attention layer. A graph convolutional residual module is designed to fuse shallow local features and deep global features through skip connections, and output a multi-modal joint embedding vector. The joint embedding vector is subjected to temporal alignment processing, and the long short-term memory network is used to capture the time dependence of the emotion state, generating a temporal emotion feature sequence. An adversarial training mechanism is introduced, and the discriminator is used to distinguish the real emotion label from the generated feature distribution, optimizing the generalization ability of the graph convolutional neural network, and finally generating an emotion state feature vector.
[0051] The generated emotion state feature vector is input into the meta-reinforcement learning model. This model adopts a two-layer policy network structure, including a meta-policy network and a base policy network. The meta-policy network outputs meta-parameters shared by tasks, and the base policy network generates specific teaching actions based on the meta-parameters. A dynamic reward function is designed, including a knowledge mastery reward term, an emotion stability reward term, and an interaction activity reward term. The curriculum progressive training method is adopted to update the network parameters in stages from simple teaching scenarios to complex scenarios, and the training task distribution is dynamically adjusted through the task sampler. A policy distillation loss function is constructed to perform KL divergence constraint on the parameter distribution output by the meta-policy network and the action distribution of the base policy network, realizing policy consistency optimization, and thus generating teaching policy parameters.
[0052] A dynamic curriculum generation model is constructed according to the generated teaching policy parameters. This model takes the maximization of knowledge coverage and the balance of cognitive load as optimization goals, and uses an improved genetic algorithm to optimize the sequence of teaching content. Based on the knowledge graph, a chromosome encoding rule is constructed, mapping the teaching unit sequence to a gene chain, and each gene contains a knowledge point identifier, a difficulty coefficient, and an association weight. A multi-objective fitness function is designed, including knowledge coherence measurement, cognitive load balance degree, and emotion matching score. An adaptive genetic operator is introduced to dynamically adjust the crossover probability and mutation probability according to the population diversity index. The elite retention strategy is adopted to directly retain the optimal solution of each generation to the next generation, and the mutation operation is restricted by the knowledge graph to ensure that the new solution conforms to the prior knowledge dependence relationship, and finally the optimal teaching sequence data is output.
[0053] Construct a hierarchical teaching control model based on the optimal teaching sequence data. This model includes a content planning layer, an interaction feedback layer, and an evaluation and adjustment layer. The content planning layer constructs a knowledge topology graph, calculates the weights of knowledge points through the PageRank algorithm, designs a path generation model, and uses the Viterbi algorithm to solve the maximum probability path. A dynamic weight adjustment mechanism is introduced to update the priority of knowledge points according to the real-time emotion state feature vector. Monte Carlo tree search is used to expand and simulate candidate paths, and the path with the highest cumulative reward is selected as the global planning result. The interaction feedback layer constructs a multi-channel interaction decision model, receives the emotion state feature vector, the attributes of the current teaching unit, and the environmental context data, and outputs the interaction mode encoding. A fuzzy logic controller is used to finely adjust the interaction parameters, and a real-time feedback loop is designed to dynamically update the fuzzy rule weights based on the learner's operation delay and attention focus area. The evaluation and adjustment layer defines a set of hidden states and a set of observation states, and dynamically evaluates the learning state based on the hidden Markov model. Finally, teaching control instructions are generated through the hierarchical teaching control model to achieve personalized online teaching optimization.
[0054] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0055] Embodiment 1:
[0056] This embodiment mainly describes the specific process of multi-modal feature fusion of real-time emotion data based on a graph convolutional neural network to generate an emotion state feature vector. Its function is to improve the ability to accurately represent the learner's emotion state, specifically including:
[0057] Construct a multi-modal emotion graph structure: Use facial expression key points, speech spectrum features, eye movement trajectory coordinates, and physiological signal waveforms as graph nodes. Assume there are n1 facial expression key points, the dimension of the speech spectrum feature is n2, the dimension of the eye movement trajectory coordinate is n3, and the dimension of the physiological signal waveform is n4. Then the total number of nodes N = n1 + n2 + n3 + n4. Based on the inter-modal correlation, construct an edge weight matrix W. For different modal nodes i and j, the edge weight W ij can be determined by calculating their similarity, for example, using cosine similarity:
[0058]
[0059] where and are the feature vectors corresponding to nodes i and j respectively.
[0060] Adopt the graph attention mechanism: Use the graph attention mechanism to dynamically aggregate node features. Calculate the weight distribution between nodes through a multi-head attention layer. Assume there are K heads. For each head k, calculate the attention weight of node i to node j as:
[0061]
[0062] Among them, W k is the weight matrix, is the attention mechanism parameter vector, is the set of neighbor nodes of node i. The finally aggregated node feature is:
[0063]
[0064] Among them, σ is the activation function.
[0065] Design the graph convolution residual module: Design the graph convolution residual module, which fuses shallow local features and deep global features through skip connections. Assume that the input of the graph convolution layer is The output after convolution operation is Then the output of the graph convolution residual module is:
[0066]
[0067] In this way, the problem of gradient disappearance can be effectively avoided, while retaining feature information at different levels, and outputting a multi-modal joint embedding vector.
[0068] Temporal alignment and feature sequence generation: Perform temporal alignment processing on the joint embedding vector, and use the long short-term memory network (LSTM) to capture the time dependence of the emotional state. The calculation formula of the LSTM cell is as follows:
[0069]
[0070] Among them, i t 、f t 、o t are the input gate, forget gate, and output gate respectively, is the candidate memory cell, is the memory cell, is the output feature, W is the weight matrix, b is the bias vector, and ⊙ represents element-wise multiplication. After processing by LSTM, a temporal emotional feature sequence is generated.
[0071] Adversarial training mechanism: Introduce the adversarial training mechanism, and use the discriminator to distinguish between real emotion labels and generated feature distributions. The goal of the discriminator D is to maximize the ability to distinguish real samples and generated samples, and the goal of the generator (graph convolutional neural network) G is to minimize the probability of being distinguished by the discriminator. The adversarial process between the two can be represented by the following loss function:
[0072]
[0073] Among them, p real is the distribution of real emotion labels, p gen is the distribution of generated features, y is the real sample, and x is the generated sample. Through adversarial training, the generalization ability of the graph convolutional neural network is optimized, and finally a more accurate emotion state feature vector is generated.
[0074] Embodiment 2:
[0075] This embodiment elaborates in detail the construction and training process of the meta-reinforcement learning model, whose role is to improve the adaptability and effectiveness of teaching effects by optimizing teaching strategies to better meet the needs of different learners. The specific methods include:
[0076] Construct a meta-policy network and a base policy network. The meta-policy network receives information such as the emotion state feature vector and outputs the meta-parameters θ meta shared by the tasks. The base policy network generates specific teaching actions a based on the meta-parameters θ meta Assume that the parameters of the meta-policy network are φ and the parameters of the base policy network are ω. Then the output of the meta-policy network can be expressed as where is the input state, and the process of the base policy network generating actions is
[0077] Design a dynamic reward function R, including a knowledge mastery reward term R know , an emotional stability reward term R emotion , and an interaction activity reward term R interact . The knowledge mastery reward term is calculated through the knowledge point correlation matrix M and the answer correct rate p. Assume that the number of knowledge points is m, then The emotional stability reward term is measured by the emotional entropy value H, R emotiom =-λ1H, where λ1 is the weight coefficient. The interaction activity reward term is generated based on the composite function of the question asking frequency f and the response delay d. For example, where λ2 is the weight coefficient. The total dynamic reward function is R = R know +R emotion +R interact .
[0078] Adopt a curriculum progressive training method to update the network parameters in stages from simple teaching scenarios to complex scenarios. First, define a series of teaching scenario difficulty levels l1 < l2 < … < l n . At each stage k, tasks are dynamically sampled from the task set T k corresponding to the difficulty level for training. Assume that the task sampled at stage k is T k, when training on this task, update the parameters of the meta-policy network and the base policy network to maximize the cumulative reward on this task. The algorithm formula is as follows:
[0079]
[0080] where α and β are learning rates, s t , a t are the state and action at time step t, and T is the number of time steps of the task.
[0081] Construct a policy distillation loss function, and perform KL divergence constraint on the parameter distribution P(θ meta ) output by the meta-policy network and the action distribution Q(a) of the base policy network to achieve policy consistency optimization.
[0082] By minimizing this loss function, the actions generated by the base policy network are made consistent with the parameter output of the meta-policy network, improving the stability and generalization ability of the model.
[0083] Example 3:
[0084] This example mainly illustrates the specific steps of using an improved genetic algorithm to optimize the sequence of teaching content. Its purpose is to better match the emotional state of learners while satisfying the knowledge coverage and cognitive load balance, and optimize the presentation order of teaching content. The specific steps include:
[0085] Based on the knowledge graph, construct a chromosome encoding rule to map the teaching unit sequence to a gene chain. Assume the number of teaching units is n, and each gene contains a knowledge point identifier id, a difficulty coefficient d, and an association weight w. For example, a chromosome can be represented as [(id1, d1, w1), (id2, d2, w2), …, (id n , d n , w n )]. This encoding method can intuitively represent the order and related attributes of teaching content, facilitating subsequent operations of the genetic algorithm.
[0086] Design a multi-objective fitness function F, including a knowledge coherence metric C, a cognitive load balance B, and an emotion matching score M. The knowledge coherence metric can be determined by calculating the association strength between adjacent knowledge points. Assume the association strength between knowledge points i and i + 1 is r i,i+1 , then The cognitive load balance can be measured by calculating the distribution uniformity of knowledge points with different difficulties. For example where is the average difficulty coefficient, and σ 2 is the variance of the difficulty coefficient. The emotion matching score is obtained through the emotion state feature vector and the knowledge point emotion label Cosine similarity calculation The total fitness function is F = λ3C + λ4B + λ5M, where λ3, λ4, and λ5 are weight coefficients
[0087] Introduce an adaptive genetic operator to dynamically adjust the crossover probability P according to the population diversity index c and the mutation probability P m . The population similarity can be measured by calculating the average distance between individuals in the population. Assume the number of population individuals is N, and the distance between individuals uses the Euclidean distance d ij , then the population similarity The crossover probability decays exponentially as the population similarity increases, that is where is the maximum crossover probability, and γ is the decay coefficient. The mutation probability increases in a piecewise linear manner as the number of iterations t increases. For example, when t < t1 when t1 ≤ t < t2 when t ≥ t2 where are the minimum and maximum mutation probabilities respectively
[0088] Adopt the elite retention strategy to directly retain the optimal solution of each generation to the next generation. During the mutation operation, ensure that the new solution conforms to the prior knowledge dependency relationship through the knowledge graph constraint. For example, for a certain gene locus i, if the mutated knowledge point id' conflicts with the previous knowledge point id i-1 (such as violating the prerequisite relationship in the knowledge graph), then use the greedy algorithm to select the gene locus with the least conflict for replacement, and reselect a knowledge point identifier that conforms to the topological constraint of the knowledge graph to ensure the logical rationality of the new solution generated by the genetic algorithm
[0089] Example 4
[0090] This example focuses on the global knowledge path planning of the content planning layer and the real-time teaching interaction mode adjustment of the interaction feedback layer, aiming to plan an adapted learning path for learners and flexibly adjust the teaching interaction method according to their real-time status to improve the learning experience and learning effectiveness
[0091] The global knowledge path planning of the content planning layer includes
[0092] Construct a knowledge topology graph and calculate the weights of knowledge points: Construct a knowledge topology graph, set knowledge points as nodes, and the dependency relationship between knowledge points as directed edges. After multiple iterations, make the weights of knowledge points tend to be stable, and the knowledge points with higher weights are more important in teaching
[0093] Design path generation model and solving the maximum probability path: Represent the learning path as an implicit Markov chain.
[0094] Dynamic weight adjustment mechanism: Introduce a dynamic weight adjustment mechanism to update the knowledge point priority according to the real-time emotion state feature vector.
[0095] Monte Carlo tree search: Use Monte Carlo tree search to expand and simulate candidate paths. Initialize the search tree and set the current learning state as the root node. In each iteration, select an under-explored node to expand, generate a complete path starting from this node by simulating a random policy, and calculate its cumulative reward. Finally, select the path with the highest cumulative reward as the global planning result.
[0096] The real-time interaction adjustment of the interaction feedback layer includes:
[0097] Construct a multi-channel interaction decision model: Construct a multi-channel interaction decision model that receives the emotion state feature vector, the current teaching unit attributes and environmental context data, and outputs an interaction mode encoding. Adopt a multi-task learning framework, share the underlying feature extraction network, and parallelly output speech, visual, and text interaction strategies. Design a collaborative loss function between channels to constrain the consistency of decision results of different channels through contrastive learning.
[0098] Fuzzy logic controller: Use a fuzzy logic controller to finely adjust the interaction parameters. Taking the speech broadcast rate as an example, first define the membership functions of the input variables (such as the excitement index in the emotion state feature vector, the difficulty of the current teaching content), such as the membership function μ low (x), μ medium (x), μ high (x), and the membership functions μ easy (y), μ medium (y), μ hard (y) of the teaching content difficulty y. Then define the fuzzy rule base, such as "if the excitement is low and the teaching content is easy, then the speech broadcast rate is slow". Finally, solve the fuzzy to obtain the precise control quantity z through the centroid method. The formula is:
[0099]
[0100] where z i is the discrete output value, and μ(z i ) is its corresponding membership degree.
[0101] Real-time feedback loop: Design a real-time feedback loop to dynamically update the fuzzy rule weights according to the learner's operation delay d and the attention focus area Set the initial weight of the fuzzy rule k as w k, according to the operation delay and the change amount Δd of the attention focus area, Use the formula to update the weights, where is the learning rate function, which dynamically adjusts the learning rate according to the magnitudes of Δd and , and Δw k is the amount of weight adjustment.
[0102] Example 5:
[0103] This example details the dynamic evaluation process of the learning state by the evaluation and adjustment layer based on the hidden Markov model, as well as the specific implementation method of the knowledge graph constraint solution space, aiming to accurately evaluate the learner's learning state and ensure that the teaching content and strategies conform to the knowledge logical structure.
[0104] The dynamic evaluation of the learning state by the evaluation and adjustment layer includes:
[0105] Define the hidden state set and the observation state set: Define the hidden state set, which includes the cognitive level state S cognitive ={low, medium, high}, the emotional fluctuation state S emotiom ={stable, unstable}, and the metacognitive ability state S meta-cogmitive ={weak, strong}. Construct the observation state set, which covers the answer correct rate o accuracy , the interaction response time o response-time , and the physiological signal statistic o physiological (such as the average heart rate, the change rate of skin conductance response, etc.).
[0106] Evaluation based on the hidden Markov model: Evaluate the learning state based on the hidden Markov model. The model includes the state transition probability matrix A (A ij represents the probability of transitioning from the hidden state i to the hidden state j), the observation probability matrix B (B i j represents the probability of observing the observation state j under the hidden state i), and the initial state distribution π. When obtaining new observation data, use the forward-backward algorithm to calculate the posterior probability of the hidden state. The forward probability α t (i) represents the probability of being in the state i at time step t and observing the first t observation values, and the backward probability β t (i) represents the probability of being in the state i at time step t and observing the observation values from t + 1 to T. Based on these posterior probabilities, the current hidden state of the learner can be inferred, providing a reference for subsequent teaching strategy adjustment.
[0107] The implementation of the knowledge graph constraint solution space includes:
[0108] Construct logical constraint expressions: Construct logical constraint expressions to encode the prerequisite conditions, exclusion relationships, and optimal learning intervals among knowledge points into first-order predicate logic. For example, if knowledge point A is a prerequisite for knowledge point B, it can be expressed as where Study(k, x) represents that learner x studies knowledge point k.
[0109] Feasibility verification: Use a constraint satisfaction problem solver to verify the feasibility of the candidate solutions generated by the genetic algorithm. For each chromosome (teaching content sequence), check whether the combination of knowledge points therein satisfies the above logical constraint expressions. If a candidate solution violates any one of the constraints, it is determined as an infeasible solution and excluded.
[0110] Repair operator: Design a repair operator to locally adjust the chromosomes that partially violate the constraints. Use the greedy algorithm to select the gene locus with the least conflict for replacement. For example, for a chromosome that violates the prerequisite constraint, starting from the gene locus that violates the constraint, check the subsequent gene loci one by one to find a knowledge point identifier that can minimize the conflict after replacement.
[0111] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0112] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An online teaching optimization method based on emotion recognition, characterized in that, The method includes: Collecting real-time emotion data of learners through multi-modal sensors, where the multi-modal sensors include a facial expression acquisition module, a speech emotion analysis module, an eye movement tracking module, and a physiological signal detection module; performing multi-modal feature fusion on the real-time emotion data based on a graph convolutional neural network to generate an emotion state feature vector; inputting the emotion state feature vector into a meta-reinforcement learning model, where the meta-reinforcement learning model adopts a two-layer policy network structure and optimizes teaching strategies across tasks based on a dynamic reward mechanism to generate teaching strategy parameters; Constructing a dynamic curriculum generation model according to the teaching strategy parameters, where the dynamic curriculum generation model takes maximizing knowledge coverage and balancing cognitive load as optimization objectives, and uses an improved genetic algorithm to optimize the sequence of teaching content, where the improved genetic algorithm introduces an adaptive crossover probability and a mutation probability, and constrains the solution space through a knowledge graph; outputting optimal teaching sequence data based on the dynamic curriculum generation model; Constructing a hierarchical teaching control model according to the optimal teaching sequence data, where the hierarchical teaching control model includes a content planning layer, an interaction feedback layer, and an evaluation and adjustment layer, where the content planning layer performs global knowledge path planning based on the teaching strategy parameters, the interaction feedback layer adjusts the teaching interaction mode in real time based on the optimal teaching sequence data, and the evaluation and adjustment layer dynamically evaluates the learning state based on a hidden Markov model; generating teaching control instructions through the hierarchical teaching control model to achieve personalized online teaching optimization.
2. The online teaching optimization method based on emotion recognition according to claim 1, wherein The performing multi-modal feature fusion on the real-time emotion data based on a graph convolutional neural network to generate an emotion state feature vector includes: Constructing a multi-modal emotion graph structure, taking facial expression key points, speech spectrum features, eye movement trajectory coordinates, and physiological signal waveforms as graph nodes, and constructing an edge weight matrix based on the inter-modal relevance; using a graph attention mechanism to dynamically aggregate node features, and calculating the weight distribution between nodes through a multi-head attention layer; Designing a graph convolutional residual module, fusing shallow local features and deep global features through skip connections, and outputting a multi-modal joint embedding vector; performing temporal alignment processing on the joint embedding vector, and using a long short-term memory network to capture the temporal dependence of the emotion state to generate a temporal emotion feature sequence; Introducing an adversarial training mechanism, using a discriminator to distinguish between real emotion labels and generated feature distributions, and optimizing the generalization ability of the graph convolutional neural network.
3. The online teaching optimization method based on emotion recognition according to claim 1, wherein The meta-reinforcement learning model adopts a two-layer policy network structure, including: Constructing a meta-policy network and a base policy network, where the meta-policy network outputs meta-parameters shared by tasks, and the base policy network generates specific teaching actions based on the meta-parameters; designing a dynamic reward function, where the dynamic reward function includes a knowledge mastery reward term, an emotion stability reward term, and an interaction activity reward term, where the knowledge mastery reward term is calculated through a knowledge point association matrix and the answer correct rate, the emotion stability reward term is measured through an emotion entropy value, and the interaction activity reward term is generated based on a composite function of the question asking frequency and the response delay; Adopt a curriculum progressive training method to update network parameters in stages from simple teaching scenarios to complex scenarios, and dynamically adjust the training task distribution through a task sampler; construct a policy distillation loss function to perform KL divergence constraint on the parameter distribution output by the meta-policy network and the action distribution of the base policy network to achieve policy consistency optimization.
4. The online teaching optimization method based on emotion recognition according to claim 1, characterized in that The improved genetic algorithm for sequence optimization of teaching content includes: Construct chromosome encoding rules based on the knowledge graph, map the teaching unit sequence to a gene chain, and each gene contains a knowledge point identifier, a difficulty coefficient, and an association weight; design a multi-objective fitness function, and the fitness function includes knowledge coherence measurement, cognitive load balance, and emotion matching score, where the emotion matching score is calculated by the cosine similarity between the emotion state feature vector and the knowledge point emotion label; Introduce an adaptive genetic operator to dynamically adjust the crossover probability and mutation probability according to the population diversity index, where the crossover probability decays exponentially as the population similarity increases, and the mutation probability increases linearly in segments as the number of iterations increases; adopt an elite retention strategy to directly retain the optimal solution of each generation to the next generation, and constrain the mutation operation through the knowledge graph to ensure that the new solution conforms to the prior knowledge dependency relationship.
5. The online teaching optimization method based on emotion recognition according to claim 1, characterized in that The content planning layer for global knowledge path planning includes: Construct a knowledge topology graph, use knowledge points as nodes and dependency relationships as directed edges, and calculate the knowledge point weights through the PageRank algorithm; design a path generation model, represent the learning path as an implicit Markov chain, and solve the maximum probability path based on the Viterbi algorithm; Introduce a dynamic weight adjustment mechanism to update the knowledge point priority according to the real-time emotion state feature vector, and balance the contribution of historical learning data and the current emotion state through a weight decay factor; use Monte Carlo tree search to expand and simulate candidate paths, and select the path with the highest cumulative reward as the global planning result.
6. The online teaching optimization method based on emotion recognition according to claim 1, wherein The interaction feedback layer for real-time adjustment of teaching interaction modes includes: Construct a multi-channel interaction decision model, which receives the emotion state feature vector, the current teaching unit attributes, and the environmental context data, and outputs an interaction mode encoding; the interaction mode encoding includes the speech broadcast rate, the interface element layout strategy, and the Q&A difficulty gradient; use a fuzzy logic controller to finely adjust the interaction parameters, define the membership function of the input variables and the fuzzy rule base, and obtain the precise control quantity through the centroid method of defuzzification; design a real-time feedback loop to dynamically update the fuzzy rule weights based on the learner's operation delay and the attention focus area.
7. The online teaching optimization method based on emotion recognition according to claim 1, wherein The evaluation and adjustment layer for dynamic evaluation of the learning state based on the hidden Markov model includes: Define a set of hidden states, including cognitive level state, emotion fluctuation state, and metacognitive ability state; construct a set of observation states, including the answer correct rate, the interaction response time, and the physiological signal statistics.
8. The online teaching optimization method based on emotion recognition according to claim 1, characterized in that, The specific implementation of the knowledge graph to constrain the solution space includes: Construct a logical constraint expression to encode the prerequisite, exclusion relationship, and optimal learning interval between knowledge points into first-order predicate logic; use a constraint satisfaction problem solver to verify the feasibility of the candidate solutions generated by the genetic algorithm and eliminate individuals that violate the logical constraints; design a repair operator to locally adjust some chromosomes that violate the constraints, and select the gene locus with the least conflict for replacement through a greedy algorithm to ensure that the new solutions satisfy the topological constraints of the knowledge graph.
9. The online teaching optimization method based on emotion recognition according to claim 6, wherein The specific implementation of the multi-channel interaction decision-making model includes: Adopt a multi-task learning framework to share the underlying feature extraction network and parallelly output speech, visual, and text interaction strategies; design a collaborative loss function between channels to constrain the consistency of decision results of different channels through contrastive learning; Introduce a memory enhancement module to store historical interaction patterns and their effect evaluations, retrieve the best strategies for similar scenarios through an attention mechanism, and achieve experience transfer optimization.
10. An online teaching optimization system based on emotion recognition, which is used to implement the method described in any one of claims 1-9, characterized in that, Include: A data acquisition module for obtaining learners' facial expressions, speech signals, eye movement trajectories, and physiological signals through multi-modal sensors; A feature fusion module for jointly representing and learning multi-modal data based on a graph convolutional neural network to generate emotion state feature vectors; A strategy generation module that uses a meta-reinforcement learning model to output cross-task teaching strategy parameters and drives a dynamic curriculum generation model to optimize the teaching content sequence; A hierarchical control module, including a content planning unit, an interaction feedback unit, and an evaluation and adjustment unit, which respectively implement knowledge path planning, real-time interaction adjustment, and learning state evaluation; An execution and output module that converts the optimized teaching instructions into multi-modal outputs of speech, interface elements, and Q&A content.
Citation Information
Cited By
Intelligent display terminal interaction method and system applying AI model
CN120523335A
Intelligent display terminal interaction method and system applying AI model
CN120523335B
Semantic fingerprint adaptive training method for teaching service robot
CN120653994A
Multi-modal preference driven graph convolution combinatorial optimization learning path generation method
CN120894204A
A multimodal preference driven graph convolution combined optimization learning path generation method
CN120894204B