Teaching information processing method and system based on digital twinning

By constructing a digital twin model of teachers and students, synchronizing their behavioral data in real time, using multi-modal data collection and analysis, generating learning behavior feature vectors, based on this, comprehensive scoring of learning emotions and concentration, triggering the mode conversion mechanism, and dynamically adjusting the teaching mode, solving the problem of inability to perceive and feedback students' emotions and concentration in the existing technology in real time, and achieving personalized and adaptive teaching effects.

CN120146367AActive Publication Date: 2025-06-13YUNTIAN HENGZHI (QINGDAO) TECH CO LTD

Patent Information

Application Number
CN202411973363.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-06-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing teaching information processing methods based on digital twins mainly focus on a single data collection and feedback mechanism. The combination of deep learning and dynamic adjustment of teaching models has not been effectively realized. It is impossible to perceive and feedback students' emotions and concentration in real time, and the teaching model cannot be flexibly adjusted according to students' specific emotions and learning status.

Method used

By constructing a digital twin model of teachers and students, synchronizing their behavioral data in real time, using multimodal data acquisition and analysis, generating learning behavior feature vectors, based on this, comprehensive scoring of learning emotions and concentration, triggering the mode conversion mechanism, dynamically adjusting the teaching mode, and providing adaptive learning paths under the student-led interactive learning mode.

Benefits of technology

Real-time monitoring and feedback on students' learning emotions and concentration are achieved, and teaching mode is dynamically adjusted, which significantly improves personalized learning effects and teaching adaptability, ensuring that different students receive the best support in the learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146367A_ABST
    Figure CN120146367A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information processing, in particular to a teaching information processing method and system based on digital twinning, and the method comprises the following steps: constructing a teacher twinning model and a student twinning model, the teacher twinning model is used for synchronizing the teaching content, the teaching progress and the interaction instruction of a teacher, the student twinborn model is used for synchronizing the learning emotion and concentration degree comprehensive score of the student; based on the comprehensive score of the learning emotion and the concentration degree, triggering an interaction instruction in a teacher twinborn model, and adjusting a teaching mode; in an interactive learning mode dominated by students, the teacher twinborn model generates a guided discussion instruction to guide the students to perform deep learning exploration and discussion, and meanwhile, an adaptive learning path is provided in the student twinborn model. According to the invention, the teaching content and mode are dynamically optimized according to the real-time state of each student, and the personalized learning effect and the teaching adaptability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly to a teaching information processing method and system based on digital twin. Background Art

[0002] Digitalized and intelligent teaching models have been more and more widely applied in the education field. The traditional teaching model usually centers around teachers, and the learning process of students is relatively single, lacking personalized and targeted adjustments, resulting in great differences in the learning effects of different students.

[0003] In recent years, as an emerging technical means, digital twin technology has been applied in multiple industries and gradually introduced into the education field. Digital twin means that the state, behavior and changes of a physical object are reflected in real time through a virtual model, and the virtual model is continuously optimized and adjusted according to the actual feedback. In the education field, digital twin technology can construct digital twin models of teachers and students by synchronizing and feeding back teacher-student interaction data in real time, providing data support and intelligent decision-making for personalized teaching. However, the existing teaching information processing methods based on digital twin mainly focus on a single data collection and feedback mechanism, and have not effectively realized the combination of deep learning and the dynamic adjustment of teaching models.

[0004] In traditional education systems, although the collection and analysis of students' learning behavior data, such as academic performance, participation, exam performance, etc., have been introduced, these data mainly focus on the static analysis of the teaching process, lacking real-time perception and feedback on complex learning behaviors such as students' emotions and concentration. More importantly, the existing technologies often ignore the dynamic changes of teaching models in different learning states, cannot flexibly adjust teaching strategies according to students' specific emotions and learning states, and have not formed effective interaction and collaborative feedback between teachers and students. Summary of the Invention

[0005] The present invention provides a teaching information processing method and system based on digital twin.

[0006] The teaching information processing method based on digital twin includes the following steps:

[0007] S1: Construct teacher-student digital twin models, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teaching content, teaching progress, and interaction instructions of the teacher, and the student twin model is used to synchronize the comprehensive score of the student's learning emotion and concentration. The comprehensive score of the learning emotion and concentration is obtained based on the learning behavior feature vector. Facial expressions, voices, eye movements, and action data of the student are collected and analyzed to generate the learning behavior feature vector. The teacher twin model and the student twin model are connected to each other and updated in real time to dynamically reflect the teacher-student interaction situation;

[0008] S2: Mode conversion trigger mechanism. Based on the comprehensive score of learning emotion and concentration, it triggers the interaction instructions in the teacher twin model to adjust the teaching mode. When it detects that the comprehensive score of learning emotion and concentration is lower than the preset threshold, it triggers the teacher-led teaching mode; when the comprehensive score of learning emotion and concentration is higher than the preset threshold, it triggers the student-led interactive learning teaching mode, giving students more permissions for active feedback.

[0009] S3: In the student-led interactive learning mode, the teacher twin model will generate guiding discussion instructions to guide students to conduct in-depth learning exploration and discussion. At the same time, it provides an adaptive learning path in the student twin model to guide students to complete self-study tasks according to their own progress.

[0010] Optionally, the construction of the teacher twin model in S1 includes:

[0011] Teaching content synchronization, which is used to store and manage the teacher's teaching content, including the curriculum syllabus, chapters, and knowledge points, and synchronizes the display progress of the teaching content to the student twin model in real time to ensure that the content learned by students is consistent with the content taught by the teacher.

[0012] Teaching progress tracking, which is used to record the real-time progress of the teacher's teaching, track the content explanation speed and teaching rhythm in the classroom, adjust the advancement speed of the teaching content according to the progress data, and send a prompt to the teacher when it detects that the student's understanding progress is lower than the teaching progress to assist the teacher in adjusting the teaching speed.

[0013] Interaction instruction generation, which is used to generate interaction instructions according to the classroom situation, including questions, discussion topics, and task guidelines. The interaction instructions are triggered according to the teacher's teaching needs or the student's learning status, are transmitted to the student twin model in real time, and provide the teacher with the student feedback situation to help the teacher adjust the interaction content and frequency in a timely manner.

[0014] Optionally, the teacher twin model T model is represented as: T model =(C(t), P(t), I(t)), where C(t) is the teaching content matrix, including the curriculum syllabus, chapters, and knowledge point information, P(t) is the teacher's teaching progress vector, recording the advancement speed of the teaching content, I(t) is the teacher's interaction instruction vector, representing the classroom interaction instructions of different teaching modes, and t is the time step.

[0015] The interactive instruction vector is represented as a conditional trigger function: I(t) = f(R(t), T, D), where T is a preset threshold for setting the trigger frequency of the interactive instruction, D is a preset teaching mode, including a teacher-led teaching mode and a student-led interactive learning teaching mode, f represents a rule function for determining whether to trigger the interactive instruction, and R(t) is a comprehensive score of learning emotion and concentration for dynamically adjusting the teaching mode. If R(t) is lower than the preset threshold, a new instruction is triggered, and I(t) is updated to a new instruction vector and synchronized to the student twin model;

[0016] The teaching content matrix C(t) = {c 1 , c 2 ,..., c n}, where c n represents different units (outline, knowledge points, examples) of the course content. The progress status of each c n unit is synchronized with the student twin model at time t, so that the content of C(t) matches the student's learning path;

[0017] The teaching progress vector where represents the teacher's explanation speed of the content c n at time t.

[0018] Optionally, the generation of the learning behavior feature vector in S1 includes:

[0019] Multi-modal data collection: including facial expression recognition, speech analysis, eye movement tracking, and motion capture, which respectively collect students' facial expressions, speech, eye movements, and body movement data in real time through cameras, microphones, eye trackers, and motion sensors;

[0020] Multi-modal data analysis: Input the collected data into a pre-trained multi-branch neural network model to extract features from different data sources. Among them, facial expression data is used to extract emotion features, speech data is used to analyze intonation and emotion fluctuations, eye movement data is used to judge students' visual concentration, and motion data is used to identify students' participation and activity status;

[0021] Learning behavior feature vector generation: Integrate facial expression, speech, eye movement, and motion features to generate a multi-dimensional learning behavior feature vector S(t);

[0022] Comprehensive scoring: Based on the weighted calculation of each feature in the learning behavior feature vector, generate a comprehensive score representing the student's learning emotion and concentration.

[0023] Optionally, the multi-branch neural network model specifically includes:

[0024] Facial Expression Feature Branch Network: It takes a sequence of facial expression images as input. The network structure includes multiple convolutional layers, batch normalization layers, pooling layers, and fully connected layers, and the output is the facial expression feature vector E(t). The facial expression features of the student are extracted through this branch to identify emotional changes;

[0025] Speech Feature Branch Network: It takes the feature sequence (Mel spectrogram) of speech audio data as input, and the output is the speech feature vector V(t). This branch is used to extract speech pitch and emotional fluctuation features;

[0026] Eye Movement Feature Branch Network: It takes a sequence of eye movement trajectories or eye movement images (such as a sequence of fixation point positions) as input. The network structure includes convolutional layers and pooling layers, and the output is the eye movement feature vector G(t). Features related to visual attention are extracted through this branch;

[0027] Action Feature Branch Network: It takes a sequence of action data (such as skeleton points or accelerometer data) as input, and the output is the action feature vector M(t);

[0028] The output feature vectors of each branch are concatenated to obtain the learning behavior feature vector S(t):

[0029] S(t) = [E(t); V(t); G(t); M(t)].

[0030] Optionally, the comprehensive score R(t) of learning emotion and concentration is obtained by performing weighted calculation on the learning behavior feature vector S(t). The calculation expression of R(t) is:

[0031] where R(t) is the comprehensive score of the student, which is used to reflect the overall state of learning emotion and concentration, s i the i-th eigenvalue in the learning behavior feature vector, w i is the weight of each eigenvalue s i indicating the contribution degree of this feature to the comprehensive score, satisfying

[0032] Optionally, the provision of an adaptive learning path in the student twin model in S3 is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning state, specifically including:

[0033] S31, State Definition: The state is defined as the learning behavior feature vector S(t), including learning emotion and concentration, and the state is updated continuously over time;

[0034] S32, Action Definition: According to the current state of the student, select a suitable learning task or discussion guide. The learning task and discussion guide are regarded as the action A(t), which is a specific learning activity arranged for the student;

[0035] S33, Reward Mechanism: Rewards are given based on the changes in the students' states. When the comprehensive score of the students' learning mood and concentration is high, high rewards are given; otherwise, low rewards are given. The reward function is implemented, and the reward function is calculated based on the changes in the students' learning states, and the rewards reflect the students' learning effects;

[0036] S34, Design Strategy π: The strategy π(S(t)) is a strategy that selects the best action from the current student state S(t). A deep Q-network is used to approximate this strategy to maximize the long-term reward. After the deep Q-network is trained, given the current learning state S(t) of the student, the optimal action A(t) is selected through the deep Q-network to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real time according to the students' understanding levels and interest preferences.

[0037] Optionally, the deep Q-network specifically includes: Initializing a Q-network to estimate the value of each action taken by the student in the current state, that is, the Q-value. Creating an experience replay buffer that stores the states, actions, rewards, and next states during the interaction between the student and the learning content. The experience replay buffer is used to provide randomly sampled data for subsequent training. Through the feedback obtained from each interaction with the student, the Q-values in the Q-network are updated to optimize the strategy for selecting learning tasks. The goal of Q-value update is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q-value of the next state. During each student interaction, a learning task is selected according to the current state, and the ε-greedy strategy is used for selection. The ε-greedy strategy is: Most of the time, select the optimal action (maximize the Q-value), and occasionally randomly select an action for exploration. After executing the selected learning task, the state changes, and feedback is given according to the new state and reward value. Randomly sample historical experiences from the experience replay buffer and use the historical experiences to train the Q-network. Update the network parameters by minimizing the loss function, and repeat the iteration to continuously learn and optimize the Q-network, and select the optimal learning task in each learning state to maximize the long-term learning effect of the students.

[0038] A digital-twin-based teaching information processing system for implementing the above digital-twin-based teaching information processing method, including the following modules:

[0039] A teacher twin model for real-time synchronization of the teacher's teaching content, teaching progress, and interaction instructions, and dynamically reflecting the teacher's teaching behavior and teaching adjustments;

[0040] A student twin model is used to synchronize the learning emotions, concentration, and learning progress of students in real time. Based on multi-modal data (including students' facial expressions, voices, eye movements, and motion data), it generates learning behavior feature vectors to judge the learning state of students.

[0041] A mode conversion trigger module is used to trigger different teaching modes according to the comprehensive score of students' learning emotions and concentration, including a teacher-led teaching mode and a student-led interactive learning mode.

[0042] A collaborative learning instruction generation module generates guiding discussion instructions based on the learning state of students in the student-led interactive learning mode.

[0043] An adaptive learning path module constructs a personalized learning path based on a deep Q-network to guide students to complete self-study tasks.

[0044] The beneficial effects of the present invention:

[0045] In the present invention, by dynamically constructing and updating the digital models of teachers and students based on digital twins, it can monitor and feedback the learning emotions, concentration, and learning progress of students in real time, automatically adjust the teaching mode, analyze the emotional changes of students through learning behavior feature vectors, and trigger teacher-led or student-led interactive modes based on this feedback. This mechanism can dynamically optimize teaching content and methods according to the real-time state of each student, significantly improving the personalized learning effect and the adaptability of teaching, avoiding the "one-size-fits-all" problem in traditional teaching modes, and ensuring that different students receive the best support during the learning process.

[0046] In the present invention, by combining a role mode conversion mechanism and a deep learning model, an adaptive learning path module is introduced. In the student-led interactive learning mode, it will dynamically generate a personalized learning path according to the real-time learning progress and feedback of students, providing deep learning tasks suitable for the current state of students. This way not only encourages students to think and explore actively but also helps students complete learning tasks on a path suitable for their own rhythm, avoiding the limitations of previous single progress control.

[0047] In the present invention, by constructing digital twin models of teachers and students and synchronizing their behavior data in real time, the system can establish a two-way feedback closed-loop between teachers and students. When the teaching mode switches, the teacher twin model and the student twin model can respond and adjust quickly to ensure that the teaching strategies of teachers always match the actual learning state of students. This not only enhances the interactivity and timely feedback in the teaching process but also can flexibly adjust teaching strategies according to the learning state of students, effectively improving the teaching quality and teaching efficiency. Description of the Drawings

[0048] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 Schematic flowchart of the processing method according to an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the functional modules of the processing system according to an embodiment of the present invention. Detailed implementation manners

[0051] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0052] It should be pointed out that in the specification, the mention of "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicates that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0053] Generally, terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, can allow for the existence of other factors that may not be explicitly described.

[0054] As Figure 1 shown, the teaching information processing method based on digital twins includes the following steps:

[0055] S1: Construct teacher-student digital twin models, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teaching content, teaching progress, and interaction instructions of the teacher. The student twin model is used to synchronize the comprehensive score of the student's learning mood and concentration. The comprehensive score of the learning mood and concentration is obtained based on the learning behavior feature vector. Collect and analyze the student's facial expressions, speech, eye movements, and action data to generate the learning behavior feature vector. The teacher twin model and the student twin model are interconnected and updated in real time to dynamically reflect the teacher-student interaction situation;

[0056] S2: Mode conversion trigger mechanism. Based on the comprehensive score of the learning mood and concentration, trigger the interaction instructions in the teacher twin model and adjust the teaching mode. When it is detected that the score of the learning mood and concentration is lower than the preset threshold, trigger the teacher-led teaching mode; when the score of the learning mood and concentration is higher than the preset threshold, trigger the student-led interactive learning teaching mode and give the student more rights to provide active feedback;

[0057] S3: In the student-led interactive learning mode, the teacher twin model will generate guiding discussion instructions to guide the students to conduct in-depth learning exploration and discussion. At the same time, provide an adaptive learning path in the student twin model to guide the students to complete self-study tasks according to their own progress.

[0058] The construction of the teacher twin model in S1 includes:

[0059] Teaching content synchronization, which is used to store and manage the teaching content of the teacher, including the curriculum syllabus, chapters, and knowledge points, and synchronize the display progress of the teaching content to the student twin model in real time to ensure that the content learned by the students is consistent with the content taught by the teacher;

[0060] Teaching progress tracking, which is used to record the real-time progress of the teacher's teaching, track the content explanation speed and teaching rhythm in the classroom, adjust the advancement speed of the teaching content according to the progress data, and send a prompt to the teacher when it is detected that the student's understanding progress is lower than the teaching progress to assist the teacher in adjusting the teaching speed;

[0061] Interaction instruction generation, which is used to generate interaction instructions according to the classroom situation, including questions, discussion topics, and task guidance. The interaction instructions are triggered according to the teacher's teaching needs or the student's learning status, transmitted to the student twin model in real time, and provide the teacher with the student feedback situation to help the teacher adjust the interaction content and frequency in a timely manner.

[0062] Teacher twin model T model Denoted as: T model=(C(t), P(t), I(t)), where C(t) is the teaching content matrix, including syllabus, chapters, and knowledge point information; P(t) is the teaching progress vector of the teacher, recording the progress speed of the teaching content; I(t) is the interaction instruction vector of the teacher, representing the classroom interaction instructions of different teaching modes; and t is the time step.

[0063] The interaction instruction vector is expressed as a conditional trigger function: I(t) = f(R(t), T, D), where T is a preset threshold for setting the trigger frequency of the interaction instruction, D is the preset teaching mode, including the teacher-led teaching mode and the student-led interactive learning teaching mode, f represents a rule function for judging whether to trigger the interaction instruction, and R(t) is the comprehensive score of learning emotion and concentration, used to dynamically adjust the teaching mode. If R(t) is lower than the preset threshold, a new instruction is triggered, and I(t) is updated to a new instruction vector and synchronized to the student twin model.

[0064] The teaching content matrix C(t) = {c 1 , c 2 ,..., c n}, where c n represents different units (syllabus, knowledge points, examples) of the course content. The progress status of each c n unit is synchronized with the student twin model at time t to match the content of C(t) with the student's learning path.

[0065] The teaching progress vector where represents the teacher's explanation speed of the content c n at time t, continuously updating the progress information to adapt to the teaching rhythm according to classroom feedback, and triggering adjustment suggestions through logical conditions when the student's understanding progress is lower than the teacher's progress.

[0066] The generation of the learning behavior feature vector in S1 includes:

[0067] Multi-modal data collection: including facial expression recognition, speech analysis, eye movement tracking, and motion capture, which respectively collect the student's facial expressions, speech, eye movements, and limb movement data in real time through cameras, microphones, eye trackers, and motion sensors.

[0068] Multi-modal data analysis: Input the collected data into a pre-trained multi-branch neural network model to extract features from different data sources. Among them, facial expression data is used to extract emotion features, speech data is used to analyze intonation and emotional fluctuations, eye movement data is used to judge the student's visual concentration, and motion data is used to identify the student's participation and activity status.

[0069] Learning behavior feature vector generation: Integrate facial expressions, speech, eye movements, and motion features to generate a multi-dimensional learning behavior feature vector S(t);

[0070] Comprehensive scoring: Based on the features in the learning behavior feature vector, perform weighted calculations to generate a comprehensive score representing the learning mood and concentration of students.

[0071] The multi-branch neural network model specifically includes:

[0072] Facial expression feature branch network: Input a sequence of facial expression images. The network structure includes multiple convolutional layers, batch normalization layers, pooling layers, and fully connected layers. The output is a facial expression feature vector E(t). Through this branch, the facial expression features of students are extracted to identify emotional changes;

[0073] Speech feature branch network: Input the feature sequence (Mel spectrogram) of speech audio data, and the output is a speech feature vector V(t). This branch is used to extract speech pitch and emotional fluctuation features;

[0074] Eye movement feature branch network: Input an eye movement trajectory sequence or eye movement images (such as a sequence of fixation point positions). The network structure includes convolutional layers and pooling layers, and the output is an eye movement feature vector G(t). Through this branch, features related to visual concentration are extracted;

[0075] Motion feature branch network: Input a sequence of motion data (such as skeleton points or accelerometer data), and the output is a motion feature vector M(t);

[0076] Each feature vector can be represented as:

[0077] E(t) = {e 1 , e 2 ,..., e l}; V(t) = {v 1 , v 2 ,..., v j}; G(t) = {g 1 , g 2 ,..., g p}; M(t) = {m 1 , m 2 ,... m q};

[0078] Facial expression feature vector E(t): Represents the emotional state of students;

[0079] Speech feature vector V(t): Represents the pitch and tone fluctuations of students;

[0080] Eye movement feature vector G(t): representing the visual concentration of students;

[0081] Action feature vector M(t): representing the physical activities and participation of students.

[0082] Concatenate the output feature vectors of each branch to obtain the learning behavior feature vector S(t):

[0083] S(t) = [E(t); V(t); G(t); M(t)].

[0084] The comprehensive score R(t) of learning emotion and concentration is obtained by weighted calculation of the learning behavior feature vector S(t), and the calculation expression of R(t) is:

[0085] where R(t) is the comprehensive score of students, used to reflect the overall state of learning emotion and concentration, s i The i-th eigenvalue in the learning behavior feature vector, w i is the weight of each eigenvalue s i indicating the contribution degree of this feature to the comprehensive score, satisfying

[0086] Input R(t) into the conditional trigger function, and standardize the range of the comprehensive score R(t) to between 0 and 1. The preset threshold T is set as a percentage, 70%, that is, 0.7. The situation where R(t) exceeds 0.7 is regarded as the better learning emotion and concentration of students, and they can enter the student-led interaction mode. When it is lower than 0.7, the teacher-led mode is triggered so that the teacher can provide more help. It can also be differentiated for different classes. For example, class A is set to 0.7 and class B is set to 0.5.

[0087] Providing an adaptive learning path in the student twin model in S3 is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning state, specifically including:

[0088] S31, State definition: Define the state as the learning behavior feature vector S(t), including learning emotion and concentration, and the state is continuously updated over time;

[0089] S32, Action definition: According to the current state of the student, select a suitable learning task or discussion guide. The learning task and discussion guide are regarded as the action A(t), which is the specific learning activity arranged for the student. The action set A includes different learning units or topics, and the system can select appropriate learning tasks or discussion guides according to the student's state;

[0090] S33, Reward Mechanism: Rewards are given based on the changes in the students' states. When the comprehensive score of the students' learning mood and concentration is high, high rewards are given; otherwise, low rewards are given. It is implemented through the reward function The reward function is calculated based on the changes in the students' learning states. The rewards reflect the students' learning effects. When students show a high level of understanding, participation, or positive feedback, high rewards are given; when students show low interest or learning difficulties, low rewards are given. The reward function is defined as:

[0091] where K t+1 and P t+1 are the knowledge mastery and interest preferences after performing actions, D represents the difficulty of the learning task, and when the difficulty is high, the rewards decrease; w 1 , w 2 , w 3 are the weights of different factors. The purpose of the rewards is to motivate students to improve their learning effects;

[0092] S34, Design Strategy π: The strategy π(S(t)) is a strategy that selects the best action from the current student state S(t). A deep Q-network is used to approximate this strategy to maximize the long-term rewards. After the deep Q-network is trained, given the current learning state S(t) of the student, the optimal action A(t) is selected through the deep Q-network to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real-time according to the students' understanding levels and interest preferences to ensure that students can obtain the maximum learning effect in in-depth learning exploration and discussion, provide personalized learning paths, guide students to explore deeper content based on their own learning progress and feedback, and promote students' active learning and independent exploration.

[0093] The Deep Q-Network specifically includes: initializing a Q-Network to estimate the value of each action that a student takes in the current state, i.e., the Q-value; creating an experience replay buffer that stores the states, actions, rewards, and next states during the interaction between the student and the learning content. The experience replay buffer is used to provide randomly sampled data for subsequent training. Through the feedback obtained from each interaction with the student, update the Q-value in the Q-Network to optimize the strategy for selecting learning tasks. The goal of Q-value update is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q-value of the next state. During each student interaction, select a learning task according to the current state, and use the ε-greedy strategy for selection. The ε-greedy strategy is as follows: most of the time, select the optimal action (maximize the Q-value), and occasionally randomly select an action for exploration. After executing the selected learning task, the state will change, and feedback is given according to the new state and reward value. Randomly sample historical experiences from the experience replay buffer, and use the historical experiences to train the Q-Network. Update the network parameters by minimizing the loss function, and repeat the iteration to enable the Q-Network to continuously learn and optimize, and select the optimal learning task in each learning state, so as to maximize the long-term learning effect of the student.

[0094] In the student-led interactive learning mode, not only learning tasks are provided, but also guiding discussion instructions are generated to encourage students to think deeply and participate actively. These tasks are customized according to the student's learning progress, interest preferences, and emotional feedback to improve the student's sense of participation and learning motivation.

[0095] Through the continuously optimized learning path, it is possible to guide students from superficial understanding to in-depth knowledge exploration, stimulate students to actively give feedback and discuss, and ultimately help students complete learning tasks at their own pace, thus realizing a personalized and in-depth learning process.

[0096] The specific calculation of the Deep Q-Network (DQN) is as follows:

[0097] 1. Initialize the Q-Network: Initialize a neural network Q(S(t), A(t); θ), which is used to estimate the Q-value function with parameter θ.

[0098] 2. Experience replay buffer: Establish an experience replay buffer D for storing the experiences of student state transitions for random sampling during training.

[0099] 3. Q-value update formula: Use the following Q-value update formula for network training:

[0100]

[0101] where γ is the discount factor representing the decay of future rewards, and A′ represents the possible next action.

[0102] 4. Selection Action (Learning Task): At the current state S(t) of the student, use an ε-greedy strategy to select the action A(t):

[0103] Randomly select an action with probability ε to ensure exploration;

[0104] Select the action that maximizes Q(S(t), A) with probability 1 - ε to ensure exploitation.

[0105] 5. Execute Action and Feedback: Execute the selected learning task, and the student twin model provides feedback on the new learning state S(t + 1) and the corresponding reward.

[0106] 6. Network Training: Randomly sample a batch of data from the experience buffer. Use the mean squared error as the loss function to train the Q-network.

[0107] 7. Iterative Update: Repeat the above processes 1 - 6 so that the Q-network can continuously optimize the strategy of selecting learning tasks, enabling the student to obtain the maximum long-term reward in each state, i.e., the optimal learning path.

[0108] As Figure 2 shown, the digital-twin-based teaching information processing system for implementing the above digital-twin-based teaching information processing method includes the following modules:

[0109] Teacher Twin Model, used to synchronize the teaching content, teaching progress, and interaction instructions of the teacher in real time, and dynamically reflect the teaching behavior and teaching adjustments of the teacher;

[0110] Student Twin Model, used to synchronize the learning emotions, attention, and learning progress of the student in real time, generate learning behavior feature vectors based on multi-modal data (including the student's facial expressions, speech, eye movements, and action data), and thus judge the learning state of the student;

[0111] Mode Conversion Trigger Module, used to trigger different teaching modes according to the comprehensive score of the student's learning emotions and attention, including the teacher-led teaching mode and the student-led interactive learning mode;

[0112] Collaborative Learning Instruction Generation Module, which generates guiding discussion instructions based on the learning state of the student in the student-led interactive learning mode;

[0113] Adaptive Learning Path Module, which constructs a personalized learning path based on the deep Q-network to guide the student to complete self-study tasks.

[0114] The present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention. In order to enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components and circuits, etc. are not described in detail in order to avoid unnecessary confusion to the essence of the present invention.

[0115] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A teaching information processing method based on digital twins, characterized in that: The following steps are involved: S1: Construct a digital twin model of teachers and students, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teacher's teaching content, teaching progress, and interactive instructions. The student twin model is used to synchronize the student's learning emotion and concentration comprehensive score. The learning emotion and concentration comprehensive score is obtained based on the learning behavior feature vector. The student's facial expression, voice, eye movement, and action data are collected and analyzed to generate a learning behavior feature vector. The teacher twin model and the student twin model are interconnected and updated in real time to dynamically reflect the interaction between teachers and students. S2: Mode conversion trigger mechanism, based on the comprehensive score of learning emotion and concentration, triggers the interactive instructions in the teacher twin model and adjusts the teaching mode. When it is detected that the learning emotion and concentration score are lower than the preset threshold, the teacher-led teaching mode is triggered; when the learning emotion and concentration score are higher than the preset threshold, the student-led interactive learning teaching mode is triggered; S3: In the student-led interactive learning mode, the teacher twin model will generate guided discussion instructions to guide students to conduct in-depth learning exploration and discussion, while providing an adaptive learning path in the student twin model.

2. The teaching information processing method based on digital twin according to claim 1 is characterized in that: The construction of the teacher twin model in S1 includes: Teaching content synchronization is used to store and manage teachers' teaching content, including course outlines, chapters, and knowledge points, and synchronize the presentation progress of teaching content to the student twin model in real time; Teaching progress tracking is used to record the real-time progress of teachers’ teaching, track the speed of content explanation and teaching rhythm in class, and adjust the advancement speed of teaching content according to progress data; Interactive instruction generation is used to generate interactive instructions based on classroom situations, including questions, discussion topics, and task instructions. Interactive instructions are triggered based on the teacher's teaching needs or the student's learning status and transmitted to the student twin model in real time.

3. The teaching information processing method based on digital twin according to claim 2 is characterized in that: The teacher twin model T model Expressed as: T model =(C(t), P(t), I(t)), where C(t) is the teaching content matrix, including the course outline, chapters, and knowledge point information; P(t) is the teacher's teaching progress vector, recording the advancement speed of the teaching content; I(t) is the teacher's interactive instruction vector, indicating the classroom interactive instructions of different teaching modes; and t is the time step; The interactive instruction vector is expressed as a conditional trigger function: I(t)=f(R(t),T,D), where T is a preset threshold used to set the trigger frequency of the interactive instruction, D is a preset teaching mode, including a teacher-led teaching mode and a student-led interactive learning teaching mode, f represents a rule function used to determine whether to trigger the interactive instruction, R(t) is a comprehensive score of learning emotion and concentration, used to dynamically adjust the teaching mode, if R(t) is lower than the preset threshold, a new instruction is triggered, I(t) is updated to a new instruction vector and synchronized to the student twin model; The teaching content matrix C(t) = {c1, c2, ..., c n }, where c n Represents different units of course content, each c n The progress status of the unit is synchronized with the student twin model at time t, so that the content of C(t) matches the student's learning path; The teaching progress vector in, It means that at time t, the teacher has n speed of explanation.

4. The teaching information processing method based on digital twin according to claim 3 is characterized in that: The generation of the learning behavior feature vector in S1 includes: Multimodal data collection: including facial expression recognition, speech analysis, eye tracking and motion capture, which collects students' facial expressions, speech, eye movements and body movements in real time through cameras, microphones, eye trackers and motion sensors; Multimodal data analysis: The collected data is input into a pre-trained multi-branch neural network model to extract features from different data sources; facial expression data is used to extract emotional features, voice data is used to analyze intonation and emotional fluctuations, eye movement data is used to determine students' visual concentration, and action data is used to identify students' participation and activity status; Learning behavior feature vector generation: Facial expression, voice, eye movement, and action features are integrated to generate a multi-dimensional learning behavior feature vector S(t); Comprehensive score: Based on the weighted calculation of each feature in the learning behavior feature vector, a comprehensive score representing the student's learning mood and concentration is generated.

5. The teaching information processing method based on digital twin according to claim 4 is characterized in that: The multi-branch neural network model specifically includes: Facial expression feature branch network: input facial expression image sequence, the network structure includes multiple convolutional layers, batch normalization layers, pooling layers and fully connected layers, and the output is the facial expression feature vector E(t); Speech feature branch network: inputs the feature sequence of speech audio data and outputs the speech feature vector V(t); Eye movement feature branch network: input eye movement trajectory sequence or eye movement image, the network structure includes convolution layer and pooling layer, and the output is eye movement feature vector G(t); Action feature branch network: input action data sequence and output action feature vector M(t); The output feature vectors of each branch are concatenated to obtain the learning behavior feature vector S(t): S(t)=[E(t); V(t); G(t); M(t)].

6. The teaching information processing method based on digital twin according to claim 5 is characterized in that: The comprehensive score of learning emotion and concentration R(t) is obtained by weighted calculation of the learning behavior feature vector S(t). The calculation expression of R(t) is: Among them, R(t) is the comprehensive score of the students, which is used to reflect the overall state of learning mood and concentration, and s i The i-th eigenvalue in the learning behavior eigenvector, w i For each eigenvalue s i The weight of represents the contribution of this feature to the comprehensive score, satisfying 7. The teaching information processing method based on digital twin according to claim 4 is characterized in that: The adaptive learning path provided in the student twin model in S3 is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning status, specifically including: S31, state definition: The state is defined as the learning behavior feature vector S(t), including learning emotions and concentration, and the state is continuously updated over time; S32, action definition: select learning tasks or discussion guidance according to the current status of students. Learning tasks and discussion guidance are regarded as actions A(t), which are specific learning activities arranged for students; S33, Reward mechanism: rewards are given according to the changes in students' status. When the comprehensive score of students' learning mood and concentration is high, high rewards are given; conversely, low rewards are given through the reward function. Implementation, reward function Calculated based on the changes in students’ learning status, rewards reflect students’ learning outcomes; S34, design strategy π: Strategy π(S(t)) is a strategy for selecting the best action from the current student state S(t). A deep Q network is used to approximate the strategy to maximize long-term rewards. After the deep Q network is trained, given the student’s current learning state S(t), the deep Q network selects the optimal action A(t) to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real time according to the student’s understanding level and interest preferences.

8. The teaching information processing method based on digital twin according to claim 7 is characterized in that: The deep Q network specifically includes: initializing a Q network to estimate the value of each action taken by the student in the current state, that is, the Q value, creating an experience replay buffer that stores the state, action, reward and next state of the student during the interaction with the learning content, the experience replay buffer is used to provide randomly sampled data for subsequent training, and the Q value in the Q network is updated through the feedback obtained from each interaction with the student to optimize the strategy for selecting learning tasks.

9. The teaching information processing method based on digital twin according to claim 8 is characterized in that: The goal of the Q-value update is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q-value of the next state. At each student interaction, a learning task is selected based on the current state, and the ε-greedy strategy is used for selection. After executing the selected learning task, the state will change, and feedback will be given based on the new state and reward value. Historical experience is randomly sampled from the experience replay buffer, and the Q network is trained with historical experience. The network parameters are updated by minimizing the loss function, and iterations are repeated so that the Q network continues to learn and optimize.

10. A teaching information processing system based on digital twins, used to implement the teaching information processing method based on digital twins as described in any one of claims 1 to 9, characterized in that: Includes the following modules: The teacher twin model is used to synchronize the teacher's teaching content, teaching progress and interactive instructions in real time, dynamically reflecting the teacher's teaching behavior and teaching adjustments; The student twin model is used to synchronize students’ learning emotions, concentration, and learning progress in real time, and to generate learning behavior feature vectors based on multimodal data to determine students’ learning status; Mode conversion trigger module, which is used to trigger different teaching modes according to the comprehensive scores of students' learning emotions and concentration, including teacher-led teaching mode and student-led interactive learning mode; The collaborative learning instruction generation module generates guided discussion instructions based on students’ learning status in a student-led interactive learning mode; The adaptive learning path module, based on the deep Q network, builds personalized learning paths to guide students to complete self-study tasks.

Citation Information

Patent Citations

  • Teaching information processing method and system based on digital twinning

    CN117252047A

  • Management teaching information processing method and system based on digital twinning

    CN119151743A

  • Automated training of failure diagnosis models for application in self-organizing networks

    US20240172001A1

Cited By

  • Multi-scene fusion learning tutoring system and method, storage medium and program product

    CN121685205A