Teaching information processing method and system based on digital twin

By building a digital twin model of teachers and students and a deep learning model, and dynamically adjusting the teaching model, the problem of lack of personalization and flexibility in the teaching model in existing technologies is solved, and personalized learning effects and teaching efficiency are improved.

CN120146367BActive Publication Date: 2025-09-16YUNTIAN HENGZHI (QINGDAO) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411973363.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-09-16
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing digital twin-based teaching information processing methods fail to effectively realize the integration of deep learning and dynamic adjustment of teaching models. They lack real-time perception and feedback of students' emotions and concentration, resulting in a lack of personalization and flexibility in teaching models.

Method used

Build a digital twin model of teachers and students, generate learning behavior feature vectors through multimodal data collection and analysis, combine mode conversion trigger mechanism and deep learning model, dynamically adjust the teaching mode, and provide personalized learning paths and interactive learning modes.

Benefits of technology

It realizes real-time adjustment of teaching mode and personalized learning effect, enhances interactivity and feedback in the teaching process, ensures that teaching strategies are consistent with students' status, and improves teaching quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146367B_ABST
    Figure CN120146367B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information processing technology, and specifically to a teaching information processing method and system based on digital twins, comprising the following steps: constructing a teacher twin model and a student twin model, wherein the teacher twin model is used to synchronize the teacher's teaching content, teaching progress, and interactive instructions, and the student twin model is used to synchronize the student's learning mood and comprehensive score of concentration; based on the comprehensive score of learning mood and concentration, the interactive instructions in the teacher twin model are triggered to adjust the teaching mode; in the student-led interactive learning mode, the teacher twin model will generate guided discussion instructions to guide students to conduct in-depth learning exploration and discussion, while providing an adaptive learning path in the student twin model. The present invention dynamically optimizes the teaching content and method according to the real-time status of each student, significantly improving the personalized learning effect and the adaptability of teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a teaching information processing method and system based on digital twins. Background Art

[0002] Digital and intelligent teaching models are being used more and more widely in the field of education. Traditional teaching models are usually teacher-centered, and the students' learning process is relatively simple, lacking personalized and targeted adjustments, resulting in large differences in learning outcomes among different students.

[0003] In recent years, digital twin technology, as an emerging technology, has been applied in multiple industries and is gradually being introduced into education. Digital twins are virtual models that reflect the state, behavior, and changes of physical objects in real time, and continuously optimize and adjust these virtual models based on actual feedback. In education, digital twin technology can construct digital twin models of teachers and students by synchronizing and providing feedback on teacher-student interaction data in real time, providing data support and intelligent decision-making for personalized teaching. However, existing digital twin-based teaching information processing methods mainly focus on single data collection and feedback mechanisms, and have not yet effectively integrated deep learning and dynamically adjusted teaching models.

[0004] In the traditional education system, although the collection and analysis of students' learning behavior data, such as academic performance, participation, and test performance, have been introduced, these data are mostly concentrated on the static analysis of the teaching process and lack real-time perception and feedback of complex learning behaviors such as students' emotions and concentration. More importantly, existing technologies often ignore the dynamic changes of teaching models under different learning conditions, cannot flexibly adjust teaching strategies according to students' specific emotions and learning status, and fail to form effective interaction and collaborative feedback between teachers and students. Summary of the Invention

[0005] The present invention provides a teaching information processing method and system based on digital twins.

[0006] The teaching information processing method based on digital twins includes the following steps:

[0007] S1: Construct a digital twin model of teachers and students, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teacher's teaching content, teaching progress, and interactive instructions. The student twin model is used to synchronize the student's learning emotion and concentration comprehensive score. The learning emotion and concentration comprehensive score is obtained based on the learning behavior feature vector. The student's facial expression, voice, eye movement, and action data are collected and analyzed to generate a learning behavior feature vector. The teacher twin model and the student twin model are interconnected and updated in real time to dynamically reflect the interaction between teachers and students.

[0008] S2: Mode switching trigger mechanism, based on the comprehensive score of learning emotion and concentration, triggers the interactive instructions in the teacher twin model and adjusts the teaching mode. When the learning emotion and concentration scores are detected to be lower than the preset threshold, the teacher-led teaching mode is triggered; when the learning emotion and concentration scores are higher than the preset threshold, the student-led interactive learning and teaching mode is triggered, giving students more authority to actively provide feedback;

[0009] S3: In the student-led interactive learning mode, the teacher twin model will generate guided discussion instructions to guide students to conduct in-depth learning exploration and discussion. At the same time, it will provide adaptive learning paths in the student twin model to guide students to complete self-study tasks according to their own progress.

[0010] Optionally, the construction of the teacher twin model in S1 includes:

[0011] Teaching content synchronization is used to store and manage teachers' teaching content, including course outlines, chapters, and knowledge points, and synchronize the teaching content display progress to the student twin model in real time to ensure that what students learn is consistent with what the teacher teaches;

[0012] Teaching progress tracking is used to record the real-time progress of teachers' teaching, track the speed of content explanation and teaching rhythm in class, adjust the speed of teaching content according to progress data, and issue a prompt to the teacher when it is detected that the student's understanding progress is lower than the teaching progress, so as to assist the teacher in adjusting the teaching speed;

[0013] Interactive instruction generation is used to generate interactive instructions based on classroom situations, including questions, discussion topics, and task instructions. Interactive instructions are triggered based on the teacher's teaching needs or the student's learning status, and are transmitted to the student twin model in real time. Student feedback is also provided to the teacher to help the teacher adjust the content and frequency of the interaction as appropriate.

[0014] Optionally, the teacher twin model T model Expressed as: T model =(C(t), P(t), I(t)), where C(t) is the teaching content matrix, including the course outline, chapters, and knowledge point information; P(t) is the teacher's teaching progress vector, recording the advancement speed of the teaching content; I(t) is the teacher's interactive instruction vector, representing the classroom interactive instructions of different teaching modes; and t is the time step;

[0015] The interactive instruction vector is expressed as a conditional trigger function: I(t) = f(R(t), T, D), where T is a preset threshold used to set the trigger frequency of the interactive instruction, D is a preset teaching mode, including a teacher-led teaching mode and a student-led interactive learning teaching mode, f represents a rule function used to determine whether to trigger the interactive instruction, and R(t) is a comprehensive score of learning emotion and concentration, which is used to dynamically adjust the teaching mode. If R(t) is lower than the preset threshold, a new instruction is triggered, and I(t) is updated to the new instruction vector and synchronized to the student twin model.

[0016] The teaching content matrix C(t)={c1,c2,...,c n}, where c n Indicates different units of course content (outline, knowledge points, examples), each c n The progress status of the unit is synchronized with the student twin model at time t, so that the content of C(t) matches the student's learning path;

[0017] The teaching progress vector in, Indicates the teacher's understanding of content c at time t n speed of explanation.

[0018] Optionally, the generating of the learning behavior feature vector in S1 includes:

[0019] Multimodal data collection: including facial expression recognition, speech analysis, eye tracking, and motion capture, using cameras, microphones, eye trackers, and motion sensors to collect students' facial expressions, speech, eye movements, and body movement data in real time;

[0020] Multimodal data analysis: The collected data is fed into a pre-trained multi-branch neural network model to extract features from different data sources. Facial expression data is used to extract emotional features, voice data is used to analyze intonation and emotional fluctuations, eye movement data is used to determine students' visual focus, and motion data is used to identify students' engagement and activity status.

[0021] Learning behavior feature vector generation: Facial expression, voice, eye movement, and action features are integrated to generate a multi-dimensional learning behavior feature vector S(t);

[0022] Comprehensive score: Based on the weighted calculation of each feature in the learning behavior feature vector, a comprehensive score representing the student's learning mood and concentration is generated.

[0023] Optionally, the multi-branch neural network model specifically includes:

[0024] Facial expression feature branch network: The input is a facial expression image sequence. The network structure includes multiple convolutional layers, batch normalization layers, pooling layers, and fully connected layers. The output is the facial expression feature vector E(t). This branch extracts the student's facial expression features to identify emotional changes.

[0025] Speech feature branch network: It inputs the feature sequence (Mel spectrum) of speech audio data and outputs the speech feature vector V(t). This branch is used to extract speech pitch and emotional fluctuation features.

[0026] Eye movement feature branch network: The input is an eye movement trajectory sequence or an eye movement image (such as a sequence of gaze point positions). The network structure includes a convolutional layer and a pooling layer. The output is an eye movement feature vector G(t). This branch is used to extract features related to visual concentration.

[0027] Motion feature branch network: input motion data sequence (such as skeleton points or accelerometer data) and output motion feature vector M(t);

[0028] The output feature vectors of each branch are concatenated to obtain the learning behavior feature vector S(t):

[0029] S(t)=[E(t); V(t); G(t); M(t)].

[0030] Optionally, the comprehensive score of learning emotion and concentration R(t) is obtained by weighted calculation of the learning behavior feature vector S(t), and the calculation expression of R(t) is:

[0031] Among them, R(t) is the comprehensive score of the students, which is used to reflect the overall state of learning mood and concentration, s i The i-th eigenvalue in the learning behavior eigenvector, w i For each eigenvalue s i The weight of , which indicates the contribution of this feature to the comprehensive score, satisfies

[0032] Optionally, the adaptive learning path provided in the student twin model in S3 is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning status, specifically including:

[0033] S31, state definition: The state is defined as the learning behavior feature vector S(t), including learning emotions and concentration. The state is continuously updated over time;

[0034] S32, Action Definition: Select appropriate learning tasks or discussion guides based on the student's current state. Learning tasks and discussion guides are considered actions A(t), which are specific learning activities arranged for students.

[0035] S33, Reward Mechanism: Rewards are given according to the changes in the student's status. When the comprehensive score of the student's learning mood and concentration is high, a high reward is given; on the contrary, the reward is low. Implementation, reward function Calculated based on changes in students’ learning status, rewards reflect students’ learning outcomes;

[0036] S34, design strategy π: Strategy π(S(t)) is a strategy for selecting the best action from the current student state S(t). The deep Q network is used to approximate the strategy to maximize the long-term reward. After the deep Q network training is completed, given the student's current learning state S(t), the deep Q network selects the optimal action A(t) to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real time according to the student's understanding level and interest preferences.

[0037] Optionally, the deep Q network specifically includes: initializing a Q network to estimate the value of each action taken by the student in the current state, i.e., the Q value; creating an experience replay buffer to store the state, action, reward and next state of the student during the interaction with the learning content; the experience replay buffer is used to provide randomly sampled data for subsequent training; the Q value in the Q network is updated through the feedback obtained from each interaction with the student to optimize the strategy for selecting learning tasks; the goal of Q value updating is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q value of the next state; each time the student interacts, the learning task is selected according to the current state, and the ε-greedy strategy is used for selection; the ε-greedy strategy is: most of the time, the optimal action is selected (maximizing the Q value), and occasionally random actions are selected for exploration; after executing the selected learning task, the state will change, and feedback is provided based on the new state and reward value; historical experience is randomly sampled from the experience replay buffer, the Q network is trained with historical experience, the network parameters are updated by minimizing the loss function, and iteration is repeated so that the Q network continuously learns and optimizes, and the optimal learning task is selected in each learning state, thereby maximizing the student's long-term learning effect.

[0038] The digital twin-based teaching information processing system is used to implement the above-mentioned digital twin-based teaching information processing method, and includes the following modules:

[0039] The teacher twin model is used to synchronize the teacher's teaching content, teaching progress and interactive instructions in real time, dynamically reflecting the teacher's teaching behavior and teaching adjustments;

[0040] The student twin model is used to synchronize students' learning emotions, concentration, and learning progress in real time. It generates learning behavior feature vectors based on multimodal data (including students' facial expressions, voice, eye movements, and movement data) to determine students' learning status.

[0041] The mode switching trigger module is used to trigger different teaching modes based on the comprehensive scores of students' learning mood and concentration, including teacher-led teaching mode and student-led interactive learning mode;

[0042] The collaborative learning instruction generation module generates guided discussion instructions based on students' learning status in a student-led interactive learning mode;

[0043] The adaptive learning path module, based on the deep Q network, builds personalized learning paths to guide students to complete self-study tasks.

[0044] Beneficial effects of the present invention:

[0045] The present invention dynamically constructs and updates digital models of teachers and students through a digital twin-based system, which can monitor and provide feedback on students' learning emotions, concentration and learning progress in real time, automatically adjust the teaching mode, analyze students' emotional changes through learning behavior feature vectors, and trigger teacher-led or student-led interaction modes based on this feedback. This mechanism can dynamically optimize teaching content and methods according to the real-time status of each student, significantly improving personalized learning effects and teaching adaptability, avoiding the "one-size-fits-all" problem in traditional teaching models, and ensuring that different students receive the best support in the learning process.

[0046] This invention introduces an adaptive learning path module by combining the role mode conversion mechanism and the deep learning model. In the student-led interactive learning mode, it will dynamically generate personalized learning paths based on students' real-time learning progress and feedback, and provide in-depth learning tasks suitable for students' current status. This approach not only encourages students to think and explore actively, but also helps students complete learning tasks on a path that suits their own pace, avoiding the limitations of previous single progress control.

[0047] This system, by constructing digital twin models of teachers and students and synchronizing their behavioral data in real time, establishes a two-way feedback loop between teachers and students. When teaching modes change, the teacher twin model and the student twin model can quickly respond and adjust, ensuring that the teacher's teaching strategy remains consistent with the student's actual learning status. This not only enhances interactivity and timely feedback during the teaching process, but also allows flexible adjustment of teaching strategies based on the student's learning status, effectively improving teaching quality and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0049] Figure 1 A schematic flow chart of a processing method according to an embodiment of the present invention;

[0050] Figure 2 Schematic diagram of the functional modules of the processing system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.

[0052] It should be noted that references in the specification to "one embodiment," "an embodiment," "an exemplary embodiment," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes such specific features, structures, or characteristics. In addition, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).

[0053] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.

[0054] like Figure 1 As shown in FIG, the teaching information processing method based on digital twins includes the following steps:

[0055] S1: Construct a digital twin model of teachers and students, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teacher's teaching content, teaching progress, and interactive instructions. The student twin model is used to synchronize the student's learning emotion and concentration comprehensive score. The learning emotion and concentration comprehensive score is obtained based on the learning behavior feature vector. The student's facial expression, voice, eye movement, and action data are collected and analyzed to generate the learning behavior feature vector. The teacher twin model and the student twin model are interconnected and updated in real time to dynamically reflect the interaction between teachers and students.

[0056] S2: Mode switching trigger mechanism, based on the comprehensive score of learning emotion and concentration, triggers the interactive instructions in the teacher twin model and adjusts the teaching mode. When the learning emotion and concentration scores are detected to be lower than the preset threshold, the teacher-led teaching mode is triggered; when the learning emotion and concentration scores are higher than the preset threshold, the student-led interactive learning and teaching mode is triggered, giving students more authority to actively provide feedback;

[0057] S3: In the student-led interactive learning mode, the teacher twin model will generate guided discussion instructions to guide students to conduct in-depth learning exploration and discussion. At the same time, it will provide adaptive learning paths in the student twin model to guide students to complete self-study tasks according to their own progress.

[0058] The construction of the teacher twin model in S1 includes:

[0059] Teaching content synchronization is used to store and manage teachers' teaching content, including course outlines, chapters, and knowledge points, and synchronize the teaching content display progress to the student twin model in real time to ensure that what students learn is consistent with what the teacher teaches;

[0060] Teaching progress tracking is used to record the real-time progress of teachers' teaching, track the speed of content explanation and teaching rhythm in class, adjust the speed of teaching content according to progress data, and issue a prompt to the teacher when it is detected that the student's understanding progress is lower than the teaching progress, so as to assist the teacher in adjusting the teaching speed;

[0061] Interactive instruction generation is used to generate interactive instructions based on classroom situations, including questions, discussion topics, and task instructions. Interactive instructions are triggered based on the teacher's teaching needs or the student's learning status, and are transmitted to the student twin model in real time. Student feedback is also provided to the teacher to help the teacher adjust the content and frequency of the interaction as appropriate.

[0062] Teacher Twin Model T model Expressed as: T model=(C(t), P(t), I(t)), where C(t) is the teaching content matrix, including the course outline, chapters, and knowledge point information; P(t) is the teacher's teaching progress vector, recording the advancement speed of the teaching content; I(t) is the teacher's interactive instruction vector, representing the classroom interactive instructions of different teaching modes; and t is the time step;

[0063] The interactive instruction vector is represented as a conditional trigger function: I(t) = f(R(t), T, D), where T is the preset threshold used to set the trigger frequency of the interactive instruction, D is the preset teaching mode, including the teacher-led teaching mode and the student-led interactive learning teaching mode, f is the rule function used to determine whether to trigger the interactive instruction, and R(t) is the comprehensive score of learning emotion and concentration, which is used to dynamically adjust the teaching mode. If R(t) is lower than the preset threshold, a new instruction is triggered, and I(t) is updated to the new instruction vector and synchronized to the student twin model.

[0064] Teaching content matrix C(t)={c1,c2,...,c n}, where c n Indicates different units of course content (outline, knowledge points, examples), each c n The progress status of the unit is synchronized with the student twin model at time t, so that the content of C(t) matches the student's learning path;

[0065] Teaching progress vector in, Indicates the teacher's understanding of content c at time t n The system continuously updates the progress information to adapt the teaching rhythm according to classroom feedback and triggers adjustment suggestions through logical conditions when the student's understanding progress is lower than the teacher's progress.

[0066] The generation of learning behavior feature vectors in S1 includes:

[0067] Multimodal data collection: including facial expression recognition, speech analysis, eye tracking, and motion capture, using cameras, microphones, eye trackers, and motion sensors to collect students' facial expressions, speech, eye movements, and body movement data in real time;

[0068] Multimodal data analysis: The collected data is fed into a pre-trained multi-branch neural network model to extract features from different data sources. Facial expression data is used to extract emotional features, voice data is used to analyze intonation and emotional fluctuations, eye movement data is used to determine students' visual focus, and motion data is used to identify students' engagement and activity status.

[0069] Learning behavior feature vector generation: Facial expression, voice, eye movement, and action features are integrated to generate a multi-dimensional learning behavior feature vector S(t);

[0070] Comprehensive score: Based on the weighted calculation of each feature in the learning behavior feature vector, a comprehensive score representing the student's learning mood and concentration is generated.

[0071] The multi-branch neural network model specifically includes:

[0072] Facial expression feature branch network: The input is a facial expression image sequence. The network structure includes multiple convolutional layers, batch normalization layers, pooling layers, and fully connected layers. The output is the facial expression feature vector E(t). This branch extracts the student's facial expression features to identify emotional changes.

[0073] Speech feature branch network: It inputs the feature sequence (Mel spectrum) of speech audio data and outputs the speech feature vector V(t). This branch is used to extract speech pitch and emotional fluctuation features.

[0074] Eye movement feature branch network: The input is an eye movement trajectory sequence or an eye movement image (such as a sequence of gaze point positions). The network structure includes a convolutional layer and a pooling layer. The output is an eye movement feature vector G(t). This branch is used to extract features related to visual concentration.

[0075] Motion feature branch network: input motion data sequence (such as skeleton points or accelerometer data) and output motion feature vector M(t);

[0076] Each eigenvector can be expressed as:

[0077] E(t)={e1,e2,...,e l};

[0078] V(t)={v1,v2,...,v j};

[0079] G(t)={g1,g2,...,g p};

[0080] M(t)={m1,m2,...m q};

[0081] Facial expression feature vector E(t): represents the emotional state of the student;

[0082] Speech feature vector V(t): represents the student's intonation and tone fluctuations;

[0083] Eye movement feature vector G(t): represents the student's visual concentration;

[0084] Motion feature vector M(t): represents the student’s physical activity and participation.

[0085] The output feature vectors of each branch are concatenated to obtain the learning behavior feature vector S(t):

[0086] S(t)=[E(t); V(t); G(t); M(t)].

[0087] The comprehensive score of learning emotion and concentration R(t) is obtained by weighted calculation of the learning behavior feature vector S(t). The calculation expression of R(t) is:

[0088] Among them, R(t) is the comprehensive score of the students, which is used to reflect the overall state of learning mood and concentration, s i The i-th eigenvalue in the learning behavior eigenvector, w i For each eigenvalue s i The weight of , which indicates the contribution of this feature to the comprehensive score, satisfies

[0089] Input R(t) into the conditional trigger function, and standardize the comprehensive score R(t) to a range between 0 and 1. The preset threshold T is set to a percentage, 70%, or 0.7. When R(t) exceeds 0.7, it is considered that the student's learning mood and concentration are good, and the student-led interaction mode can be entered. When it is lower than 0.7, the teacher-led mode is triggered so that the teacher can provide more help. It can also be differentiated for different classes, for example, Class A is set to 0.7 and Class B is set to 0.5.

[0090] The adaptive learning path provided in the student twin model in S3 is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning status. Specifically, it includes:

[0091] S31, state definition: The state is defined as the learning behavior feature vector S(t), including learning emotions and concentration. The state is continuously updated over time;

[0092] S32, Action Definition: Select appropriate learning tasks or discussion guides based on the student's current state. Learning tasks and discussion guides are considered actions A(t), which are specific learning activities assigned to students. The action set A includes different learning units or topics. The system can select appropriate learning tasks or discussion guides based on the student's state.

[0093] S33, Reward Mechanism: Rewards are given according to the changes in the student's status. When the comprehensive score of the student's learning mood and concentration is high, a high reward is given; on the contrary, the reward is low. Implementation, reward function The reward is calculated based on the changes in the student's learning status. The reward reflects the student's learning effect. When the student shows high understanding, participation or positive feedback, the reward is high; when the student shows low interest or has learning difficulties, the reward is low. The reward function is defined as:

[0094] Among them, K t+1 and P t+1 is the knowledge mastery and interest preference after performing the action, D represents the difficulty of the learning task. When the difficulty is higher, the reward is reduced; w1, w2, and w3 are the weights of different factors. The purpose of the reward is to motivate students to improve their learning effect;

[0095] S34, design strategy π: Strategy π(S(t)) is a strategy for selecting the best action from the current student state S(t). The deep Q network is used to approximate the strategy to maximize the long-term reward. After the deep Q network training is completed, given the student's current learning state S(t), the deep Q network selects the optimal action A(t) to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real time according to the student's understanding level and interest preferences to ensure that students can maximize their learning effects through in-depth learning exploration and discussion. It provides personalized learning paths to guide students to explore more in-depth content based on their own learning progress and feedback, and promote students' active learning and independent exploration.

[0096] The deep Q-network specifically includes: initializing a Q-network to estimate the value of each action taken by the student in the current state, that is, the Q-value; creating an experience replay buffer that stores the state, action, reward, and next state of the student during the interaction with the learning content; the experience replay buffer is used to provide randomly sampled data for subsequent training; updating the Q-value in the Q-network through feedback obtained from each interaction with the student to optimize the strategy for selecting learning tasks; the goal of Q-value updating is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q-value of the next state; each time the student interacts, the learning task is selected based on the current state, using the ε-greedy strategy for selection, which is: most of the time, the optimal action is selected (maximizing the Q-value), and occasionally random actions are selected for exploration; after executing the selected learning task, the state will change, and feedback is provided based on the new state and reward value; historical experience is randomly sampled from the experience replay buffer, and the Q-network is trained with historical experience; the network parameters are updated by minimizing the loss function; repeated iterations allow the Q-network to continuously learn and optimize, selecting the optimal learning task in each learning state, thereby maximizing the student's long-term learning effect.

[0097] In the student-led interactive learning model, not only learning tasks are provided, but also guided discussion instructions are generated to encourage students to think deeply and actively participate. These tasks are customized according to students' learning progress, interest preferences and emotional feedback to enhance students' sense of participation and learning motivation.

[0098] Through continuously optimized learning paths, students can be guided from shallow understanding to deep knowledge exploration, stimulate students' active feedback and discussion, and ultimately help students complete learning tasks at their own pace, thus achieving a personalized and in-depth learning process.

[0099] The specific calculation of Deep Q Network (DQN) is as follows:

[0100] 1. Initialize the Q network: Initialize a neural network Q(S(t), A(t); θ), which is used to estimate the Q value function with parameter θ.

[0101] 2. Experience replay buffer: Establish experience replay buffer D to store the experience of student state transition To allow for random sampling during training.

[0102] 3. Q value update formula: Use the following Q value update formula for network training:

[0103]

[0104] Among them, γ is the discount factor, which represents the decay of future rewards, and A′ represents the possible next action.

[0105] 4. Select action (learning task): Under the student's current state S(t), use the greedy strategy to select action A(t):

[0106] Randomly select actions with probability to ensure exploration;

[0107] The action that maximizes Q(S(t), A) is selected with probability 1 to ensure development.

[0108] 5. Action execution and feedback: Execute the selected learning task, and the student twin model feedbacks the new learning state S(t+1) and the corresponding reward

[0109] 6. Network training: Randomly sample a batch of data from the experience buffer The Q network is trained using mean squared error as the loss function.

[0110] 7. Iterative update: Repeat the above steps 1-6 so that the Q network can continuously optimize the strategy of selecting learning tasks, so that the student can obtain the maximum long-term reward in each state, that is, the optimal learning path.

[0111] like Figure 2 As shown, the digital twin-based teaching information processing system is used to implement the above-mentioned digital twin-based teaching information processing method, including the following modules:

[0112] The teacher twin model is used to synchronize the teacher's teaching content, teaching progress and interactive instructions in real time, dynamically reflecting the teacher's teaching behavior and teaching adjustments;

[0113] The student twin model is used to synchronize students' learning emotions, concentration, and learning progress in real time. It generates learning behavior feature vectors based on multimodal data (including students' facial expressions, voice, eye movements, and movement data) to determine students' learning status.

[0114] The mode switching trigger module is used to trigger different teaching modes based on the comprehensive scores of students' learning mood and concentration, including teacher-led teaching mode and student-led interactive learning mode;

[0115] The collaborative learning instruction generation module generates guided discussion instructions based on students' learning status in a student-led interactive learning mode;

[0116] The adaptive learning path module, based on the deep Q network, builds personalized learning paths to guide students to complete self-study tasks.

[0117] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.

[0118] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. The teaching information processing method based on digital twin is characterized by: The following steps are involved: S1: Construct a digital twin model of teachers and students, including a teacher twin model and a student twin model. The teacher twin model is used to synchronize the teacher's teaching content, teaching progress, and interactive instructions. The student twin model is used to synchronize the student's learning emotion and concentration comprehensive score. The learning emotion and concentration comprehensive score is obtained based on the learning behavior feature vector. The student's facial expression, voice, eye movement, and action data are collected and analyzed to generate a learning behavior feature vector. The teacher twin model and the student twin model are interconnected and updated in real time to dynamically reflect the interaction between teachers and students. S2: Mode switching trigger mechanism, based on the comprehensive score of learning emotion and concentration, triggers the interactive instructions in the teacher twin model and adjusts the teaching mode. When the learning emotion and concentration scores are detected to be lower than the preset threshold, the teacher-led teaching mode is triggered; when the learning emotion and concentration scores are higher than the preset threshold, the student-led interactive learning and teaching mode is triggered; S3: In the student-led interactive learning mode, the teacher twin model will generate guided discussion instructions to guide students to conduct in-depth learning exploration and discussion, while providing adaptive learning paths in the student twin model; The adaptive learning path provided in the student twin model is dynamically planned based on reinforcement learning. Reinforcement learning adjusts the learning path through real-time feedback on the student's learning status, specifically including: S31, state definition: define the state as a learning behavior feature vector , including learning mood, concentration, and status that is continuously updated over time; S32, Action Definition: Select learning tasks or discussion guidance based on the student's current status. Learning tasks and discussion guidance are considered actions. , are specific learning activities arranged for students; S33, Reward Mechanism: Rewards are given according to the changes in the student's status. When the comprehensive score of the student's learning mood and concentration is high, a high reward is given; on the contrary, the reward is low. Implementation, reward function Calculated based on changes in students’ learning status, rewards reflect students’ learning outcomes; S34, Design Strategy :Strategy Is from current student status The strategy of selecting the best action in the learning process is approximated by a deep Q network to maximize the long-term reward. After the deep Q network is trained, the optimal action A(t) is selected by the deep Q network given the student's current learning state S(t) to guide the student to complete the next learning unit. As the learning state is updated, the learning path is adjusted in real time according to the student's understanding level and interest preferences. The deep Q-network specifically includes: initializing a Q-network to estimate the value of each action taken by the student in the current state, i.e., the Q-value; creating an experience replay buffer that stores the state, action, reward, and next state of the student during the interaction with the learning content; the experience replay buffer is used to provide randomly sampled data for subsequent training; and updating the Q-value in the Q-network based on the feedback obtained from each interaction with the student to optimize the strategy for selecting learning tasks; The goal of the Q-value update is to calculate the value of the action in the current state by adding the current reward to the maximum expected Q-value of the next state. During each student interaction, a learning task is selected based on the current state and the ε-greedy strategy is used for selection. After executing the selected learning task, the state will change. Feedback is given based on the new state and reward value. Historical experience is randomly sampled from the experience replay buffer, and the Q network is trained with historical experience. The network parameters are updated by minimizing the loss function, and the iteration is repeated so that the Q network continues to learn and optimize.

2. The teaching information processing method based on digital twin according to claim 1 is characterized in that: The construction of the teacher twin model in S1 includes: Teaching content synchronization is used to store and manage teachers' teaching content, including course outlines, chapters, and knowledge points, and synchronize the teaching content display progress to the student twin model in real time; Teaching progress tracking is used to record the real-time progress of teachers' teaching, track the speed of content explanation and teaching rhythm in class, and adjust the advancement speed of teaching content based on progress data; Interactive instruction generation is used to generate interactive instructions based on classroom situations, including questions, discussion topics, and task instructions. Interactive instructions are triggered based on the teacher's teaching needs or the student's learning status and transmitted to the student twin model in real time.

3. The teaching information processing method based on digital twin according to claim 2 is characterized in that: The teacher twin model Expressed as: ,in, It is the teaching content matrix, including course outline, chapters, knowledge points information, It is the teacher's teaching progress vector, which records the advancement speed of the teaching content. is the teacher's interactive instruction vector, which represents the classroom interactive instructions of different teaching modes. is the time step; The interactive instruction vector is represented as a conditional trigger function: ,in, It is a preset threshold used to set the trigger frequency of interactive instructions. It is a preset teaching mode, including teacher-led teaching mode and student-led interactive learning teaching mode. Represents a rule function, used to determine whether to trigger an interactive instruction. Comprehensive score of learning mood and concentration, used to dynamically adjust teaching mode. If the value is lower than the preset threshold, a new instruction will be triggered. Update to the new instruction vector and synchronize to the student twin model; The course content matrix ,in, Represents different units of course content, each The progress status of the unit at time Synchronize with the student twin model Matching content to students’ learning paths; The teaching progress vector ,in, Indicates time When the teacher is on the content speed of explanation.

4. The teaching information processing method based on digital twin according to claim 3 is characterized in that: The generation of the learning behavior feature vector in S1 includes: Multimodal data collection: including facial expression recognition, speech analysis, eye tracking, and motion capture, using cameras, microphones, eye trackers, and motion sensors to collect students' facial expressions, speech, eye movements, and body movement data in real time; Multimodal data analysis: The collected data is fed into a pre-trained multi-branch neural network model to extract features from different data sources. Facial expression data is used to extract emotional features, voice data is used to analyze intonation and emotional fluctuations, eye movement data is used to determine students' visual focus, and motion data is used to identify students' engagement and activity status. Learning behavior feature vector generation: Fusing facial expressions, voice, eye movements, and action features to generate multi-dimensional learning behavior feature vectors ; Comprehensive score: Based on the weighted calculation of each feature in the learning behavior feature vector, a comprehensive score representing the student's learning mood and concentration is generated.

5. The teaching information processing method based on digital twin according to claim 4 is characterized in that: The multi-branch neural network model specifically includes: Facial expression feature branch network: input facial expression image sequence, the network structure includes multiple convolutional layers, batch normalization layers, pooling layers and fully connected layers, and the output is facial expression feature vector ; Speech feature branch network: inputs the feature sequence of speech audio data and outputs the speech feature vector ; Eye movement feature branch network: input eye movement trajectory sequence or eye movement image, the network structure includes convolution layer and pooling layer, and the output is eye movement feature vector ; Action feature branch network: input action data sequence output as action feature vector ; The output feature vectors of each branch are spliced ​​to obtain the learning behavior feature vector : 。 6. The teaching information processing method based on digital twin according to claim 5 is characterized in that: The comprehensive score of learning mood and concentration By learning the behavioral feature vector Perform weighted calculation to obtain The calculation expression is: ,in, A comprehensive score for students, used to reflect their overall learning mood and concentration. The first eigenvalues, For each eigenvalue The weight of , which indicates the contribution of this feature to the comprehensive score, satisfies .

7. A digital twin-based teaching information processing system, used to implement the digital twin-based teaching information processing method according to any one of claims 1 to 6, characterized in that: Includes the following modules: The teacher twin model is used to synchronize the teacher's teaching content, teaching progress and interactive instructions in real time, dynamically reflecting the teacher's teaching behavior and teaching adjustments; The student twin model is used to synchronize students' learning emotions, concentration, and learning progress in real time, and to generate learning behavior feature vectors based on multimodal data to determine students' learning status; The mode switching trigger module is used to trigger different teaching modes based on the comprehensive scores of students' learning mood and concentration, including teacher-led teaching mode and student-led interactive learning mode; The collaborative learning instruction generation module generates guided discussion instructions based on students' learning status in a student-led interactive learning mode; The adaptive learning path module, based on the deep Q network, builds personalized learning paths to guide students to complete self-study tasks.

Citation Information

Patent Citations

  • Teaching information processing method and system based on digital twinning

    CN117252047A

  • Management teaching information processing method and system based on digital twinning

    CN119151743A