Intelligent teaching management method and device and electronic equipment

By constructing a nonlinear learning state evolution model and optimal control principle, combining lag compensation and local game models, and dynamically adjusting teaching strategies, the problems of inaccurate learning state identification and delayed strategy response in personalized teaching are solved, and efficient and reasonable allocation of teaching resources and personalized learning are achieved.

CN120634809APending Publication Date: 2025-09-12BEIJING HUACAN ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510960835.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing personalized teaching systems lack accuracy in learning status identification, strategy response, and multi-student scenarios. Strategy updates lag and lack coordination mechanisms, leading to resource allocation conflicts.

Method used

By acquiring multimodal learning behavior data, constructing a nonlinear learning state evolution model, using the optimal control principle to optimize the teaching strategy, and combining lag compensation and local game models for strategy coordination, the teaching content and task difficulty can be dynamically adjusted.

Benefits of technology

It achieves high-precision dynamic modeling of students' learning status, solves the problems of inaccurate learning status identification and delayed strategy response, optimizes the rational allocation of teaching resources, and reduces control conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634809A_ABST
    Figure CN120634809A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of personalized teaching, and discloses an intelligent teaching management method and device and electronic equipment, and the method comprises the steps: collecting learning behavior data, constructing a state vector, building an evolution model, optimizing a control strategy, and dynamically pushing teaching content. The device comprises a data acquisition module, a learning state construction module, a learning state evolution model module, an optimization target construction module, a control strategy optimization module and the like, and supports strategy closed-loop adjustment and multi-student coordination control. The electronic equipment comprises a processor, a communication module, an interactive interface and other structures, and is used for executing the method steps and realizing strategy display and data feedback. According to the technical scheme of constructing the nonlinear learning state evolution model through the multi-modal behavior data, the high-precision modeling effect on the dynamic change of the learning state of the student is achieved, and the problems that the state representation is single and the change trend is difficult to capture are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of personalized teaching technology, and in particular to a smart teaching management method, device and electronic equipment. Background Art

[0002] With the continuous advancement of intelligent education technology, more and more teaching systems are beginning to introduce data-driven methods to track and provide feedback on students' learning process. Some systems use behavioral data, such as answer status and study time, to build preliminary status assessment models, and use this data to push and adjust teaching content. In terms of strategy formulation, common practices are mostly preset recommendations or dynamic adjustments based on rules. At the same time, some platforms have begun to experiment with building personalized strategy templates based on student characteristics to improve content matching and teaching response speed.

[0003] While these approaches have improved the adaptability of teaching content to a certain extent, they still have limitations in state recognition, strategy evolution, and group coordination. For example, learning states often rely on static labels or single indicators, making it difficult to reflect complex dimensions such as emotion and attention. Strategy updates lack a closed-loop regulation mechanism and are significantly affected by delayed feedback, making it difficult to balance real-time performance and stability. In multi-student scenarios, resource allocation and individual strategies often lack coordinated modeling, which can easily lead to control conflicts. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a smart teaching management method, device and electronic equipment, which solve the problems of inaccurate learning status recognition, delayed strategy control response and difficulty in coordinating conflicts among multiple students in personalized teaching.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A smart teaching management method, comprising the following steps: S1. Obtain multimodal learning behavior data of multiple students and construct corresponding learning state vectors; S2. Based on the learning state vector, establish a nonlinear learning state evolution model for each student, where the model describes the dynamic process of the student's learning state changing over time; S3. Based on the nonlinear learning state evolution model of each student, construct an optimization objective function that represents the student's learning effect and teaching control cost; S4. Utilizing the optimal control principle and based on the optimization objective function, optimizing the teaching control strategy to obtain a personalized teaching control trajectory for each student, and making the trajectory capable of maximizing the student's learning effect; S5. Based on the time lag characteristics in the feedback process, a hysteresis compensation model is constructed to predict the feedback delay and adjust the control strategy; S6. Considering the impact of strategic interactions among student groups, the control strategies of multiple students are coordinated based on the local game model to optimize the overall learning effect of the group; S7. Based on the final personalized optimal teaching control strategy, the teaching content, task type and difficulty are dynamically pushed and adjusted to ensure that each student's learning process can be carried out efficiently and personalized.

[0006] Preferably, in step S1, the multimodal learning behavior data includes the student's facial expression, eye concentration, voice output, mouse trajectory and answering time, and the learning state vector includes the student's attention level, knowledge mastery and emotional state.

[0007] Preferably, in step S2, the nonlinear learning state evolution model is described by a differential equation, which takes into account multiple learning state factors of the students and models the students' attention, emotions and knowledge mastery over time.

[0008] Preferably, in step S3, the optimization objective function includes a comprehensive cost term of the student's attention level, knowledge mastery, and teaching cost, wherein the cost term is: ; in, For students The comprehensive cost function value of is the weight coefficient of the attention bias term; For students attention level; is the weight coefficient of the knowledge mastery bias term; For students knowledge mastery; is the weight coefficient of the teaching cost item; For students the cost of teaching.

[0009] Preferably, in step S4, the optimization teaching control strategy is implemented by solving the Euler-Lagrange equation in the variational method, wherein the Euler-Lagrange equation is given by the following form: ; in, is the Lagrangian function; For control variables ; For control variables About Time The derivative of , that is, the rate of change of the control variable; is the Lagrangian function For control variables The partial derivative of is the Lagrangian function Derivative of the control variable The partial derivative of is the derivative of the above partial derivative with respect to time.

[0010] Preferably, in step S5, the time lag characteristics include feedback lag, control strategy implementation lag, system delay, student response lag and strategy adjustment lag; The hysteresis compensation model predicts the feedback lag of the control strategy by constructing a delayed feedback function, and adjusts the control strategy according to the prediction result, so that the adjusted control strategy satisfies the following relationship: ; in, For the moment Control variables when The value of To control the time lag, that is, the delay time of the system response; For control variables In the lag time The value of when is the control change rate function within the time period; From time arrive The integral value of the control rate of change.

[0011] Preferably, in step S6, the local game model describes the students' learning strategies by introducing individual difference parameters of students, so its cost function for: ; in, For students The multi-round comprehensive cost function value of ; is the discrete time step index; is the total number of time steps; student In the The attention level at each time step; For students In the The knowledge mastery of each time step; For students In the Control strategy input for time steps; For students Control strategy input for other students except students The average attention level of other students; is the weight coefficient of the attention bias term; is the weight coefficient of the knowledge mastery bias term; For students The control strategy consistency adjustment coefficient for students The attention contrast adjustment coefficient.

[0012] Preferably, in step S7, the process of dynamically pushing and adjusting teaching content, task type and difficulty includes the following operations: Real-time monitoring of students' learning status, including their attention level, knowledge mastery, emotional state, and learning progress; Dynamically adjust the difficulty of learning tasks based on students' real-time learning status and select appropriate task types based on students' ability levels; Adjust the order and time of learning tasks based on students' learning progress and emotional state to ensure that the assignment of tasks matches students' personalized learning progress; Dynamically push personalized teaching content and adjust the type, difficulty and order of tasks based on students' real-time feedback.

[0013] A smart teaching management device, comprising: Data collection module, used to obtain multimodal learning behavior data of multiple students; A learning state construction module, connected to the data acquisition module, for constructing a corresponding learning state vector based on the learning behavior data; a learning state evolution model module, connected to the learning state construction module, for establishing a nonlinear learning state evolution model for each student based on the learning state vector; The optimization target construction module is connected to the learning state evolution model module and is used to construct the optimization target function that represents the student learning effect and the teaching control cost; A control strategy optimization module is connected to the optimization target construction module and is used to optimize the teaching control strategy based on the optimization target function using the optimal control principle to obtain a personalized teaching control trajectory for each student; A hysteresis compensation module is connected to the control strategy optimization module and is used to construct a hysteresis compensation model and adjust the control strategy according to the feedback lag characteristics; A game coordination module, which is connected to the hysteresis compensation module and is used to coordinate the control strategies of multiple students based on the local game model; The dynamic push module is connected to the game coordination module and is used to dynamically push and adjust the teaching content, task type and difficulty according to the optimized control strategy.

[0014] An electronic device, comprising: a processor, configured to execute a computer program stored in the electronic device; Memory, used to store programs and related data required for execution; Communication module, used for data transmission with external devices, supporting data collection and teaching content push; Display module, used to display the teaching content, task type and difficulty adjustment information generated by the dynamic push module; The interactive interface is used to interact with students, receive students' feedback and transmit it to the data collection module.

[0015] The present invention provides a smart teaching management method, device, and electronic device. It has the following beneficial effects: 1. This invention utilizes multimodal behavioral data to construct a nonlinear learning state evolution model, achieving high-precision modeling of the dynamic changes in students' learning states. Compared to existing approaches that use staged performance or static labels for modeling, this approach solves the problem of single-state representation and difficulty capturing changing trends.

[0016] 2. This invention utilizes the principle of optimal control, combined with state feedback and hysteresis compensation mechanisms, to dynamically adjust and modify the teaching strategy, achieving continuous closed-loop control. This approach, unlike traditional fixed or semi-fixed strategy delivery models, effectively overcomes the problem of strategy deviation caused by feedback delay.

[0017] 3. This invention utilizes a local game model for multi-student strategy coordination, enabling rational allocation of limited teaching resources. Compared to existing individual optimization approaches that lack interactive influence modeling, it addresses the issues of control strategy conflict and intervention imbalance in collaborative tasks.

[0018] 4. By building a feedback evolution mechanism and an individualized parameter adjustment process, this invention achieves multi-round iterative optimization of the control strategy. Compared with existing solutions that rely on manual adjustment or pre-set rules, this solves the technical shortcomings of strategy adjustment that lack data-driven control and delayed response. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of the method steps of the present invention; Figure 2 This is a diagram of the device module architecture of the present invention; Figure 3 Schematic diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] Example 1: Please see the attached Figure 1 , an embodiment of the present invention provides a smart teaching management method, comprising the following steps: S1. Obtain multimodal learning behavior data of multiple students and construct corresponding learning state vectors; In step S1, the multimodal learning behavior data includes the student's facial expression, eye focus, voice output, mouse trajectory, and answer time, and the learning state vector includes the student's attention level, knowledge mastery, and emotional state; Specifically, in this embodiment, the system includes a data acquisition unit, a preprocessing unit, and a learning state construction unit. The data acquisition unit is connected to the preprocessing unit for signal processing, and the preprocessing unit is connected to the learning state construction unit.

[0022] The data acquisition unit includes an image acquisition submodule, a voice acquisition submodule, an interaction behavior recording submodule, and a response time recording submodule. Each submodule is connected to the preprocessing unit via a bus or wireless communication.

[0023] The image acquisition submodule collects facial image data through a camera and estimates the expression feature values ​​in combination with a face recognition algorithm. Furthermore, the module combines a convolutional neural network (CNN) with an expression classifier to output the probability distribution of emotion labels.

[0024] The voice acquisition submodule records the students' voice signals during the answering process through a microphone, and uses the voice activity detection (VAD) algorithm to extract valid voice segments, and combines the changes in intonation to judge the cognitive load status.

[0025] The interactive behavior recording submodule records mouse trajectory data, including position coordinate sequence, click events, pause duration, etc., and uses time series analysis method to extract dynamic behavior characteristics such as mouse speed and trajectory curvature.

[0026] The response time recording submodule times the answer time for each question and calculates the change ratio of its deviation from the student's historical average answer time.

[0027] All collected multimodal raw data is uniformly encoded by the preprocessing unit, which includes normalization, feature fusion, and time synchronization. This process ensures data consistency along the time axis while compressing redundant features.

[0028] Subsequently, the learning state construction unit constructs the learning state vector based on the uniformly preprocessed feature sequence. , the vector form is as follows: ; in, For students The learning state vector within a certain time period; It is an attention level indicator based on the degree of eye concentration and mouse operation activity; It is an indicator of knowledge mastery, inferred through response time, correct answer rate and speech fluency; It is an indicator of emotional state, and the emotional tendency is determined by combining facial expression recognition and voice intonation.

[0029] Attention level indicator The calculation of adopts the weighted fusion model, specifically: ; in, is the gaze concentration, which is taken from the gaze stability index estimated in the image acquisition submodule; Mouse activity index, including the total length of the track and the normalized value of the click frequency; , is the experience weight coefficient.

[0030] Knowledge mastery Estimated by the following function: ; in, is the correct answer rate; The answer time for this question is: is the average time it takes students to answer history questions; score the expressive fluency of speech output; is the weighting parameter; is the normalization constant.

[0031] emotional state Through the construction of multimodal fusion model, the decision-making layer weighting mechanism is used to: ; in, The negative emotion probability value output by the expression recognition module; It is an indicator of the proportion of low and high active intonation in speech emotion recognition; is the weighting coefficient.

[0032] The above state vector Updated over time to form a state sequence , which will serve as the input basis for subsequent evolutionary modeling.

[0033] Structured data is transmitted between the data acquisition unit, preprocessing unit and state construction unit through a logical interface. All modules can be embedded and integrated in the teaching platform, with real-time and traceability.

[0034] This step method can accurately represent the current learning status of individual students, provide basic input for subsequent personalized control and strategy optimization, and has high-resolution and high-adaptability technical effects.

[0035] S2. Based on the learning state vector, a nonlinear learning state evolution model is established for each student. This model describes the dynamic process of the student's learning state changing over time. In step S2, the nonlinear learning state evolution model is described by a differential equation, which considers the student's multiple learning state factors and models the student's attention, emotion, and knowledge mastery over time; Specifically, in this embodiment, the system includes a state modeling unit and a parameter updating unit. The state modeling unit is connected to the learning state construction unit and is used to construct a nonlinear learning state evolution model of an individual student based on a state vector sequence.

[0036] First, a learning state evolution model is established to describe the dynamic changes of students' learning state over time. The model uses a nonlinear differential equation to express the trajectory of student state evolution, which is defined as follows: ; in, For students At the moment The learning state vector includes attention level, knowledge mastery and emotional state; is the teaching control strategy for the corresponding moment; is the parameter set of individual differences among students; is the state evolution function with nonlinear characteristics.

[0037] State evolution function The expression of contains multiple influencing factors and can be obtained based on empirical modeling or data-driven training. In a preferred implementation, the following combination structure is adopted: ; in, is the individual difference adjustment matrix; is the activation function, preferably ReLU or tanh; The state input weight matrix and the control input weight matrix; is the bias term.

[0038] The above model can be trained by time series supervised learning, using the state observation sequence and control input sequence As input, minimize the error between the predicted state and the true state; Among them, in the state observation sequence, Indicates the Students in the time steps The learning state vector at time t; Indicates the Sampling time points; Indicates the number of time steps in the total duration of state observation; In the control input sequence, Indicates in The time point is applied to Teaching control input for each student.

[0039] The parameter update part uses gradient descent or Adam optimizer to update the parameter set ,in, For students Individual adjustment matrix; Enter the weight matrix for the state vector; Input weight matrix for control; is the hidden layer bias vector.

[0040] The objective function is defined as: ; in, is the model prediction state, The actual observation state.

[0041] In order to enhance the model's ability to model time dependencies, long short-term memory networks (LSTM) or gated recurrent units (GRU) can be used to replace traditional fully connected structures to improve the accuracy of state sequence modeling.

[0042] The state evolution model output by the state modeling unit is connected to the control strategy optimization module in differential or continuous form as the basis for subsequent objective function construction and optimal control trajectory calculation.

[0043] In the system, the state modeling unit is connected to the control strategy module through the API interface. The model structure is stored in the database and updated in real time to ensure that the model dynamically adapts to the student's current behavioral status.

[0044] The modeling process in this step can achieve high-precision characterization of the changing trend of student status, and has the technical effects of high state prediction accuracy and strong adjustability of model parameters, providing a dynamic basis for subsequent control strategy optimization.

[0045] S3. Based on the nonlinear learning state evolution model of each student, construct an optimization objective function that represents the student's learning effect and teaching control cost; In step S3, the optimization objective function includes a comprehensive cost term of the student's attention level, knowledge mastery, and teaching cost, where the cost term is: ; in, For students The comprehensive cost function value of is the weight coefficient of the attention bias term; For students attention level; is the weight coefficient of the knowledge mastery bias term; For students knowledge mastery; is the weight coefficient of the teaching cost item; For students The cost of teaching; Specifically, in this embodiment, step S3 includes an objective function construction part, which connects the state modeling part and the strategy generation part, and its function is to construct the optimization goal of the individualized teaching strategy based on the output of the state evolution model.

[0046] First, based on the dynamic evolution trajectory of the student's learning state, a difference function between the expected learning performance and the actual state is constructed to measure the degree of optimization of the teaching control strategy. Among them, the cost term is: ; in, For students The comprehensive cost function value of is the weight coefficient of the attention bias term; For students attention level; is the weight coefficient of the knowledge mastery bias term; For students knowledge mastery; is the weight coefficient of the teaching cost item; For students the cost of teaching.

[0047] The corresponding objective function includes two weighted terms: The first item measures the deviation between the student’s current state and the target state; The second constraint controls the energy of the input to avoid excessive control or frequent intervention that may lead to a decline in the teaching experience.

[0048] Furthermore, the ideal state trajectory It can be defined as the following function form: ; in, It is an estimate of the average learning status of students in this category; To gradually increase the amplitude coefficient; It is a time evolution function, which can be a linear growth, sigmoid or exponential function to control the target improvement speed.

[0049] This construction method ensures that the target state grows smoothly over time, allowing students to adapt gradually rather than making sudden jumps.

[0050] The objective function construction unit converts the predicted state provided by the state modeling module into and preset target status Align, calculate the current control strategy Contribution to goal achievement.

[0051] This objective function is suitable for use as a loss function in the gradient descent or dynamic programming method in the subsequent control strategy optimization process, and can also be used as the negative form of the reward function in reinforcement learning.

[0052] The control objective function calculation process is executed in real time in the inference engine and supports batch parallel optimization, which can adapt to the needs of multi-student heterogeneous state input and personalized goal setting.

[0053] The optimization objective function construction method in this step can realize the dynamic adaptation of the teaching control strategy to the individual state, and has technical effects such as controllable deviation, adjustable energy consumption, and plastic target, providing a quantifiable target basis for the subsequent teaching strategy generation.

[0054] S4. Utilize the optimal control principle and optimize the teaching control strategy based on the optimization objective function to obtain a personalized teaching control trajectory for each student, and make the trajectory maximize the student's learning effect; In step S4, the optimization teaching control strategy is implemented by solving the Euler-Lagrange equation in the variational method, where the Euler-Lagrange equation is given by the following form: ; in, is the Lagrangian function; For control variables ; For control variables About Time The derivative of , that is, the rate of change of the control variable; is the Lagrangian function For control variables The partial derivative of is the Lagrangian function Derivative of the control variable The partial derivative of is the derivative of the above partial derivative with respect to time; Specifically, in this embodiment, step S4 includes a strategy optimization part, which connects the objective function construction part and the state modeling part to solve the optimal teaching control input so that the student's learning state evolves towards a preset target trajectory.

[0055] First, based on the state evolution model provided by the state modeling department ; Construct an objective function to minimize state deviation and control cost ; in, For students At the moment The state vector of , with a dimension of 3, represents attention, mastery, and emotion respectively; is the target state trajectory; Input for teaching control, including prompt intensity, question type difficulty, etc. It is the control cost weight coefficient used to balance the control intensity and status improvement goals.

[0056] Subsequently, the strategy optimization department introduced adjoint variables based on the variational principle and constructed a constrained Lagrangian functional: ; in, Indicates the objects at time The actual state vector.

[0057] Indicates the objects at time The desired or target state vector.

[0058] It represents the square of the state error, which is used to measure the distance between the current state of the system and the target state. It is the most core performance item in optimization.

[0059] Represents the control input vector, corresponding to the objects at time exert control.

[0060] To control the cost item, It is the energy consumption weight coefficient used to balance control efficiency and control intensity.

[0061] is the Lagrange multiplier (also called adjoint variable), which is used to model the dynamic constraints of the system.

[0062] is the residual (or default term) of the state evolution, where For the The system dynamics function of an object, the parameter is the current state , control input and model parameters .

[0063] In the above functional, the third term introduces state evolution constraints.

[0064] The optimality of the control variables must satisfy the Euler-Lagrange condition, which can be expressed as: ; in, is the Lagrangian function; For control variables ; For control variables About Time The derivative of , that is, the rate of change of the control variable; is the Lagrangian function For control variables The partial derivative of is the Lagrangian function Derivative of the control variable The partial derivative of is the derivative of the above partial derivative with respect to time.

[0065] The strategy optimization department uses a three-stage algorithm of "forward-backward-update" to solve the problem. First, the state trajectory is forward integrated. , and then reversely integrate the adjoint trajectory , and finally update according to the Euler-Lagrange condition .

[0066] The error threshold is set during the entire optimization process and the maximum number of iterations; terminates when the following conditions are met: ; in, Indicates the Round iterative middle school students The control cost function value of Indicates the The objective function value in the round iteration; It is the preset error convergence threshold, used to control the solution accuracy; Represents the Euclidean norm, which is used to measure the absolute difference between two consecutive rounds of cost functions The optimal control sequence finally output It is transmitted to the teaching control department to achieve fine adjustment of the individual status of students.

[0067] The control sequence can be applied to teaching tasks such as recommended content push, personalized feedback frequency control, and interaction rhythm scheduling.

[0068] This implementation introduces the Euler-Lagrange optimality condition and combines it with the back-propagation mechanism of the adjoint variables to obtain the optimal control strategy in the entire time period while satisfying the state evolution constraints. It has the technical effects of strong convergence, complete theory, and fine control granularity.

[0069] S5. Based on the time lag characteristics in the feedback process, a hysteresis compensation model is constructed to predict the feedback delay and adjust the control strategy; In step S5, the time lag characteristics include feedback lag, control strategy implementation lag, system delay, student response lag, and strategy adjustment lag; The hysteresis compensation model predicts the feedback lag of the control strategy by constructing a delayed feedback function and adjusts the control strategy according to the prediction result so that the adjusted control strategy satisfies the following relationship: ; in, For the moment Control variables when The value of To control the time lag, that is, the delay time of the system response; For control variables In the lag time The value of when is the control change rate function within the time period; From time arrive The integral value of the control change rate between Specifically, in this embodiment, the system includes a teaching control execution unit and a state feedback update unit. The teaching control execution unit is connected to the strategy optimization unit and is used to execute teaching intervention operations based on the optimal control strategy obtained. The state feedback update unit is connected to the learning state construction unit and the strategy optimization unit and is used to dynamically modify the control strategy based on the state evolution results.

[0070] First, the teaching control execution unit receives the optimal control sequence output from the strategy optimization unit: ; in, Indicates students At time step The optimal teaching control input that should be applied at all times.

[0071] The teaching control execution unit includes a content presentation submodule, an interactive adjustment submodule and a feedback control submodule. The content presentation submodule is connected to the teaching database. In the task recommendation instructions, select the topic or video content.

[0072] The interaction adjustment submodule adjusts user interface parameters such as page dwell time and answer countdown length based on the rhythm control parameters in the control vector. The feedback control submodule selects graphic prompts, voice explanations, or emotional encouragement based on the feedback frequency and form set in the strategy.

[0073] The system simultaneously records the students' actual behavior data under control, including answering performance, interactive reactions, and emotional state, and updates it into a new learning state sequence: ; in, Indicates that after the control is executed, the student At the time point The actual observed learning state vector, whose components include attention level, knowledge mastery and emotion indicators.

[0074] After executing the strategy, the state feedback update unit receives the latest observed state sequence from the learning state construction unit, and differentiates it from the predicted state to construct a dynamic error function: ; in, Indicates students At the time point The learning state prediction error vector of It represents the actual observed learning state collected by the system through physiological data recognition, interactive behavior analysis, etc., including three dimensions: attention level, knowledge mastery and emotional stability; Represents the theoretical state predicted by the current control input and state evolution model; time step The discrete time point at which the current policy is executed.

[0075] The error function is used to characterize the deviation response of the strategy control input to the actual state and is the direct basis for feedback regulation.

[0076] When the error continues to deviate from the threshold, the system activates the feedback adjustment mechanism to dynamically correct the current control input to cope with the state evolution deviation.

[0077] In an exemplary regulation strategy, the adjustment of the control input introduces a time lag compensation and an error integration structure so that the modified control strategy satisfies the following relationship: ; in, For the moment The control variable The value of To control the time lag, that is, the delay time of the system response; For control variables In the lag time The value of when; is the control change rate function within the time period; From time arrive The integral value of the control rate of change.

[0078] This structure ensures that the control adjustment has historical smoothness, avoids control oscillation caused by instantaneous abnormal errors, and improves system stability.

[0079] This dynamic integral adjustment of the control strategy is implemented by the state feedback update unit and transmitted to the teaching control execution unit through the buffer control queue. The modules are connected through a standard API interface, enabling high-speed transmission of parameter streams and concurrent execution.

[0080] S6. Considering the impact of strategic interactions among student groups, the control strategies of multiple students are coordinated based on the local game model to optimize the overall learning effect of the group; In step S6, the local game model describes the students’ learning strategies by introducing individual difference parameters of students, so its cost function is for: ; in, For students The multi-round comprehensive cost function value of ; is the discrete time step index; is the total number of time steps; student In the The attention level at each time step; For students In the The knowledge mastery of each time step; For students In the Control strategy input for time steps; For students Control strategy input from other students; For students The average attention level of other students; is the weight coefficient of the attention bias term; is the weight coefficient of the knowledge mastery deviation item; For students Control strategy consistency adjustment coefficient; For students The attention contrast adjustment coefficient of Specifically, in this embodiment, step S6 includes an individualized parameter adjustment unit, which connects the strategy optimization unit, the state feedback update unit and the student portrait storage module, and is used to dynamically update the individual parameter configuration in the control strategy based on historical learning behavior and feedback error, thereby realizing a multi-round iterative control mechanism.

[0081] First, the individualized parameter adjustment unit receives multiple rounds of error sequences from the state feedback update unit: ; in, Indicates students In the Round control execution The state deviation of each time step; Indicates the current cumulative number of intervention-feedback iterations.

[0082] The adjustment unit performs a sliding average operation on the cross-wheel error to smooth the individual feedback trend: ; in, For students at the time The average deviation vector of is used to characterize the long-term error trend of individuals; Indicates the current cumulative number of intervention-feedback iterations; Indicates at a point in time Time, dimension The average state change on .

[0083] The system then updates the adjustment weight matrix in the individual control strategy based on the error distribution trend and control adjustment sensitivity. , adjust the expression as follows: ; in, For the Individualized adjustment gain matrix in round strategy; is the learning rate constant, which controls the parameter update amplitude; is the sign function, corresponding to the error direction; It indicates that a copy operation is performed on the control dimension to form a consistent adjustment vector.

[0084] The individual parameter adjustment department will update the Return to the strategy optimization unit and replace the original gain matrix for the next round of strategy calculation to achieve a closed-loop parameter adjustment.

[0085] Furthermore, the system synchronously stores each round of strategy parameters, error change curve and learning behavior records in the student portrait storage module to form an individual evolution history.

[0086] The student portrait storage module includes a feature vector recording unit, a model deviation cache unit, and an individual parameter adjustment trajectory unit. The units are logically interconnected and support time series data indexing and model backtracking comparison.

[0087] Through the collaborative structure of the individualized parameter adjustment unit and the student profiling module, the system can dynamically identify changes in student behavior trends, automatically adapt the parameter adjustment intensity in the control strategy, and achieve optimal control stability in multiple rounds of iterations; In step S6, coordination is performed based on the local game model, and the local game model describes the students' learning strategies by introducing individual difference parameters of students. Then its cost function is for: ; in, For students The multi-round comprehensive cost function value of ; is the discrete time step index; is the total number of time steps; student In the The attention level at each time step; For students In the The knowledge mastery of each time step; For students In the Control strategy input for time steps; For students Control strategy input for other students except students The average attention level of other students; is the weight coefficient of the attention bias term; is the weight coefficient of the knowledge mastery bias term; For students The control strategy consistency adjustment coefficient for students The attention contrast adjustment coefficient.

[0088] Ultimately, this mechanism can achieve a self-closed loop of cross-wheel adjustment feedback, with the technical effects of transparent adjustment path, fast adaptation speed and clear storage structure.

[0089] S7. Based on the ultimately obtained personalized optimal teaching control strategy, the teaching content, task type, and difficulty are dynamically pushed and adjusted to ensure that each student's learning process is efficient and personalized. In step S7, the process of dynamically pushing and adjusting teaching content, task types and difficulty includes the following operations: Real-time monitoring of students' learning status, including their attention level, knowledge mastery, emotional state, and learning progress; Dynamically adjust the difficulty of learning tasks based on students' real-time learning status and select appropriate task types based on students' ability levels; Adjust the order and time of learning tasks based on students' learning progress and emotional state to ensure that the assignment of tasks matches students' personalized learning progress; Dynamically push personalized teaching content and adjust the type, difficulty and order of tasks based on students' real-time feedback.

[0090] Specifically, in this embodiment, step S7 includes a group analysis attribution unit, which is connected to a state feedback update unit, an individualized parameter adjustment unit, and a student portrait storage module, and is used to perform evolutionary trend modeling and causal attribution analysis based on the state data of all learners after multiple rounds of strategy execution.

[0091] First, the system aggregates the state feedback sequence of each student after each round of intervention to construct a group-level state evolution trajectory tensor: ; in, For students In the Round The observed state at a time point; is the total number of students; is the number of intervention rounds; is the number of time steps per round.

[0092] The group analysis attribution unit extracts the covariance spectrum of the main change direction by performing statistical dimensionality reduction on the tensor along the time axis and the individual axis: ; in, is the overall state average vector; Represents the co-evolution relationship between state variables; For tensors Perform statistical covariance calculations for the entire sample; is the total number of tensor samples, that is, the total number of combinations of all students, all rounds, and all time steps.

[0093] The system then constructs a causal graph model to identify the statistical relationship network between state indicator changes and control variables. The graph takes the control vector as input and the state transition gradient as output to construct a mapping: ; in, Indicates the The control variables The sensitivity of the state variables; Indicates time At this moment, the system The object's Observation state quantity in dimensions; It represents the rate of change of the state variable with respect to the control variable. The derivative operation is obtained through regression fitting and disturbance response modeling, avoiding direct analytical form.

[0094] Attribution Graph It is used to identify the input channel that has the greatest impact on the target state in teaching control, which serves as the basis for subsequent strategy inversion tuning.

[0095] The group analysis and attribution department feeds back the highly sensitive channels to the strategy optimization department, which compresses the control variable space of the optimal control input generation module, retaining only the dimensions with significant impact: ; in, Based on causal graph The projection function is used to map the control variables to the main influence space; is the original control vector; is the control vector after projection.

[0096] Through the group attribution tuning mechanism, the system can effectively reduce ineffective control intervention during strategy execution and improve the efficiency of control-state mapping.

[0097] The attribution department further stores the impact factor rankings, change trend labels and sensitivity thresholds into the student portrait storage module to support subsequent long-term learning path planning.

[0098] The entire group evolution attribution mechanism constitutes a five-stage closed-loop process of control-observation-attribution-counter-control, ensuring the interpretability of individual intervention strategies and the robustness of group strategy configuration.

[0099] This step, by introducing high-order covariance modeling and causal directional reasoning methods, can achieve automatic compression of control variables and reverse correction of models in group evolution tuning, and has the technical effects of data dimension noise reduction, model generalization enhancement and strategy dimension optimization.

[0100] Example 2: Please see the attached Figure 2 , a smart teaching management device includes: Data collection module, used to obtain multimodal learning behavior data of multiple students; A learning state construction module, which is connected to the data acquisition module and is used to construct a corresponding learning state vector based on the learning behavior data; A learning state evolution model module, which is connected to the learning state construction module and is used to establish a nonlinear learning state evolution model for each student based on the learning state vector; The optimization target construction module is connected to the learning state evolution model module and is used to construct the optimization target function that represents the student learning effect and the teaching control cost; The control strategy optimization module is connected to the optimization target construction module and is used to optimize the teaching control strategy based on the optimization objective function and the optimal control principle to obtain the personalized teaching control trajectory for each student; A hysteresis compensation module is connected to the control strategy optimization module and is used to construct a hysteresis compensation model and adjust the control strategy according to the feedback lag characteristics; A game coordination module, which is connected to the hysteresis compensation module and is used to coordinate the control strategies of multiple students based on the local game model; The dynamic push module is connected to the game coordination module and is used to dynamically push and adjust the teaching content, task type and difficulty according to the optimized control strategy.

[0101] Specifically, the device first uses the data acquisition module to collect multimodal behaviors of each student and generate an original learning behavior data stream, including operation records, physiological parameters, voice expressions and other information.

[0102] The above data is transmitted to the learning status construction module in real time. The module preprocesses and fuses the input data to construct the student's learning status vector in the current time period, which is used to reflect dimensions such as individual attention level, mastery level and emotional response.

[0103] The learning state vector is continuously input into the learning state evolution model module, which establishes a nonlinear state evolution process model for each student to characterize the dynamic change trend of their state under teaching intervention.

[0104] The learned state evolution model is then passed to the optimization objective construction module, which combines the learned state with the teaching control behavior to construct an optimization objective function that simultaneously minimizes the state error and control cost.

[0105] The optimization objective function is transmitted to the control strategy optimization module, which solves the personalized teaching control trajectory based on the optimal control principle and determines the teaching control input that each student should take at different time points.

[0106] The control strategy optimization results are passed to the lag compensation module, which detects and models the time-lag response characteristics in the feedback path, and performs time alignment and dynamic compensation on the control instructions to avoid the mismatch between control and response that affects the teaching effect.

[0107] The compensated control strategy is synchronously transmitted to the game coordination module. When there are collaborative tasks or resource competition relationships among the student group, the module adjusts the control strategy of each individual based on the local game mechanism to achieve coordination and consistency among the overall strategies.

[0108] Finally, the coordinated control strategy is input into the dynamic push module. Based on the parameters set in the control vector, such as task intensity, content type, and feedback rhythm, personalized teaching content is generated in real time and pushed to the student end for execution, achieving closed-loop control. This section is designed based on the aforementioned method. The aforementioned method has fully disclosed the key control models, state expressions, optimal control paths and feedback correction methods, so it will not be repeated here.

[0109] Example 3: Please see the attached Figure 3 , an electronic device, comprising: a processor for executing a computer program stored in the electronic device; Memory, used to store programs and related data required for execution; Communication module, used for data transmission with external devices, supporting data collection and teaching content push; Display module, used to display the teaching content, task type and difficulty adjustment information generated by the dynamic push module; The interactive interface is used to interact with students, receive students' feedback and transmit it to the data collection module.

[0110] Specifically, when the system is started, the computer program stored in the memory is loaded and executed by the processor, initializing each functional module and establishing a data interaction process with the management device.

[0111] The communication module first receives data input from the teaching management device, including teaching strategy control instructions, push content, and behavior data type instructions that need to be collected.

[0112] The display module presents the corresponding teaching content, task interface and individual difficulty adjustment information according to the control parameters generated by the dynamic push module for students to understand and execute.

[0113] When completing learning tasks, students answer questions, select feedback, or perform language interactions through the interactive interface. All interaction data is collected by the interactive interface and transmitted to the data acquisition module in real time.

[0114] The data acquisition module uses the communication module to synchronously upload the above-mentioned collection results to the teaching management device, realizing the real-time return of student-side data and providing an input basis for subsequent state construction and strategy optimization.

[0115] At the same time, the processor of the electronic device continuously calls the optimized push parameters and responds to the scheduling of the management device, periodically refreshes the teaching interface and task settings, and supports individualized, dynamic adjustment and real-time feedback of teaching rhythm control.

[0116] The memory in this electronic device is used to store local intermediate cache data, behavior logs, and feedback sampling information, supporting the operational requirements of breakpoint resume teaching and long-term trajectory tracking.

[0117] Through the joint operation of the above modules, the electronic equipment realizes the display of teaching tasks, the collection of student feedback, and the issuance and response of teaching instructions, forming a control-feedback closed-loop interface with the student end as the execution carrier.

[0118] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A smart teaching management method, characterized in that: The following steps are involved: S1. Obtain multimodal learning behavior data of multiple students and construct corresponding learning state vectors; S2. Based on the learning state vector, establish a nonlinear learning state evolution model for each student, where the model describes the dynamic process of the student's learning state changing over time; S3. Based on the nonlinear learning state evolution model of each student, construct an optimization objective function that represents the student's learning effect and teaching control cost; S4. Utilizing the optimal control principle and based on the optimization objective function, optimizing the teaching control strategy to obtain a personalized teaching control trajectory for each student, and making the trajectory capable of maximizing the student's learning effect; S5. Based on the time lag characteristics in the feedback process, a hysteresis compensation model is constructed to predict the feedback delay and adjust the control strategy; S6. Considering the impact of strategic interactions among student groups, the control strategies of multiple students are coordinated based on the local game model to optimize the overall learning effect of the group; S7. Based on the final personalized optimal teaching control strategy, the teaching content, task type and difficulty are dynamically pushed and adjusted to ensure that each student's learning process can be carried out efficiently and personalized.

2. A smart teaching management method according to claim 1, characterized in that: In step S1, the multimodal learning behavior data includes the student's facial expression, eye concentration, voice output, mouse trajectory and answering time, and the learning state vector includes the student's attention level, knowledge mastery and emotional state.

3. The intelligent teaching management method according to claim 1, characterized in that: In step S2, the nonlinear learning state evolution model is described by a differential equation, which takes into account multiple learning state factors of the students and models the students' attention, emotions and knowledge mastery over time.

4. The intelligent teaching management method according to claim 1, characterized in that: In step S3, the optimization objective function includes a comprehensive cost term of the student's attention level, knowledge mastery, and teaching cost, wherein the cost term is: ; in, For students The comprehensive cost function value of is the weight coefficient of the attention bias term; For students attention level; is the weight coefficient of the knowledge mastery bias term; For students knowledge mastery; is the weight coefficient of the teaching cost item; For students the cost of teaching.

5. The intelligent teaching management method according to claim 1, characterized in that: In step S4, the optimization teaching control strategy is implemented by solving the Euler-Lagrange equation in the variational method, wherein the Euler-Lagrange equation is given by the following form: ; in, is the Lagrangian function; For control variables ; For control variables About Time The derivative of , that is, the rate of change of the control variable; is the Lagrangian function For control variables The partial derivative of is the Lagrangian function Derivative of the control variable The partial derivative of is the derivative of the above partial derivative with respect to time.

6. A smart teaching management method according to claim 1, characterized in that: In step S5, the time lag characteristics include feedback lag, control strategy implementation lag, system delay, student response lag and strategy adjustment lag; The hysteresis compensation model predicts the feedback lag of the control strategy by constructing a delayed feedback function, and adjusts the control strategy according to the prediction result, so that the adjusted control strategy satisfies the following relationship: ; in, For the moment The control variable The value of To control the time lag, that is, the delay time of the system response; For control variables In the lag time The value of when; is the control change rate function within the time period; From time arrive The integral value of the control rate of change.

7. The intelligent teaching management method according to claim 1, characterized in that: In step S6, the local game model describes the students’ learning strategies by introducing individual difference parameters of students, so its cost function for: ; in, For students The multi-round comprehensive cost function value of ; is the discrete time step index; is the total number of time steps; student In the The attention level at each time step; For students In the The knowledge mastery of each time step; For students In the Control strategy input for time steps; For students Control strategy input from other students; For students The average attention level of other students; is the weight coefficient of the attention bias term; is the weight coefficient of the knowledge mastery deviation item; For students Control strategy consistency adjustment coefficient; For students The attention contrast adjustment coefficient.

8. The intelligent teaching management method according to claim 1, characterized in that: In step S7, the process of dynamically pushing and adjusting the teaching content, task type and difficulty includes the following operations: Real-time monitoring of students' learning status, including their attention level, knowledge mastery, emotional state, and learning progress; Dynamically adjust the difficulty of learning tasks based on students' real-time learning status and select appropriate task types based on students' ability levels; Adjust the order and time of learning tasks based on students' learning progress and emotional state to ensure that the assignment of tasks matches students' personalized learning progress; Dynamically push personalized teaching content and adjust the type, difficulty and order of tasks based on students' real-time feedback.

9. A smart teaching management device, according to a smart teaching management method according to any one of claims 1 to 8, characterized in that: include: Data collection module, used to obtain multimodal learning behavior data of multiple students; A learning state construction module, connected to the data acquisition module, for constructing a corresponding learning state vector based on the learning behavior data; a learning state evolution model module, connected to the learning state construction module, for establishing a nonlinear learning state evolution model for each student based on the learning state vector; The optimization target construction module is connected to the learning state evolution model module and is used to construct the optimization target function that represents the student learning effect and the teaching control cost; A control strategy optimization module is connected to the optimization target construction module and is used to optimize the teaching control strategy based on the optimization target function using the optimal control principle to obtain a personalized teaching control trajectory for each student; A hysteresis compensation module is connected to the control strategy optimization module and is used to construct a hysteresis compensation model and adjust the control strategy according to the feedback lag characteristics; A game coordination module, which is connected to the hysteresis compensation module and is used to coordinate the control strategies of multiple students based on the local game model; The dynamic push module is connected to the game coordination module and is used to dynamically push and adjust the teaching content, task type and difficulty according to the optimized control strategy.

10. An electronic device, according to the intelligent teaching management device of claim 9, characterized in that: include: a processor, configured to execute a computer program stored in the electronic device; Memory, used to store programs and related data required for execution; Communication module, used for data transmission with external devices, supporting data collection and teaching content push; Display module, used to display the teaching content, task type and difficulty adjustment information generated by the dynamic push module; The interactive interface is used to interact with students, receive students' feedback and transmit it to the data collection module.