A student psychological state dynamic evaluation and teaching intervention method and system

By using multimodal signal fusion and multi-scale temporal networks, the problems of semantic comprehension bias and strategy rigidity in student psychological state assessment and teaching intervention were solved, enabling precise teaching intervention and student psychological state assessment, and improving teaching effectiveness and student learning experience.

CN122347274APending Publication Date: 2026-07-07HEBEI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610613770.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing technologies for assessing students' psychological state and intervening in teaching suffer from problems such as semantic comprehension bias, limited information dimensions, weak environmental robustness, inability to adapt to modal conflicts, lagging trend prediction, inaccurate intervention timing, and rigid and unpersonalized strategies, resulting in poor teaching effectiveness.

Method used

A multimodal signal fusion framework is adopted, which combines visual, audio, behavioral and physiological signals. Through dynamic weighted feature fusion using a cross-attention mechanism, a multi-scale hybrid temporal network is constructed to capture second-level state fluctuations and minute-level trends. Intervention decisions are made in conjunction with teaching events and student profiles, and a feedback closed-loop optimization strategy is established.

Benefits of technology

It has enabled more accurate assessment of students' psychological state, improved the accuracy of identification and prediction, ensured the precision of the timing of teaching interventions and the personalization of strategies, and enhanced classroom teaching effectiveness and students' learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347274A_ABST
    Figure CN122347274A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of student psychological state evaluation and teaching intervention, and specifically discloses a student psychological state dynamic evaluation and teaching intervention method and system, which comprises the following steps: collecting various signals of students in the teaching process; calculating an emotional probability distribution; combining course difficulty, student portraits and teaching events to construct enhanced input features; using a multi-scale mixed time sequence network architecture to capture the second-level state fluctuation and minute-level trend evolution of the enhanced input features, and fusing to obtain unified time sequence representation; predicting the mean and variance of the future K time step prediction distribution; and if it is judged to trigger teaching intervention, selecting the optimal teaching intervention strategy according to the matching degree of the strategy and the current emotional state of the student, etc. The present application can capture the second-level state fluctuation and minute-level trend evolution, accurately fit the state evolution law in the education scene, accurately determine the student state, trigger teaching intervention, and improve the classroom teaching effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of student psychological state assessment and teaching intervention technology, and in particular to a method and system for dynamic assessment of student psychological state and teaching intervention. Background Technology

[0002] With the widespread application of online education, real-time perception of students' psychological states and timely adjustment of teaching strategies have become the core foundation for improving teaching effectiveness. Accurate and robust student status monitoring and intervention have become a key challenge for intelligent education systems. Currently, existing technologies typically employ the following technical approaches in areas such as status perception, feature fusion, trend prediction, and intervention decision-making, but these approaches have corresponding shortcomings: 1. In terms of state perception and modeling, these methods mainly rely on unimodal or general emotion models. These methods primarily depend on facial images captured by cameras, treating the student's state as a static sequence of visual features, and applying deep learning models to automatically extract features from facial expressions, directly mapping them to a general basic emotion classification system, such as Ekman's six emotions.

[0003] Existing drawbacks: (a) Semantic gaps in educational scenarios: General emotion classification systems cannot distinguish educational-specific states, such as cognitive confusion and anxiety, which have similar facial features but different causes, leading to semantic comprehension bias in state recognition and making it difficult to support precise teaching.

[0004] (b) Limited information dimensions: It completely ignores key clues such as tone of voice, interactive behavior and physiological signals, and cannot fully depict the psychological state of students in the complex learning process, resulting in insufficient completeness of state assessment.

[0005] (c) Weak environmental robustness: The model performance is highly dependent on the quality of a single signal. Once it encounters scenarios such as side face occlusion or drastic changes in lighting, the feature extraction capability will be severely impaired, making the system unusable.

[0006] 2. Regarding feature fusion mechanisms, fixed weights or simple multimodal splicing methods are mainly used. Although these methods introduce multi-source signals such as visual, audio, and behavioral signals, they usually use early feature splicing or late-stage decision voting for fusion, with the weights of each modality pre-set and fixed.

[0007] Existing drawbacks: (a) Modal conflict cannot be adaptive: The fusion mechanism lacks awareness of signal quality. When a certain modal signal fails due to environmental interference, such as excessive microphone noise or camera obstruction, the model still assigns it a high weight, causing the overall recognition performance to plummet due to the drag from the low-quality modality.

[0008] (b) Ignoring feature space heterogeneity: The feature distributions of different modalities vary greatly. Simple splicing leads to a shift in the feature space distribution, affecting the learning of classification boundaries and limiting the model's generalization ability.

[0009] (c) Insufficient utilization of complementary information, failure to explore deep semantic relationships between modalities, and only the superposition of data at the data level, which cannot achieve information complementarity and noise suppression between modalities.

[0010] 3. In terms of trend prediction capabilities, these methods are mainly based on static identification or simple time series models. Most of these methods only classify and identify the state at the current moment, or use standard recurrent neural networks to model historical sequences and output a single state prediction value.

[0011] Existing drawbacks: (a) Delayed intervention timing: It only reflects the current state and cannot detect the trend of attention decline in advance, so the intervention is often triggered after the student has already become distracted, thus missing the best intervention window.

[0012] (b) Insufficient multi-scale dynamic capture: It is difficult to capture both sudden state changes at the second level (such as sudden distraction) and trend evolution at the minute level (such as gradual fatigue) at the same time, resulting in limited ability to fit the evolution of complex states.

[0013] (c) Missing prediction confidence: Only point estimates are output, failing to inform the system of the reliability of the prediction. During periods of drastic state fluctuations, the model may still output incorrect predictions with high confidence, leading to unintended interventions.

[0014] 4. In terms of intervention decision-making logic, it is mainly based on static rules or human experience. After state identification, such methods rely on hard rules preset by education experts or rely entirely on teachers to manually adjust the teaching pace after observation.

[0015] Existing drawbacks: (a) Rigid and impersonal strategies: Fixed rules cannot adapt to students' individual differences and dynamic preferences. The same intervention strategy may have drastically different effects on different students, resulting in a decrease in the effectiveness of the intervention over time.

[0016] (b) Lack of feedback loop: After the system implements the intervention, it does not collect data on students' responses to the intervention, making it impossible to determine whether the strategy is effective, and the model has difficulty in achieving self-optimization and evolution of strategy weights.

[0017] (c) Weak context awareness: Without considering the context of course difficulty and teaching process, intervention may be triggered at inappropriate times, causing negative interference, interrupting the learning flow, and placing a heavy burden on teachers for manual intervention. Summary of the Invention

[0018] This invention aims to solve the aforementioned problems. To this end, this invention provides a method and system for dynamic assessment of student psychological states and teaching intervention. It can capture complementary information from visual, audio, behavioral, and physiological signals to achieve comprehensive psychological state assessment. Based on dynamic weighting of modal quality perception and cross-fusion of multimodal features through attention mechanisms, it significantly improves the accuracy of student emotion recognition. A multi-scale hybrid temporal network architecture is constructed to capture second-level state fluctuations and minute-level trend evolution. Combined with uncertainty quantification, it achieves accurate fitting of state evolution patterns in educational scenarios, accurately determines student states, triggers teaching interventions in advance, and improves classroom teaching effectiveness.

[0019] This invention provides a method for dynamic assessment and teaching intervention of students' psychological state, and the technical solution adopted is as follows: including: S1: During the teaching process, collect students' visual signals, audio signals, behavioral signals and physiological signals; S2: Extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors; perform feature fusion through a cross-attention mechanism to calculate multimodal fusion features; and calculate the emotion probability distribution based on the multimodal fusion features. S3: Construct an emotion probability sequence based on multiple emotion probability distributions, and combine it with course difficulty, student profiles and teaching events to construct enhanced input features; S4: Use a multi-scale hybrid temporal network architecture to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features, and fuse them to obtain a unified temporal representation; predict the mean and variance of the predicted distribution for the next K time steps based on the unified temporal representation; determine whether to trigger teaching intervention based on the mean and variance of the predicted distribution for the next K time steps; if teaching intervention is triggered, proceed to S5. S5: Select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.

[0020] Furthermore, a quality assessment vector is calculated based on signal-to-noise ratio, occlusion rate, data integrity, and feature variance.

[0021] Furthermore, the labels for emotions include: cognitive dimension, affective dimension, and behavioral dimension; The cognitive dimension includes focus, cognitive confusion, insight, and cognitive overload; The emotional dimension includes learning pleasure, anxiety, burnout, and curiosity; The behavioral dimension includes active participation, passive following, distraction and detachment, and meditative states.

[0022] Furthermore, in step S3, a gating weight vector is generated using course difficulty and student profiles to modulate the emotion probability sequence and teaching events, thereby obtaining enhanced input features.

[0023] Furthermore, the formula for calculating the enhanced input features is as follows: in, For context-gated vectors, For gated functions, Here is the weight matrix for gating. This is the bias vector for gating. Code the difficulty of the course. To create a student profile, To enhance input features, It is a probability sequence of emotions; Embedding sequences for teaching events, This is element-wise multiplication.

[0024] Furthermore, in step S4, the operation of the multi-scale hybrid temporal network architecture is as follows: Enhanced input features: Input dilated convolutional networks capture second-level state fluctuations and output local features; Local features are input into the Transformer encoder to capture minute-level trend evolution, and global features are output. The scale fusion layer weightedly fuses local and global features to generate a unified temporal representation.

[0025] Furthermore, in step S4, the unified temporal representation is input into the prediction head network to predict the mean and variance of the prediction distribution for the next K time steps; the structure of the prediction head network is a multilayer perceptron. The prediction head network is trained using uncertainty-weighted algorithms and dynamically calculated weight loss functions based on teaching event embeddings. ) in, For pattern weights, It is the Sigmoid activation function. For learnable weight vectors, Embedded into teaching events, These are learnable bias parameters. For loss function, For the current time step, For traversal index, For real labels, For time steps The mean of the predicted probability of attention decay, For time steps The variance of the prediction uncertainty This represents the total number of time steps to be predicted.

[0026] Furthermore, the time step with the smallest variance is selected from the variances of the predicted distributions over the next K time steps. If the mean of this time step is greater than the basic warning threshold and the variance of this time step is less than the uncertainty threshold, then the teaching intervention is triggered.

[0027] Furthermore, the degree of matching between the strategy and the student's current emotional state is represented by the similarity between the strategy and the student's emotional state vector. Student emotional state vector The calculation formula is: in, The mean of the predicted attention decay probability at the time step with the minimum variance. The uncertainty variance of the prediction at the time step with the minimum variance. For the probability distribution of emotions, Encode the difficulty level of the course.

[0028] This invention also provides a dynamic assessment and teaching intervention system for students' psychological state, the technical solution of which is as follows: including: The signal acquisition and processing module is used to acquire students' visual signals, audio signals, behavioral signals, and physiological signals during the teaching process. The emotion classification module is used to extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors. Feature fusion is performed through a cross-attention mechanism to calculate multimodal fusion features. The emotion probability distribution is then calculated based on the multimodal fusion features. The feature enhancement module is used to construct an emotion probability sequence based on multiple emotion probability distributions, and to construct enhanced input features by combining course difficulty, student profiles and teaching events; The teaching intervention judgment module is used to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features using a multi-scale hybrid temporal network architecture, and fuse them to obtain a unified temporal representation; based on the unified temporal representation, it predicts the mean and variance of the predicted distribution for the next K time steps; based on the mean and variance of the predicted distribution for the next K time steps, it determines whether to trigger a teaching intervention; if a teaching intervention is triggered, it calls the teaching intervention execution module. The teaching intervention strategy selection module is used to select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.

[0029] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. This invention constructs a multimodal fusion framework for psychological states specific to educational scenarios, capturing complementary information from visual, audio, behavioral, and physiological signals, and establishing a three-dimensional classification system for academic emotions encompassing cognition, emotion, and behavior. Compared to general emotion models, this invention significantly improves the semantic distinguishability of education-specific states, achieving a more comprehensive and accurate assessment of students' psychological states, and providing a reliable basis for precision teaching.

[0030] 2. This invention utilizes a cross-attention mechanism to weightedly fuse multimodal features, achieving dynamic evaluation and weighting of modality credibility. Compared to fixed-weight fusion methods, this invention can adaptively suppress interference from low-quality modalities in complex environments such as lighting changes, occlusion, and noise, fully utilizing high-quality modal information, significantly improving recognition accuracy, and ensuring stable system operation in real teaching scenarios.

[0031] 3. This invention constructs a multi-scale hybrid temporal network architecture, coupled with an uncertainty quantification mechanism, which solves the problems of traditional methods failing to detect trends in advance and lacking confidence assessment. By simultaneously capturing sudden changes and long-term trends through the multi-scale hybrid temporal network architecture and introducing the perception of teaching events, the prediction accuracy of attention decay trends is significantly improved; the uncertainty quantification mechanism effectively reduces the false alarm rate, ensuring that intervention is triggered only when the confidence level is high, thus achieving precise advance intervention timing.

[0032] 4. The intervention strategy of this invention can dynamically adjust the teaching content and interaction methods according to the students' psychological state. Compared with a static rule base, this invention also establishes a closed loop for intervention effect feedback, which can automatically optimize the strategy weight based on individual student differences and historical feedback, significantly improving the accuracy of strategy matching and the long-term intervention success rate, effectively reducing the burden on teachers, and improving the overall classroom teaching efficiency and students' learning experience.

[0033] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is a flowchart of the method provided by the present invention.

[0036] Figure 2 This is a schematic diagram of the multi-scale hybrid temporal network architecture provided by the present invention. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but should not be used to limit the scope of this invention.

[0038] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0039] The following is combined Figure 1 and Figure 2 The present invention will be further described in detail below, including a method and system for dynamic assessment of students' psychological states and teaching intervention: Preliminary preparation: Constructing a multimodal fusion framework for psychological states specifically for educational scenarios. This embodiment first establishes the correspondence between four student modalities and 12 categories of academic emotions.

[0040] 1. Definition and acquisition of multimodal signals.

[0041] This embodiment models a student's psychological state within a teaching time window as a multimodal temporal signal pair. It includes: (1) Visual modality: defined as an image sequence captured by a camera, including facial action units, head posture angles, and gaze direction vectors; (2) Audio modality: defined as a speech waveform sequence captured by a microphone, including fundamental frequency, energy, speech rate, and pause duration; (3) Behavioral modality: defined as a terminal interaction log sequence, including mouse click frequency, keyboard input interval, and page dwell time; (4) Physiological modality: defined as a physiological signal sequence captured by a wearable device, including heart rate variability and skin conductance response.

[0042] In this embodiment, after constructing multimodal time-series signal pairs from visual, audio, behavioral, and physiological modalities, a quality assessment vector for that modality within the current time window is generated based on the signal-to-noise ratio, occlusion rate, data integrity, and feature variance of the modality signal. , Used to quantify the reliability of each modal signal.

[0043] 2. Spatiotemporal alignment and preprocessing.

[0044] Time alignment: Based on a preset time granularity, an interpolation method is used to resample asynchronously acquired multi-source signals to generate a synchronous timestamp sequence; Spatial alignment: For the face region in the visual signal, the target detection model is used to locate key points and uniformly map them to the standard facial coordinate system.

[0045] Missing information awareness: generating modal validity mask In this context, a value of 1 indicates that the corresponding modality is valid at the current moment, while a value of 0 indicates that it is missing or unreliable. A modality is deemed unreliable when its quality assessment vector is below a threshold.

[0046] 3. Construction of an education-specific emotion classification system.

[0047] Unlike general emotion classifications, this embodiment is based on the control-value theory and academic emotion theory in educational psychology. It ensures multimodal observability for each label category and utilizes technologies such as computer vision, speech processing, and behavioral analysis to construct a three-dimensional academic emotion labeling system encompassing cognition, emotion, and behavior, totaling 12 categories of academic emotions. The cognitive dimension reflects students' processing of knowledge content and directly relates to learning outcomes; the emotional dimension reflects students' emotional experiences during the learning process, influencing learning motivation and persistence; and the behavioral dimension reflects students' behavioral tendencies in participating in learning activities, representing the outward expression of cognition and emotion. Verification has shown that it covers over 90% of high-frequency learning states, with an accuracy rate >85%.

[0048] Cognitive dimension: includes focus, cognitive confusion, insight, and cognitive overload; Emotional dimension: includes learning pleasure, anxiety, burnout, and curiosity; Behavioral dimension: includes active participation, passive following, distraction and detachment, and meditative state.

[0049] It should be noted that the labeling follows the multimodal evidence chain principle, assigning a label only when multiple modal signals support the same label, thus ensuring the objectivity of the label. For example, a label is only assigned if there are no missing signals from the visual, audio, behavioral, and physiological modalities, and all of them support cognitive confusion.

[0050] In this embodiment, as Figure 1As shown, a method for dynamic assessment and teaching intervention of students' psychological state is provided, including the following steps: S1: During the teaching process, collect students' visual signals, audio signals, behavioral signals, and physiological signals.

[0051] The visual signal is an image sequence captured by the camera, including facial motion units, head posture angle, and gaze direction vector; the audio signal is a speech waveform sequence captured by the microphone, including fundamental frequency, energy, speech rate, and pause duration; the behavioral signal is a terminal interaction log sequence, including mouse click frequency, keyboard input interval, and page dwell time; and the physiological signal is a physiological signal sequence captured by the wearable device, including heart rate variability and skin conductance response.

[0052] S2: Extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors; perform feature fusion through a cross-attention mechanism to calculate multimodal fusion features; and calculate the emotion probability distribution based on the multimodal fusion features.

[0053] Specifically, it includes the following steps: S2.1: Extract features from visual signals, audio signals, behavioral signals, and physiological signals to obtain facial temporal features, speech prosody features, interaction sequence features, and physiological signal features.

[0054] After spatiotemporal alignment, visual, audio, behavioral, and physiological signals are used for feature extraction via the encoders described below. The visual encoder employs a convolutional neural network to extract temporal facial features. Audio encoder: Employs a pre-trained speech model to extract speech prosodic features. Behavior encoder: Employs a temporal convolutional network to extract interaction sequence features. Physiological encoder: Employs a one-dimensional convolutional network to process physiological signal features. .

[0055] S2.2: Modal credibility dynamic weighting mechanism: calculate the quality assessment vector based on facial temporal features, speech prosody features, interaction sequence features and physiological signal features, and then calculate the dynamic fusion weight of the four features based on the quality assessment vector.

[0056] The quality assessment vector is calculated based on four features: signal-to-noise ratio, occlusion rate, data integrity, and feature variance. In this embodiment, a weighted summation method is used.

[0057] In this embodiment, it is also necessary to determine whether the features are missing or unreliable; filter out the valid signals, and determine the current set of valid features.

[0058] To overcome the modal failure problem caused by environmental interference, this embodiment designs a learnable weight generator to achieve quality perception and dynamic allocation: in, The dynamic fusion weights for feature m; Let m be the quality evaluation vector of feature m; For the shared quality characteristic transformation matrix, The bias is transformed to represent the shared quality characteristics; Let m be the learnable weight vector of feature m; For the current set of valid features, For traversal index, This is the matrix transpose. Feature m includes facial temporal features, speech prosodic features, interaction sequence features, and physiological signal features.

[0059] When the quality of a feature degrades due to occlusion or noise, its dynamic fusion weight is automatically reduced, and the system adaptively increases the contribution of high-quality modalities to ensure the robustness of the fused features.

[0060] S2.3: Based on the dynamic fusion weights, the facial temporal features, speech prosody features, interaction sequence features and physiological signal features are fused through the cross-attention mechanism to calculate the multimodal fusion features.

[0061] A cross-attention mechanism is used to achieve intermodal information complementarity and noise suppression. The calculation formula is as follows: , in, Let m be the query vector for feature m. Let m be the query projection weight matrix for feature m. Let m be the original feature vector. The key vector for all features. The key projection weight matrix for facial temporal features. This is the key projection weight matrix for speech prosodic features. The key projection weight matrix is ​​the feature of the interaction sequence. The key projection weight matrix represents the physiological signal characteristics; The enhanced features are those of feature m after cross-attention calculation. For normalization function, Let be the dimension of the key vector. This is an attention mask; it is 0 when the feature is valid, and negative infinity otherwise. A vector of values ​​for all features; For multimodal fusion features, This is a gate function that suppresses noise characteristics. Let m be the gated projection weight matrix of feature m. This is element-wise multiplication.

[0062] Attention masks are used to mask invalid time steps and missing features. Data for different features may be missing (e.g., students turn off their cameras or mute their microphones). If left untreated, the model will focus on these empty data, leading to errors in feature fusion.

[0063] S2.4: The probability distribution of emotions is calculated based on the multimodal fusion features.

[0064] The multimodal fusion features are processed through a fully connected classification head and then normalized using Softmax to obtain the sentiment probability distribution. .

[0065] S3: Construct an emotion probability sequence based on multiple emotion probability distributions, and combine course difficulty, student profiles and teaching events to construct enhanced input features.

[0066] Traditional time series forecasting only uses historical state sequences, ignoring the teaching context and failing to distinguish between natural decay and event-driven decay. Furthermore, the predictive models for different students and courses have poor general applicability. In educational settings, the same psychological state can have completely different semantics in different contexts. For example, "cognitive confusion" is a normal thought process in advanced courses, but may indicate a student's weak foundation in simple, basic courses.

[0067] To this end, this embodiment constructs enhanced input features, integrates the teaching scenario context into the prediction model, generates gating weight vectors using course difficulty and student profiles, modulates historical emotion probability sequences and teaching events, and can dynamically adjust feature sensitivity according to the teaching context.

[0068] This embodiment introduces context gating, dynamically suppressing or amplifying the sensitivity of specific emotional signals based on the course challenge and individual student differences. The calculation formula is as follows: in, The context gating vector is generated from the course difficulty and student profile through a fully connected layer. Here is the weight matrix for gating. The bias vector for gating; By encoding the difficulty of the courses and adjusting the difficulty sensitivity of the prediction model, the baseline probability of attention decay is improved for high-difficulty courses. Student profiles are characterized by static features such as grade level and historical grades, as well as dynamic features such as recent homework completion rate and preferred intervention types. To enhance input features, It is a probability sequence of emotions; The teaching event embedding sequence encodes teaching events such as starting a quiz and playing a video into vectors, which serve as causal signals for state transitions. This is element-wise multiplication.

[0069] S4: Use a multi-scale hybrid temporal network architecture to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features, and fuse them to obtain a unified temporal representation; predict the mean and variance of the predicted distribution for the next K time steps based on the unified temporal representation; determine whether to trigger teaching intervention based on the mean and variance of the predicted distribution for the next K time steps; if teaching intervention is triggered, proceed to S5.

[0070] S4.1: Single-scale models cannot simultaneously capture second-level mutations and minute-level trends. Dilated convolutional networks excel at local features but are insufficient at modeling long-range dependencies. Transformers excel at global dependencies but lag in responding to sudden changes. This embodiment designs a cascaded architecture, employing a hybrid architecture of local feature extraction and global dependency modeling, weightedly fusing local and global features to generate a unified temporal representation. Figure 2 As shown, a dilated convolutional network is used at a local scale to capture second-level state fluctuations, such as sudden distraction. The input is... Output local features The global scale uses a Transformer encoder to capture minute-level trend evolution, such as gradual fatigue, with the input being... Output global features The scale fusion layer weighted and fused local and global features to generate a unified temporal representation. The multi-scale hybrid temporal network architecture in this embodiment can simultaneously capture short-term mutations and long-term trends, significantly improving the ability to fit the state evolution law.

[0071] S4.2: Input the unified time series representation into the prediction head network to predict the mean and variance of the prediction distribution for the next K time steps. This embodiment no longer outputs only a single predicted value, but instead outputs the mean and variance of the prediction distribution for the next K time steps, thus quantifying uncertainty. in, The mean of the predicted attention decay probability over K future time steps; Let V be the variance of the uncertainty in the prediction over K future time steps; PredictionHead is the prediction head network, specifically a multilayer perceptron with fully connected layers and activation functions, used to map temporal features to prediction results.

[0072] In this embodiment, The sequence length is 10, and the mean and variance of the predicted distribution for the next 3 time steps are predicted, with a time granularity of 1 minute.

[0073] The prediction head network is trained using uncertainty-weighted algorithms and dynamically calculated weight loss functions based on teaching event embeddings. ) in, As a pattern weight, it decreases when an event occurs, focusing more on the variance term and encouraging the model to output large variance (acknowledging uncertainty), and increases during natural decay, focusing more on the prediction error term and encouraging the model to output accurate mean. It is the Sigmoid activation function. For learnable weight vectors, Embedding teaching events provides event context. These are learnable bias parameters; For loss function, For the current time step, For traversal index, The total number of time steps to be predicted. For real labels, For time steps The mean of the predicted probability of attention decay, For time steps The uncertainty variance of the prediction is considered. Uncertainty weighting enables the model to automatically identify difficult-to-predict samples, reduce loss weights by increasing variance, avoid overfitting noise, and improve generalization accuracy. Dynamically calculating weights based on teaching event embeddings enables the model to distinguish between two different state evolution modes: natural decay and event-driven decay, thereby significantly improving prediction reliability and intervention decision quality.

[0074] This embodiment will use the loss function ( ) with mean squared error loss function (MSE), negative log-likelihood loss function (NLL) and fixed weight A value of 0.5 was used for comparison, and the results are shown in Table 1. Evaluation metrics selected included AUC (Area Under the ROC Curve), MAE (Mean Absolute Error), ECE (Expected Calibration Error), and false trigger rate. ↑ indicates a higher value is better, and ↓ indicates a lower value is better. Among these, ECE is the core indicator for measuring the calibration quality of the probabilistic prediction model, used to assess the consistency between the model's prediction confidence and the actual accuracy.

[0075] Table 1

[0076] This embodiment also verifies the contributions of enhanced input features, multi-scale hybrid temporal network architecture, and uncertainty quantification to prediction. A comparative experiment was conducted, and the results are shown in Table 2. It can be seen that each module improves prediction accuracy. The evaluation metrics selected were AUC, F1 (F1 score), and MAE.

[0077] Table 2

[0078] S4.3: Teaching intervention triggering logic: Select the time step with the smallest variance from the variances of the predicted distributions over the next K time steps (denoted as...). If the mean of a given time step is greater than the basic warning threshold and the variance of that time step is less than the uncertainty threshold, then instructional intervention is deemed necessary. This embodiment only triggers intervention when a decrease in predicted attention is observed and the prediction confidence is high, effectively reducing the false trigger rate. In this embodiment, the basic warning threshold is 0.75, and the uncertainty threshold is 0.30.

[0079] S5: Select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.

[0080] (1) Design of intervention strategy library Establish a strategy library containing various intervention methods, such as inserting micro-videos, initiating interactive activities, pushing notifications, and adjusting task difficulty. Each strategy is associated with an applicable condition vector. Strategy matching algorithm: in, A represents the recommended optimal teaching intervention strategy; A is the strategy set. Student emotional state vector: The mean of the predicted attention decay probability at the time step with the minimum variance. The uncertainty variance of the prediction at the time step with the smallest variance is determined by step S4.3; The probability distribution of emotions is calculated in step S2.4; Code the difficulty of the course. For similarity calculation function, Let be the vector of applicable conditions for strategy a; The historical effectiveness rate of strategy a is obtained by calculating the percentage of the historical effective times of strategy a in the total historical usage times. To ensure the novelty of the strategy, an exponential decay model was used to calculate the intervention, avoiding the repeated introduction of the same strategy to students, which could lead to intervention fatigue or resistance, and maintaining the freshness of the intervention and the students' acceptance. For the applicable weights of the strategy, For strategy efficiency weights, This is the weight for strategy novelty.

[0081] (2) Feedback closed-loop optimization Record student status changes after the intervention strategy is implemented, calculate quantitative indicators of intervention effectiveness, including improvement in concentration, reduction in confusion, and reduction in anxiety. Based on the comparison results of the quantitative indicators of intervention effectiveness with preset effectiveness thresholds, update the historical effectiveness parameters of the intervention strategy online in real time. Periodically accumulate intervention record data to optimize the applicable condition vector of each intervention strategy in the strategy library. By updating statistical parameters online in real time and optimizing matching rules offline periodically, the system addresses issues such as strategy rigidity, unsustainable effects, and intervention fatigue, enabling self-evolution of strategy effectiveness, significantly improving strategy matching accuracy, and achieving continuous optimization of intervention results.

[0082] To verify the effectiveness of this method, a comparative test was conducted on two groups of students using both traditional teaching methods and this method. The test subjects were university classes with 35 students each, and the test period was 8 weeks. Specific data comparisons are shown in Table 3.

[0083] Table 3

[0084] This embodiment also provides a dynamic assessment and teaching intervention system for student psychological state, which adopts the following technical solution: including: The signal acquisition and processing module is used to acquire students' visual signals, audio signals, behavioral signals, and physiological signals during the teaching process. The emotion classification module is used to extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors. Feature fusion is performed through a cross-attention mechanism to calculate multimodal fusion features. The emotion probability distribution is then calculated based on the multimodal fusion features. The feature enhancement module is used to construct an emotion probability sequence based on multiple emotion probability distributions, and to construct enhanced input features by combining course difficulty, student profiles and teaching events; The teaching intervention judgment module is used to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features using a multi-scale hybrid temporal network architecture, and fuse them to obtain a unified temporal representation; based on the unified temporal representation, it predicts the mean and variance of the predicted distribution for the next K time steps; based on the mean and variance of the predicted distribution for the next K time steps, it determines whether to trigger a teaching intervention; if a teaching intervention is triggered, it calls the teaching intervention execution module. The teaching intervention strategy selection module is used to select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for dynamic assessment and teaching intervention of students' psychological state, characterized in that, include: S1: During the teaching process, collect students' visual signals, audio signals, behavioral signals and physiological signals; S2: Extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors; perform feature fusion through a cross-attention mechanism to calculate multimodal fusion features; and calculate the emotion probability distribution based on the multimodal fusion features. S3: Construct an emotion probability sequence based on multiple emotion probability distributions, and combine it with course difficulty, student profiles and teaching events to construct enhanced input features; S4: Use a multi-scale hybrid temporal network architecture to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features, and fuse them to obtain a unified temporal representation; predict the mean and variance of the predicted distribution for the next K time steps based on the unified temporal representation; determine whether to trigger teaching intervention based on the mean and variance of the predicted distribution for the next K time steps; if teaching intervention is triggered, proceed to S5. S5: Select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.

2. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1, characterized in that, The quality assessment vector is calculated based on signal-to-noise ratio, occlusion rate, data integrity, and feature variance.

3. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1, characterized in that, Emotional labels include: cognitive dimension, affective dimension, and behavioral dimension; The cognitive dimension includes focus, cognitive confusion, insight, and cognitive overload; The emotional dimension includes learning pleasure, anxiety, burnout, and curiosity; The behavioral dimension includes active participation, passive following, distraction and detachment, and meditative states.

4. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1, characterized in that, In step S3, a gating weight vector is generated using course difficulty and student profiles to modulate the emotion probability sequence and teaching events, thereby obtaining enhanced input features.

5. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 4, characterized in that, The formula for calculating enhanced input features is: in, For context-gated vectors, For gated functions, Here is the weight matrix for gating. This is the bias vector for gating. Code the difficulty of the course. To create a student profile, To enhance input features, It is a probability sequence of emotions; Embedding sequences for teaching events, This is element-wise multiplication.

6. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1, characterized in that, In step S4, the operation of the multi-scale hybrid temporal network architecture is as follows: Enhanced input features: Input dilated convolutional networks capture second-level state fluctuations and output local features; Local features are input into the Transformer encoder to capture minute-level trend evolution, and global features are output. The scale fusion layer weightedly fuses local and global features to generate a unified temporal representation.

7. A method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1 or 6, characterized in that, In step S4, the unified temporal representation is input into the prediction head network to predict the mean and variance of the prediction distribution for the next K time steps; the structure of the prediction head network is a multilayer perceptron. The prediction head network is trained using uncertainty-weighted algorithms and dynamically calculated weight loss functions based on teaching event embeddings. ) in, For pattern weights, It is the Sigmoid activation function. For learnable weight vectors, Embedded into teaching events, These are learnable bias parameters. For loss function, For the current time step, For traversal index, For real labels, For time step The mean of the predicted probability of attention decay, For time step The variance of the prediction uncertainty This represents the total number of time steps to be predicted.

8. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 1, characterized in that, Select the time step with the smallest variance from the variances of the predicted distributions over the next K time steps. If the mean of this time step is greater than the basic warning threshold and the variance of this time step is less than the uncertainty threshold, then the teaching intervention is triggered.

9. The method for dynamic assessment and teaching intervention of students' psychological state as described in claim 8, characterized in that, The degree of matching between the strategy and the student's current emotional state is represented by the similarity between the strategy and the student's emotional state vector. Student emotional state vector The calculation formula is: in, The mean of the predicted attention decay probability at the time step with the minimum variance. The uncertainty variance of the prediction at the time step with the minimum variance. For the probability distribution of emotions, Encode the difficulty level of the course.

10. A dynamic assessment and teaching intervention system for students' psychological state, characterized in that, A method for dynamically assessing and intervening in students' psychological states as described in any one of claims 1 to 9, comprising: The signal acquisition and processing module is used to acquire students' visual signals, audio signals, behavioral signals, and physiological signals during the teaching process. The emotion classification module is used to extract features from visual, audio, behavioral, and physiological signals, calculate quality assessment vectors, and calculate dynamic fusion weights based on the quality assessment vectors. Feature fusion is performed through a cross-attention mechanism to calculate multimodal fusion features. The emotion probability distribution is then calculated based on the multimodal fusion features. The feature enhancement module is used to construct an emotion probability sequence based on multiple emotion probability distributions, and to construct enhanced input features by combining course difficulty, student profiles and teaching events; The teaching intervention judgment module is used to capture the second-level state fluctuations and minute-level trend evolutions of the enhanced input features using a multi-scale hybrid temporal network architecture, and fuse them to obtain a unified temporal representation; based on the unified temporal representation, it predicts the mean and variance of the predicted distribution for the next K time steps; based on the mean and variance of the predicted distribution for the next K time steps, it determines whether to trigger a teaching intervention; if a teaching intervention is triggered, it calls the teaching intervention execution module. The teaching intervention strategy selection module is used to select the optimal teaching intervention strategy based on the strategy's historical effectiveness, novelty, and the degree to which the strategy matches the student's current emotional state.