Interactive classroom student dynamic portrait generation method, medium and system
By integrating facial expression and eye-tracking data with deep learning, the method creates a dynamic student portrait for precise attention evaluation, addressing the limitations of single-feature assessments and enhancing educational effectiveness.
Patent Information
- Application Number
- CN202510377200.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing technology is difficult to accurately evaluate the dynamic changes in students' classroom attention status and cannot provide teachers with real-time and accurate classroom interaction optimization suggestions.
By collecting student facial image sequence and eye movement trajectory data, we construct expression feature matrix and eye concentration indicators, and establish individual correlation models in combination with deep learning technology to achieve accurate assessment and prediction of students' attention status.
It realizes accurate, dynamic and personalized assessment of students' attention status, provides teachers with real-time classroom interaction optimization suggestions, significantly improving the application value and teaching effect of intelligent education technology.
Smart Images

Figure CN120318878A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of student dynamic portraits, and more particularly, relates to a method, medium, and system for generating student dynamic portraits in an interactive classroom. Background Art
[0002] During the classroom teaching process, accurately evaluating the attention state of students is crucial for optimizing teaching strategies. Traditional classroom attention evaluations mainly rely on methods such as subjective observation by teachers or questionnaires, making it difficult to achieve objective and quantitative evaluations. With the development of computer vision technology, attention evaluation methods based on facial expression recognition have gradually been applied to the education field, collecting students' facial data through camera devices and analyzing emotions and attentiveness to provide classroom feedback for teachers.
[0003] However, the existing technologies have obvious limitations. First, most methods only evaluate based on a single feature (such as facial expression or head pose), without comprehensively considering multi-dimensional physiological behavior parameters such as gaze tracking. Second, existing systems mostly adopt static evaluation models, making it difficult to capture the temporal dynamic changes of students' attention. Finally, there is generally a lack of an individual differentiation modeling mechanism, making it difficult to adapt to the association patterns between different students' expressions and attention.
[0004] This results in the existing technologies being difficult to accurately evaluate the dynamic changes in students' attention states in complex teaching scenarios, unable to provide real-time and accurate classroom interaction optimization suggestions for teachers, and severely restricting the application effect of intelligent education technologies. That is to say, there is a technical problem in the existing technologies that it is difficult to accurately evaluate the dynamic changes in students' classroom attention states based on a single feature. Summary of the Invention
[0005] In view of this, the present invention provides a method, medium, and system for generating student dynamic portraits in an interactive classroom, which can solve the technical problem in the existing technologies that it is difficult to accurately evaluate the dynamic changes in students' classroom attention states based on a single feature.
[0006] The present invention is implemented as follows: The first aspect of the present invention provides a method for generating a dynamic portrait of students in an interactive classroom, including: collecting a sequence of facial images of students in the classroom, extracting facial key point data to construct an expression feature matrix; calculating the similarity between the expression feature matrix and a standard expression library to generate an expression similarity vector, recording the frequency of expression changes, and establishing an expression dynamic change curve; collecting a sequence of gaze focus coordinates, calculating the area and residence time where the eyes focus, generating a fixation heat map, and quantifying the gaze concentration index; combining the expression feature matrix and the gaze concentration index to establish an individual association model; identifying specific expressions and marking abnormal learning state points; using a sliding time window to analyze the correlation between the attention concentration and the classroom content nodes, and constructing an attention fluctuation curve; applying a state transition function to model the attention fluctuation and generating an attention state score; generating a dynamic portrait feature map as the dynamic portrait of students based on a learning state evaluation network model that fuses multi-dimensional features; and further including analyzing multiple dynamic portrait feature maps of students using a clustering algorithm to form an evaluation index for the classroom interaction effect.
[0007] Among them, the frequency of expression changes is the number of significant changes in the expression similarity vector of students per unit time, which is obtained by calculating the cumulative difference between the expression similarity vectors at adjacent time points, and reflects the activity level of the change in the emotional state of students.
[0008] Among them, a specific expression is an expression change of a student in the classroom that has a significant statistical difference from the personal expression baseline or the average expression state of the class, which is identified by the Mahalanobis distance or the Z-score method, and reflects the unconventional reaction of the student to the teaching content.
[0009] Among them, the gaze concentration index is the proportion of the residence time of the student's line of sight in the specified area to the total observation time, which is obtained by calculating the spatial aggregation degree and time density of the eye movement trajectory points, and reflects the stability of the student's visual attention.
[0010] Among them, the gaze concentration threshold is the critical value for judging whether the student's line of sight is in a concentrated state, which is determined by analyzing historical data and is used to convert the continuous gaze concentration index into a discrete attention state judgment.
[0011] Among them, the gaze concentration area is the spatial range where the student's line of sight frequently stays, which is obtained by analyzing the sequence of gaze focus coordinates using a density clustering algorithm, and is characterized as a set of hot area coordinates and weight distribution on a two-dimensional plane.
[0012] Among them, the attention concentration index is a quantitative index of the cognitive investment degree of students calculated by comprehensively considering the stability of the expression feature and the gaze concentration index. By weighted fusion of multiple physiological behavior feature parameters, it reflects the degree of attention of students to the classroom content and the depth of cognitive processing.
[0013] Among them, the individual association model is a mapping function established for each student between the facial expression feature matrix and the attention concentration index, which is obtained through deep learning algorithms and is used to adjust the weights of evaluation parameters according to individual differences among students, thereby improving the accuracy of attention state judgment.
[0014] Among them, the state transition function is used to predict the trend of attention state change based on the multi-dimensional feature information and historical state of the student at the current moment. The inputs include the facial expression similarity vector, the eye concentration index, the occurrence frequency of specific expressions, the attention historical value, and the difficulty coefficient of classroom content. The output is the attention state score and the prediction of the change trend, which characterize the current attention investment degree of the student. Through recursive calculation, continuous tracking and prediction of the student's attention state are realized.
[0015] Among them, the learning state evaluation network model structure is a multi-layer bidirectional transformer network architecture, including an encoder layer and a decoder layer. The encoder layer is responsible for extracting the temporal features of the student's facial expression feature matrix and the eye concentration index, and the decoder layer is responsible for fusing the temporal features and outputting the student's attention state score. The core of the model is the multi-head attention mechanism. The number of heads is determined by the facial expression change frequency, the dimension of each attention head is determined by the eye concentration threshold, and the depth of the attention layer is calculated from the attention volatility. The model adopts a residual connection structure to alleviate the problem of gradient disappearance, and a layer normalization operation is added after each layer to improve the training stability. The high-dimensional features are mapped to the student's attention state score space through a fully connected layer.
[0016] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned method for generating a dynamic portrait of students in an interactive classroom.
[0017] The third aspect of the present invention provides a system for generating a dynamic portrait of students in an interactive classroom, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.
[0018] Compared with the prior art, the present invention provides a method, medium, and system for generating a dynamic portrait of students in an interactive classroom. The present invention proposes a method for generating a dynamic portrait of students in an interactive classroom based on multi-modal feature fusion. By simultaneously collecting and analyzing the facial expression sequence and eye movement trajectory data of students, a facial expression feature matrix and an eye concentration index are constructed, and an individual association model is established in combination with deep learning technology to achieve accurate evaluation and prediction of the student's attention state.
[0019] This method effectively solves the limitations of the prior art. Through the multi-feature fusion mechanism, it overcomes the one-sidedness of single-feature evaluation and improves the evaluation accuracy; by using the sliding time window technology and the state transition function, it realizes the continuous tracking of the dynamic changes of the attention state; by introducing the individual correlation model, it solves the problem of one-size-fits-all evaluation criteria and realizes personalized evaluation for different students.
[0020] Thus, the present invention successfully solves the technical problem of accurately evaluating the dynamic changes of students' classroom attention states, provides real-time and accurate optimization suggestions for classroom interaction for teachers, and significantly improves the application value and teaching effect of intelligent education technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0023] As Figure 1 shown, it is a flowchart of a method for generating a dynamic portrait of students in an interactive classroom provided by the first aspect of the present invention. This method includes the following steps:
[0024] S01. Collect a sequence of classroom facial images of multiple students, extract the key point data of the students' faces through facial feature point localization technology, construct an expression feature matrix, and mark the trajectory of each student's facial feature points changing with time;
[0025] S02. Calculate the similarity between each student's expression feature matrix and the standard expression library, generate an expression similarity vector, record the expression change frequency, and establish a dynamic change curve of the students' classroom expressions;
[0026] S03. Collect a sequence of the coordinates of the students' line of sight focus through eye movement tracking technology, calculate the area and residence time of the focused area of the eyes, generate a fixation heat map, and quantify the eye concentration index;
[0027] S04. Combine the students' expression feature matrix and the eye concentration index, establish an individual correlation model, and train an attention concentration evaluation function for individual students through a deep learning algorithm;
[0028] S05. Conduct statistical analysis on the students' expression feature matrices in different time periods of the classroom, identify specific expressions that deviate significantly from the average classroom expression state, and mark them as abnormal learning state points;
[0029] S06. Using the sliding time window technique, analyze the temporal correlation between the student's attention concentration index and the classroom content nodes, and construct a dynamic attention fluctuation curve;
[0030] S07. Apply the state transition function to model the student's attention fluctuation. The input parameters include the expression similarity vector, the eye gaze concentration index, the frequency of occurrence of specific expressions, the attention historical value, and the classroom content difficulty coefficient, and generate the attention state score of the student at the current moment;
[0031] S08. Based on the learning state evaluation network model, fuse multi-dimensional features. The parameters of the multi-head attention mechanism are determined by the expression change frequency, the eye gaze concentration threshold, and the attention volatility, and generate the dynamic portrait feature map of each student as the student's dynamic portrait;
[0032] S09. Optionally, apply the clustering algorithm to analyze the dynamic portrait feature maps of multiple students, identify the distribution of the class learning state, form the classroom interaction effect evaluation index, and provide real-time classroom interaction optimization suggestions for teachers;
[0033] Among them, the expression feature matrix specifically refers to the data set of the spatial position relationship of multiple key points on the student's face and its changes over time, including multi-dimensional feature data such as the positions, movement speeds, and relative position change rates of key points such as eyebrows, eyes, and mouth corners.
[0034] Among them, the expression similarity vector specifically refers to the Euclidean distance or cosine similarity calculation result between the student's current expression feature matrix and various expression templates in the pre-defined standard expression library, which is used to quantify the degree of proximity between the student's expression and standard learning states such as concentration, confusion, and understanding.
[0035] Among them, the expression change frequency specifically refers to the number of significant changes in the student's expression similarity vector per unit time, which is obtained by calculating the cumulative difference between the expression similarity vectors at adjacent time points, and reflects the activity degree of the student's emotional state change.
[0036] Among them, the specific expression specifically refers to the expression change of the student in the classroom that has a significant statistical difference from his / her personal expression baseline or the class average expression state, which is identified by the Mahalanobis distance or Z-score method, and reflects the student's unconventional reaction to the teaching content.
[0037] Among them, the frequency of occurrence of specific expressions specifically refers to the number of times the student generates specific expressions within the unit time window, which is obtained by counting statistics and serves as a quantitative index of the student's cognitive conflict or emotional fluctuation.
[0038] Among them, the eye gaze concentration index specifically refers to the proportion of the time that the student's line of sight stays in the specified area to the total observation time, which is obtained by calculating the spatial aggregation degree and time density of the eye movement trajectory points, and reflects the stability of the student's visual attention.
[0039] Among them, the eye gaze concentration threshold specifically refers to the critical value for judging whether the student's line of sight is in a concentrated state, which is determined by analyzing historical data and is used to convert the continuous eye gaze concentration index into a discrete attention state judgment.
[0040] Among them, the eye gaze concentration area specifically refers to the spatial range where the student's line of sight frequently stays, which is obtained by analyzing the sequence of line of sight focus coordinates through a density clustering algorithm, and is characterized as a set of hot spot area coordinates and their weight distributions on a two-dimensional plane.
[0041] Among them, the attention concentration index specifically refers to a quantitative index of the student's cognitive engagement degree calculated by comprehensively considering the stability of the facial expression features and the eye gaze concentration index. By weighted fusion of multiple physiological behavior feature parameters, it reflects the degree of the student's attention to the classroom content and the depth of cognitive processing.
[0042] Among them, the attention historical value specifically refers to the sequence of attention concentration indexes of the student within the past time window, which preserves the time evolution record of the student's attention state and is used for the time series modeling of the state transition function.
[0043] Among them, the attention volatility specifically refers to the change range of the student's attention concentration index within a unit time, which is obtained by calculating the first derivative or difference of the attention concentration index, and reflects the stability of the student's attention.
[0044] Among them, the difficulty coefficient of the classroom content specifically refers to the quantitative evaluation of the cognitive complexity of the current teaching content, which is preset by the teacher or automatically calculated through the student feedback data, and serves as a background parameter for adjusting the attention state evaluation standard.
[0045] Among them, the attention state score specifically refers to the value that quantitatively represents the student's current cognitive engagement degree, which is calculated by the state transition function and serves as the core dimension of the dynamic portrait feature map.
[0046] Among them, the dynamic portrait feature map specifically refers to the multi-dimensional feature set that describes the student's learning state and its time evolution pattern, including time series data in multiple dimensions such as the facial expression similarity vector, the eye gaze concentration index, and the attention state score, and presents the student's cognitive and emotional states in a visual form.
[0047] Among them, the classroom interaction effect evaluation index specifically refers to the numerical index that quantifies the effectiveness of classroom teaching activities, which is obtained by analyzing the dynamic portrait feature maps of multiple students and includes multiple dimensions such as the class average attention state score, attention synchronization, and emotion distribution.
[0048] Among them, the individual association model specifically refers to the mapping function between the facial expression feature matrix established for each student and the attention concentration index, which is obtained through deep learning algorithms and can adjust the evaluation parameter weights according to individual differences among students to improve the accuracy of attention state judgment.
[0049] Among them, the state transition function is used to predict the changing trend of the student's attention state based on the multi-dimensional feature information and historical state at the current moment of the student. The inputs include the facial expression similarity vector, the eye concentration index, the frequency of special facial expressions, the attention historical value, and the classroom content difficulty coefficient. The output is the attention state score representing the current attention investment degree of the student and the prediction of its changing trend, and continuous tracking and prediction of the student's attention state are achieved through recursive calculation.
[0050] The specific structure of the learning state evaluation network model is a multi-layer bidirectional transformer network architecture, which includes an encoder layer and a decoder layer. The encoder layer is responsible for extracting the temporal features of the student's facial expression feature matrix and the eye concentration index, and the decoder layer is responsible for fusing the temporal features and outputting the student's attention state score. The core of the model is the multi-head attention mechanism. The number of heads is determined by the facial expression change frequency, the dimension of each attention head is determined by the eye concentration threshold, and the depth of the attention layer is calculated from the attention volatility. At the same time, the model adopts a residual connection structure to alleviate the problem of gradient disappearance and adds a layer normalization operation after each layer to improve the training stability. Finally, the high-dimensional features are mapped to the student's attention state score space through a fully connected layer.
[0051] The steps for establishing the training data set in the training process of the learning state evaluation network model specifically include first collecting the facial expression sequences and eye movement data of students in different subject classes from multiple schools, and at the same time recording the student attention state labels and objective test scores evaluated by teachers as supervision signals. Then, data cleaning and preprocessing are carried out, including outlier detection and removal, data standardization and enhancement. Then, the processed data is segmented according to time windows to construct sequence samples, and each sample contains a fixed-length facial expression sequence and eye movement sequence, and is accompanied by corresponding attention state labels. Finally, the data set is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 to ensure that the data of different classes and subjects are evenly distributed in the three data sets to ensure the generalization ability of the model.
[0052] The steps for training the learning status evaluation network model specifically include first initializing the model, where the weights are initialized using the truncated normal distribution initialization method. Then, a two-stage training strategy is adopted. In the first stage, self-supervised learning is used for training, and the internal representations of the facial expression sequence and the eye movement sequence are learned by designing a temporal mask prediction task. In the second stage, supervised learning is used for fine-tuning, and the student's attention state label is used as the supervision signal to optimize the model parameters. During the training process, a cosine annealing learning rate scheduling strategy is adopted, with the initial learning rate set to 0.001, which is gradually decreased according to the performance on the validation set during the training process. At the same time, gradient clipping technology is used to prevent gradient explosion, and an early stopping strategy is adopted to avoid overfitting. Training stops when the performance metric on the validation set has not improved for 5 consecutive epochs. The entire training process is carried out in parallel on four GPUs, with the batch size for each GPU set to 32, and a total of 100 epochs are trained.
[0053] The specific implementation manners of the above steps are described in detail below.
[0054] The specific implementation manner of step S01 is to use a deep convolutional neural network model for detecting and tracking the key points of the students' faces. First, a high-definition camera is used to collect videos of the students in the classroom, with the frame rate set to 30 frames per second and the resolution to 1080P to ensure that the facial features are clearly distinguishable. After the video collection is completed, a face detection algorithm is used to locate the face regions in each frame of the image. An improved MTCNN (Multi-Task Convolutional Neural Network) algorithm is used to detect the facial regions and the positions of the main organs such as the eyes, nose, and mouth simultaneously, with the detection accuracy reaching over 98%. The 68-point facial feature point localization algorithm in the improved Dlib library is applied to the detected facial regions to extract a complete set of facial feature points including eyebrows (6 points on each side), eyes (6 points on each side), nose (9 points), mouth (20 points), and contours (17 points). After the feature points are extracted, the coordinates of each feature point are normalized to eliminate the scale differences caused by the different distances between the students and the camera, and an expression feature matrix F is constructed. The dimension of the matrix is n×68×2, where n represents the number of video frames, 68 represents the number of feature points, and 2 represents the x and y coordinates of each feature point. The Kalman filter algorithm is used to smooth the feature point trajectories to eliminate the tracking noise caused by changes in lighting or small head movements, ensuring the continuity and stability of the feature point trajectories. Finally, the displacement vectors of the feature points between adjacent frames are calculated, and the movement trajectories of each feature point over time are recorded to form a complete temporal dataset of the expression feature matrix. The computational complexity of the entire process is O(n), enabling real-time processing.
[0055] The specific implementation of step S02 is to construct an expression similarity calculation and analysis framework. First, a standard expression library is established, which contains standard expression templates corresponding to 10 typical learning states, such as concentration, confusion, understanding, boredom, tiredness, etc. Each expression template is annotated by professional expression recognition experts and verified multiple times. The principal component analysis (PCA) method is used to reduce the dimensionality of the facial feature point data, extract the main patterns of expression changes, and retain the principal components that explain 95% of the variance, usually 15 - 20 principal components. For the expression feature matrix extracted from each frame of the student's facial image, calculate its cosine similarity with each template in the standard expression library to form an expression similarity vector S with a dimension of 10. The cosine similarity calculation uses an improved cosine similarity formula, introducing a feature point weight coefficient, assigning higher weights to key regions for emotion expression such as the eyes and mouth to enhance the sensitivity of expression recognition. The similarity threshold is set to 0.75, and if it is higher than this threshold, it is determined as the corresponding expression category. Based on the sliding time window technology (window size is 3 seconds, step size is 1 second), the expression change frequency is statistically analyzed. By calculating the Euclidean distance between the expression similarity vectors in adjacent time windows, when the distance exceeds the threshold of 0.3, it is recorded as an effective expression change. The number of expression changes within a unit time (such as 1 minute) is accumulated to form an expression change frequency index F. Combining the expression similarity vector S and the expression change frequency F, a dynamic change curve C of the student's classroom expression is constructed to reflect the changing trend of the student's emotional state with the classroom process. This curve uses the spline interpolation method to ensure smooth continuity for subsequent analysis and visualization.
[0056] Optionally, the standard expression templates corresponding to the 10 typical learning states are specifically as follows: Concentration: Slightly frowned eyebrows, focused and firm eyes, slightly closed lips; Confusion: Frown tightly, obvious forehead wrinkles, slightly tilted corners of the mouth, confused eyes; Understanding: Eyebrows relaxed, bright eyes, slightly raised corners of the mouth; Boredom: Dilated eyes, half - closed eyelids, expressionless face, mouth may be slightly open; Tiredness: Half - closed eyes, drooping eyelids, may yawn, loose face; Distraction: Wandering eyes, unfocused on the teaching area, absent - minded expression; Positive: Wide - open eyes, raised eyebrows, smiling face, focused expression; Absent - mindedness: Staring blankly, empty eyes, blank expression; Anxiety: Slightly frowned eyebrows, tightly closed or bitten lips, uneasy eyes; Suddenly enlightened: Suddenly wide - open eyes, raised eyebrows, slightly open mouth, suddenly enlightened expression.
[0057] The specific implementation of step S03 is to implement an eye movement tracking and fixation analysis system. First, based on the principles of infrared illumination and corneal reflection, high-precision eye movement tracking equipment (sampling rate of 120 Hz) is used to collect the eye movement data of students. The preprocessing of eye movement data includes steps such as removing blink interference, filtering out noise points, and correcting drift. The original eye movement data is smoothed using a second-order Butterworth low-pass filter (cutoff frequency of 20 Hz) to ensure signal quality. The I-VT (Velocity Threshold Identification) algorithm is used to identify the fixation points and saccade movements in the line-of-sight trajectory. The velocity threshold is set to 30° / second. Eye movements below this threshold are classified as fixations, and those above this threshold are classified as saccades. For each identified fixation point, its spatial coordinates (x, y) and duration t are recorded, forming a sequence of line-of-sight focus coordinates {(x1, y1, t1), (x2, y2, t2), …, (x n , y n , t n )}. The DBSCAN (Density-Based Spatial Clustering) algorithm is used to perform clustering analysis on the fixation points to identify the areas where the gaze is concentrated. The algorithm parameters are set as follows: ε (neighborhood radius) = 50 pixels, MinPts (minimum number of points) = 5. The weights are calculated based on the number and duration of the fixation points in each clustering area, and a two-dimensional fixation heat map H is generated. The heat map is smoothed using a Gaussian kernel function, and the standard deviation of the kernel function is set to the pixel value corresponding to a viewing angle of 1.5°. Based on the heat map, the area ratio and fixation time ratio of the areas where the gaze is concentrated are statistically analyzed, and the gaze concentration index E is calculated. The value range of this index is [0, 1], and the larger the value, the more concentrated the gaze. In practical applications, the gaze concentration threshold E0 is set to 0.6, and a state of concentrated attention is determined if the value is higher than this value.
[0058] The specific implementation of step S04 is to construct an individual association model and an attention evaluation system. An individual association model M of the expression feature matrix and the gaze concentration index is established for each student, and a long short-term memory network (LSTM) structure is used to capture the temporal dependence relationship between the expression changes and the attention state. The input of the model is the temporal data of the expression feature matrix F (a historical window with a length of T is intercepted, T = 10 seconds), and the output is the predicted attention concentration index The LSTM network consists of an input layer, a bidirectional LSTM layer (number of hidden units = 128), a fully connected layer, and an output layer. The bidirectional LSTM layer is used to capture the forward and backward temporal dependencies of facial expression changes. The model is trained using the teacher-annotated attention states as the supervision signal, with the mean squared error (MSE) as the loss function, and the parameters are optimized by the stochastic gradient descent method. To handle individual student differences, an adaptation layer is introduced into the model, which dynamically adjusts the weight distribution according to the student's historical data. The adaptation layer adopts a self-attention mechanism to adaptively adjust the importance weights of different facial expression features. The performance of the model is evaluated by the cross-validation method, and the prediction accuracy on the validation set reaches over 87%. The attention concentration index output by the model is combined with the actual eye gaze concentration index, and a comprehensive attention concentration evaluation index A is formed through weighted averaging (weight ratio is 0.6:0.4). This index reflects the overall cognitive engagement of students in classroom content, and its value range is [0, 1].
[0059] The specific implementation of step S05 is to perform specific facial expression recognition and learning state anomaly detection. Based on the statistical characteristics of the facial expression feature matrix, an individual student facial expression baseline and a class average facial expression state reference model are established. The individual facial expression baseline is obtained by calculating the mean and covariance matrix of the facial expression feature matrix of this student in recent classes (such as the previous three classes), and the class average facial expression state is calculated in real time based on the facial expression feature matrices of all students in the current class. The Mahalanobis distance is used to measure the deviation of the student's current facial expression from the baseline or the class average state. The Mahalanobis distance calculation formula takes into account the correlation between features and can more accurately reflect the abnormal state in the multi-dimensional feature space. The Mahalanobis distance threshold is set to 2.5. When the Mahalanobis distance between the student's facial expression feature matrix and the baseline exceeds this threshold, it is marked as an individual specific facial expression; when the Mahalanobis distance from the class average state exceeds the threshold, it is marked as a group specific facial expression. The specific facial expressions are classified and analyzed, and according to the facial expression similarity vector, it is judged whether it belongs to a positive reaction (such as surprise, sudden realization, etc.) or a negative reaction (such as confusion, boredom, etc.). A sliding window (window size is 5 minutes, step size is 1 minute) is used to count the occurrence frequency of specific facial expressions. When the frequency exceeds the threshold (3 times / minute), a learning state anomaly warning is triggered. The time correspondence between specific facial expressions and classroom content is established to identify the teaching content or activities that cause abnormal reactions in students, providing a basis for teaching adjustment. The anomaly detection results are added to the student's dynamic portrait in the form of timestamp annotations, forming a set of learning state anomaly point markers.
[0060] The specific implementation of step S06 is to conduct dynamic attention analysis and content association modeling. The sliding time window technique is used to analyze the time-varying characteristics of the student's attention concentration index. The window size is set to 3 minutes, and the sliding step is 30 seconds to ensure the continuity and sensitivity of the analysis. For each time window, statistical features such as the mean, standard deviation, and change rate of the attention concentration index are calculated to form an attention state description vector. Combining with the teaching content progress information, the classroom content is divided into multiple content nodes according to knowledge points or teaching activity forms, and each node is marked with start and end timestamps. The Pearson correlation coefficient is used to calculate the time correlation between the attention concentration index and the content nodes. The value range of the correlation coefficient is [-1, 1]. A positive value indicates a positive correlation (attention improvement), and a negative value indicates a negative correlation (attention decline). Based on the attention state description vector and the content node timestamps, a dynamic attention fluctuation curve is constructed. The curve uses the cubic spline interpolation method to ensure smoothness. The attention volatility is calculated, which is defined as the maximum change amplitude of the attention concentration index within a unit time (such as 1 minute). The volatility threshold is set to 0.2, and fluctuations exceeding this threshold are recorded as significant fluctuation points. Analyze the time correspondence between the attention significant fluctuation points and the teaching content nodes to identify the key content factors affecting the student's attention concentration. Establish an attention fluctuation pattern library, which includes typical patterns such as rising, falling, fluctuating, and stable patterns. Perform pattern matching and classification on the student's attention fluctuation curve to provide a basis for adjusting teaching strategies.
[0061] The specific implementation of step S07 is to implement the state transition function and the attention state evaluation system. Construct an attention state transition function based on the Hidden Markov Model (HMM). Define the attention state space S = {S1, S2, S3, S4, S5}, which correspond to five attention levels: extremely low, low, medium, high, and extremely high, respectively. Initialize the state transition matrix P based on expert experience and historical data analysis. The matrix element P ij represents the probability of transitioning from state S i to state S j . Set the observation probability matrix B. The matrix element B ij represents the probability in state S iThe probability of observing feature vector j under the following conditions. The input parameters of the state transition function include: expression similarity vector (dimension 10), eye gaze concentration index (scalar value, range [0, 1]), frequency of occurrence of specific expressions (scalar value, unit: times / minute), attention history value (sequence of attention states in the past 5 minutes), and classroom content difficulty coefficient (scalar value, range [0, 1]). The Viterbi algorithm is used to solve the most likely hidden state sequence to obtain the most likely attention state at the current moment. The state is mapped to an attention state score A, with a score range of [0, 100], where state S1 corresponds to 0 - 20 points, S2 corresponds to 21 - 40 points, S3 corresponds to 41 - 60 points, S4 corresponds to 61 - 80 points, and S5 corresponds to 81 - 100 points. For different classroom content difficulties, a dynamic threshold adjustment mechanism is set. When the content difficulty coefficient increases, the attention state judgment threshold is appropriately reduced, and vice versa. The model parameters are optimized based on historical data through the Expectation-Maximization (EM) algorithm, and the model parameters are updated once per semester to adapt to the changes in students' cognitive development.
[0062] The specific implementation of step S08 is to construct an efficient learning state evaluation network and generate a dynamic portrait atlas as the student's dynamic portrait. A multi-layer bidirectional transformer network architecture is used to implement the learning state evaluation network model. This architecture consists of 6 encoder layers and 2 decoder layers, and each layer uses a multi-head attention mechanism. The encoder layer is responsible for extracting the temporal features of the student's expression feature matrix and eye gaze concentration index. The input sequence length is 128 (corresponding to approximately 4 minutes of classroom data), and the feature dimension is 134 (including the x and y coordinates of 68 facial feature points and multi-dimensional eye movement indicators). The encoder uses position encoding technology to process the temporal information of the sequence data, and the position encoding is generated using sine and cosine functions. The number of heads of the multi-head attention mechanism is determined by the expression change frequency, and the calculation formula is where F is the expression change frequency, and usually the number of heads is between 4 and 8. The dimension d of each attention head is determined by the eye gaze concentration threshold, and the calculation formula is where E0 is the eye gaze concentration threshold, and usually the dimension is between 38 and 45. The depth L of the attention layer is calculated from the attention volatility R, and the calculation formula is Usually the depth is between 6 and 8. The model adopts a residual connection structure to alleviate the problem of gradient disappearance, and a layer normalization operation is added after each layer to improve the training stability. The normalization parameter ε is set to 10 -6 . The decoder layer is responsible for fusing the temporal features extracted by the encoder and predicting the future state through the masked self-attention mechanism. Finally, the high-dimensional features are mapped to the student attention state score space through a fully connected layer, and the output dimension is 1, representing the attention state score at the current moment.
[0063] The specific steps for establishing the training dataset of the learning state evaluation network model include first collecting the facial expression sequences and eye movement data of students in different subjects' classes from multiple schools and classes. The collection devices include high-definition cameras (resolution 1080P) and professional eye movement trackers (sampling rate 120Hz). At the same time, record the student attention state labels rated by teachers (using a 5-level scoring: 1 - completely inattentive, 2 - slightly attentive, 3 - moderately attentive, 4 - highly attentive, 5 - extremely attentive) and objective test scores (including in-class quiz scores and after-class homework scores) as supervision signals. The collection duration is not less than 10 teaching cycles for each class to ensure that the data covers different teaching contents and learning stages. Clean and preprocess the collected raw data, including outlier detection and removal (using the 3σ principle to identify outlier data points), data standardization processing (normalizing each feature quantity to the [0, 1] interval), and data augmentation (expanding training samples through techniques such as temporal perturbation, noise injection, and feature replacement). Then, slice the processed data according to time windows to construct sequence samples. The window length is set to 128 frames (about 4 minutes), and the sliding step is 32 frames (about 1 minute). Each sample contains a fixed-length expression sequence and eye movement sequence, and is accompanied by the attention state label for the corresponding time period. To balance the sample distribution of different attention states, a stratified sampling strategy is adopted to ensure that the sample ratios of each attention state are close to 1:1:1:1:1. Finally, divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1. During the division process, a stratified random sampling method is used to ensure that the data of different classes and subjects are evenly distributed in the three datasets to ensure the generalization ability of the model.
[0064] The specific steps for training the learning state evaluation network model include first initializing the model. The weights are initialized using the truncated normal distribution initialization method, with the truncation range being [-0.02, 0.02], and the bias term is initialized to 0. Then, adopt a two-stage training strategy. In the first stage, use the self-supervised learning method for training, and learn the internal representation of the expression sequence and eye movement sequence by designing a temporal mask prediction task. The mask ratio is set to 15%, that is, randomly cover 15% of the time steps in the input sequence, and let the model predict these covered features. The loss function uses the mean square error (MSE). Train for 50 epochs in the first stage, and the batch size is 32. In the second stage, use the supervised learning method for fine-tuning, and optimize the model parameters with the student attention state label as the supervision signal. The loss function uses a weighted combination of cross-entropy loss and mean square error, with the weight ratio being 0.7:0.3. During the training process, adopt the cosine annealing learning rate scheduling strategy, with the initial learning rate set to 0.001 and the minimum learning rate set to 10 -6, it gradually decreases according to the performance on the validation set during the training process. At the same time, the gradient clipping technique is used to prevent gradient explosion, and the upper limit of the gradient norm is set to 5.0. To prevent overfitting, in addition to using L2 regularization (weight decay coefficient is 10 -4 ), a Dropout layer is added after each transformer layer, and the dropout rate is set to 0.1. An early stopping strategy is adopted to avoid overfitting, and the training stops when the performance metric on the validation set does not improve for 5 consecutive epochs. The entire training process is carried out in parallel on four GPUs, adopting a data parallel strategy. The batch size for each GPU is set to 32, and a total of 100 epochs are trained. The model checkpoint is saved every 10 epochs, and finally the model with the best performance on the validation set is selected as the final model. During the training process, multiple performance metrics are monitored, including accuracy, precision, recall, F1 score, etc., to comprehensively evaluate the classification performance and regression performance of the model.
[0065] Step S09 is an optional step, and its specific implementation method is to carry out cluster analysis of the class learning status and construct a teaching feedback system. The improved K-means++ clustering algorithm is used to analyze the dynamic portrait feature map of all students in the class. The number of clusters K is adaptively determined by the silhouette coefficient method, usually between 3 and 5. The clustering feature dimensions include core indicators such as attention state score, expression change frequency, and eye concentration index. The feature vectors are processed by Z-score standardization to eliminate the dimension difference. Calculate the distance between each cluster center and the ideal learning state (high attention, moderate expression change, high eye concentration), and define the cluster with the smallest distance as the best learning state group, and the cluster with the largest distance as the group to be improved. Count the proportion of the number of students in each cluster to form a class learning status distribution map. Based on the clustering analysis results, calculate the evaluation indicators of classroom interaction effects, including the class average attention state score (value range 0-100), attention synchronization index (defined as the standard deviation of the attention state scores of all students in the class, the smaller the higher the synchronization), and emotion distribution balance degree (defined as the entropy value of the proportion of each type of emotion, the larger the more balanced the emotion distribution). Set the threshold of the evaluation indicators, the threshold of the average attention state score is 70 points, the threshold of the attention synchronization index is 15, and the threshold of the emotion distribution balance degree is 0.8. When the indicators are lower than the thresholds, corresponding teaching adjustment suggestions are triggered, such as increasing interaction links, adjusting the teaching rhythm, and refining the explanation of difficult points. Establish a teaching suggestion push mechanism, and on the premise of not disturbing the normal teaching, provide real-time teaching optimization suggestions through the teacher's end display screen, including class overall state analysis, problem area identification, improvement strategy recommendation, etc., to assist teachers in adjusting teaching strategies in a timely manner and realizing the closed-loop optimization of the teaching process.
[0066] In a second aspect of the present invention, a computer-readable storage medium is provided. Program instructions are stored in the computer-readable storage medium, and when the program instructions run on a computer, they are used to execute the above-mentioned method for generating a dynamic portrait of students in an interactive classroom.
[0067] In a third aspect of the present invention, a system for generating a dynamic portrait of students in an interactive classroom is provided, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.
[0068] The following will describe in detail the mathematical models or calculation processes involved in the present invention.
[0069] In step S01, when collecting the sequence of student facial images and extracting the facial key point data, the construction of an expression feature matrix is involved. The specific representation of the expression feature matrix is as follows:
[0070]
[0071] In the formula, F is a sequence of expression feature matrices; F t is the expression feature matrix of the t-th frame; n is the total number of video frames.
[0072] The construction formula for a single-frame expression feature matrix is:
[0073]
[0074] In the formula, (x i,t , y i,t ) is the normalized coordinate of the i-th feature point in the t-th frame image; 68 represents the total number of facial feature points.
[0075] The calculation formula for coordinate normalization processing is:
[0076]
[0077] In the formula, (x i,t ′, y i,t ′) is the normalized coordinate; (x c,t , y c,t ) is the coordinate of the center point of the human face; d t is the face scale factor, usually taking the face width.
[0078] The feature point trajectory smoothing adopts the Kalman filtering algorithm, and its state equation and observation equation are:
[0079] s t+1 = As t + w t ;
[0080] z t = Hs t + v t ;
[0081] In the formula, is the state vector, including the feature point position and velocity; z t = [x i,t , y i,t T is the observation vector; A is the state transition matrix; H is the observation matrix; w t is the process noise, following a Gaussian distribution with mean zero and covariance Q; v t is the observation noise, following a Gaussian distribution with mean zero and covariance R.
[0082] The calculation formula for the feature point displacement vector is:
[0083] d i,t = [x i,t+1 - x i,t , y i,t+1 - y i,t ;
[0084] In the formula, d i,t is the displacement vector of the i-th feature point from time t to time t + 1.
[0085] The calculation of the expression similarity vector in step S02 involves the calculation of similarity with the standard expression library. The specific representation of the expression similarity vector is as follows:
[0086] S t = [s t,1 , s t,2 ,..., s t,10 T ;
[0087] In the formula, S t is the expression similarity vector at time t; s t,j is the similarity between the expression at time t and the j-th type of expression in the standard expression library, and the value range is [0, 1].
[0088] The calculation of expression similarity uses weighted cosine similarity, and the calculation formula is:
[0089]
[0090] In the formula, p i,t is the position vector of the i-th feature point at time t; p i,j is the position vector of the i-th feature point of the j-th type of expression in the standard expression library; w i is the weight coefficient of the i-th feature point. The weight of the feature points in the eye and mouth regions is higher than that in other regions, and the value range is [0.5, 2.0].
[0091] The calculation formula for the expression change frequency is as follows:
[0092]
[0093] In the formula, F is the expression change frequency; D(S t , S t+1 ) is the Euclidean distance of the expression similarity vector between adjacent time windows; δ(·) is an indicator function, which takes the value of 1 when the condition is satisfied, otherwise 0; τ d is the expression change threshold, and the value is 0.3; T is the number of time windows.
[0094] The calculation formula for the Euclidean distance of the expression similarity vector is as follows:
[0095]
[0096] The calculation of the eye gaze concentration index in step S03 involves the processing and analysis of eye movement data. The fixation point recognition adopts the I-VT algorithm, and the calculation formula is as follows:
[0097]
[0098] In the formula, v t is the eye movement speed at time t; (x t , y t ) is the line of sight position at time t; Δt is the sampling time interval. For an eye tracker with a sampling rate of 120 Hz, Δt = 1 / 120 second.
[0099] When v t < τ v , it is determined to be in the fixation state, where τ v is the speed threshold, and the value is 30° / s.
[0100] The generation of the fixation heat map adopts a two-dimensional Gaussian kernel function, and the calculation formula is as follows:
[0101]
[0102] In the formula, H(x, y) is the heat value at the coordinate (x, y); (x i , y i ) is the coordinate of the i-th fixation point; t i is the duration of the i-th fixation point; σ is the standard deviation of the Gaussian kernel function, corresponding to the pixel value of a viewing angle of 1.5°; N is the total number of fixation points.
[0103] The calculation formula for the eye gaze concentration index is as follows:
[0104]
[0105] In the formula, E is the eye gaze concentration index, and its value range is [0, 1]; Ω is the area in the heat map where the heat value exceeds the threshold; T f is the total time of the fixation state; T is the total observation time.
[0106] In step S05, the specific expression recognition uses the Mahalanobis distance to measure the deviation degree of the student's current expression from the baseline or the class average state. The calculation formula is:
[0107]
[0108] In the formula, D M (F t , F b ) is the Mahalanobis distance between the expression feature matrix F t and the baseline expression matrix F b ; Σ is the covariance matrix, which reflects the correlation between features.
[0109] The calculation formula for the occurrence frequency of specific expressions is:
[0110]
[0111] In the formula, f a is the occurrence frequency of specific expressions; δ(·) is the indicator function; τ m is the Mahalanobis distance threshold, and its value is 2.5; T is the time window length.
[0112] In step S07, the state transition function is based on the hidden Markov model. The calculation formulas for its state transition matrix and observation probability matrix are as follows:
[0113] P ij = P(S t+1 = j|S t = i);
[0114] In the formula, P ij is an element of the state transition matrix, representing the probability of transitioning from state i to state j; S t represents the attention state at time t.
[0115] B ij = P(O t = j|S t = i);
[0116] In the formula, B ij is an element of the observation probability matrix, representing the probability of observing the feature vector j in state i; O t represents the observation vector at time t.
[0117] The observation vector contains multiple features, expressed as:
[0118] O t =[S t ,E t ,f a,t ,A t-1 ,d t ;
[0119] In the formula, S t is the expression similarity vector; E t is the eye gaze concentration index; f a,t is the frequency of occurrence of specific expressions; A t-1 is the attention state score at the previous moment; d t is the difficulty coefficient of the classroom content.
[0120] The calculation formula for the attention state score is:
[0121]
[0122] In the formula, A t is the attention state score at time t, and its value range is [0, 100]; α i is the scoring weight corresponding to state i, α1 = 10, α2 = 30, α3 = 50, α4 = 70, α5 = 90; P(S t =i|O 1:t ) is the posterior probability that the state at time t is i under the condition of the given observation sequence O 1:t , which is calculated by the forward-backward algorithm.
[0123] The calculation formula for the dynamic threshold adjustment mechanism is:
[0124]
[0125] In the formula, τ A (d t ) is the adjusted attention state threshold; is the basic threshold, with a value of 60; β is the adjustment coefficient, with a value of 0.3; d t is the difficulty coefficient of the classroom content, and its value range is [0, 1].
[0126] In step S08, the multi-head attention mechanism of the learning state evaluation network model involves the calculation of the number of heads, dimension, and layer depth. The calculation formula is as follows:
[0127]
[0128] In the formula, h is the number of heads of the multi-head attention mechanism; F is the frequency of expression change; represents the floor operation.
[0129]
[0130] Wherein, d is the dimension of each attention head; E0 is the eye gaze concentration threshold, with a value of 0.6.
[0131]
[0132] Wherein, L is the depth of the attention layer; R is the attention volatility.
[0133] The self-attention calculation formula of the Transformer network is:
[0134]
[0135] Wherein, Q is the query matrix; K is the key matrix; V is the value matrix; d k is the dimension of the key vector.
[0136] The calculation formula of multi-head attention is:
[0137] MultiHead(Q, K, V) = Concat(head1,..., head h )W O ;
[0138] Wherein, and W O are learnable weight matrices.
[0139] The calculation of the classroom interaction effect evaluation index in step S09 involves cluster analysis and the calculation of statistical indexes, which are specifically expressed as follows:
[0140] The calculation formula of the attention synchronization index is:
[0141]
[0142] Wherein, ASI is the attention synchronization index; A i is the attention state score of the i-th student; is the class average attention state score; N is the total number of students.
[0143] The calculation formula of the emotion distribution balance degree is:
[0144]
[0145] Wherein, EDB is the emotion distribution balance degree; p j is the proportion of the j-th type of emotion in the class; 10 is the total number of emotion categories.
[0146] The K-means++ algorithm is used for cluster analysis, and its objective function is:
[0147]
[0148] In the formula, J is the clustering objective function; x i is the feature vector of the i-th student; μ j is the j-th clustering center; K is the number of clusters; N is the total number of students.
[0149] The calculation formula for determining the optimal number of clusters K by the silhouette coefficient method is:
[0150]
[0151] In the formula, s(i) is the silhouette coefficient of the i-th sample; a(i) is the average distance between sample i and other samples in the same cluster; b(i) is the average distance between sample i and samples in the nearest different cluster.
[0152] The calculation formula for the average silhouette coefficient is:
[0153]
[0154] In the formula, is the average silhouette coefficient, and its value range is [-1, 1]. The larger the value, the better the clustering effect.
[0155] The calculation formula for the distance between the ideal learning state and the clustering center is:
[0156]
[0157] In the formula, D(c j , c * ) is the weighted Euclidean distance between the j-th clustering center c j and the ideal learning state c * ; c j,k is the k-th feature of the j-th clustering center; is the k-th feature of the ideal learning state; w k is the weight of the k-th feature; d is the feature dimension.
[0158] Most of the parameters involved in these formulas are obtained through experiments. For example, the eye gaze concentration threshold E0 is set to 0.6 by analyzing eye movement data, and a state with a value higher than this is determined to be an attention-concentrated state; the Mahalanobis distance threshold τ m is set to 2.5 through statistical analysis, and an expression exceeding this threshold is marked as a specific expression; the attention state scoring threshold is set to 60 based on the experience of educational experts and serves as a reference value for judging the attention state of students. Other parameters such as the expression feature point weight w i , the state transition matrix P, and the observation probability matrix B need to be determined through statistical analysis of a large amount of historical data and machine learning methods.
[0159] Specifically, the principle of the present invention is as follows: The core technical principle of the present invention lies in establishing a multi-level and multi-dimensional student attention evaluation framework. First, through facial feature point localization technology, key point data of the student's face is extracted to construct an expression feature matrix, capturing the positions and changes of key points such as eyebrows, eyes, and the corners of the mouth, and calculating the expression similarity vector by comparing with the standard expression library to quantify the student's emotional state.
[0160] At the same time, eye movement tracking technology is used to collect the sequence of the student's line of sight focus coordinates, calculate the area of eye concentration and the retention time, and generate a fixation heat map to quantitatively characterize the stability and distribution characteristics of the student's visual attention. This dual-channel data acquisition mechanism ensures a comprehensive perception of the student's attention state.
[0161] The innovation of the present invention lies in introducing deep learning to construct an individual association model to solve the individual differences in the expression-attention mapping relationship of different students. Through a multi-layer bidirectional transformer network architecture, the extraction and fusion of the temporal features of the expression feature matrix and the eye concentration index are realized, where the parameters of the multi-head attention mechanism are dynamically determined by the expression change frequency, the eye concentration threshold, and the attention volatility, enabling the model to adapt to the behavior patterns of different students.
[0162] Another core innovation is to apply a state transition function to dynamically model the student's attention fluctuations. This function comprehensively considers multi-dimensional parameters such as the expression similarity vector, the eye concentration index, the frequency of specific expressions, the attention historical value, and the classroom content difficulty coefficient, and realizes continuous tracking and prediction of the student's attention state through recursive calculation, overcoming the limitations of static evaluation methods.
[0163] Through the organic combination of these technical principles, the present invention realizes the accurate, dynamic, and personalized evaluation of the student's attention state, providing a scientific basis for optimizing classroom interaction.
[0164] A specific embodiment 1 of the present invention is provided below, and the specific implementation manners of each step in this embodiment 1 are described in detail as follows.
[0165] The specific implementation of step S01 is to use a deep convolutional neural network model for student facial key point detection and tracking. First, video of the students in the classroom is collected through a high-definition camera. The collection frame rate is set to 30 frames per second, and the resolution is 1080P to ensure that facial features are clearly distinguishable. After the video collection is completed, a face detection algorithm is used to locate the face area in each frame of the image. An improved MTCNN (Multi-Task Convolutional Neural Network) algorithm is used to detect the facial area and the positions of the main organs such as eyes, nose, and mouth simultaneously, and the detection accuracy reaches more than 98%. The improved 68-point facial feature point localization algorithm in the Dlib library is applied to the detected facial area to extract a complete set of facial feature points including eyebrows (6 points on each side), eyes (6 points on each side), nose (9 points), mouth (20 points), and contour (17 points). The extracted facial feature points form an expression feature matrix sequence, denoted as:
[0166]
[0167] where F is the expression feature matrix sequence; F t is the expression feature matrix of the t-th frame; n is the total number of frames in the video. The single-frame expression feature matrix is constructed as:
[0168]
[0169] where (x i,t , y i,t ) is the normalized coordinate of the i-th feature point in the t-th frame of the image. The coordinate normalization process uses the following formula:
[0170]
[0171] where (x i,t ′, y i,t ′) is the normalized coordinate; (x c,t , y c,t ) is the coordinate of the center point of the face; d t is the face scale factor. The Kalman filter algorithm is used to smooth the feature point trajectory, and its state equation and observation equation are:
[0172] s t+1 = As t + w t ;
[0173] z t = Hs t + v t ;
[0174] where is the state vector; z t = [x i,t , y i,tT is the observation vector; A is the state transition matrix; H is the observation matrix; w t is the process noise; v t is the observation noise. Finally, calculate the displacement vector of the feature points:
[0175] d i,t = [x i,t+1 - x i,t , y i,t+1 - y i,t ;
[0176] In the formula, d i,t is the displacement vector of the i-th feature point from time t to time t + 1.
[0177] The specific implementation of step S02 is to construct a facial expression similarity calculation and analysis framework. First, establish a standard facial expression library, which contains standard facial expression templates corresponding to 10 typical learning states such as concentration, confusion, understanding, boredom, and fatigue. Use the principal component analysis method to reduce the dimensionality of the facial feature point data and extract the main patterns of facial expression changes. For the facial expression feature matrix extracted from each frame of the student's facial image, calculate its cosine similarity with each template in the standard facial expression library to form a facial expression similarity vector:
[0178] S t = [s t,1 , s t,2 ,..., s t,10 T ;
[0179] In the formula, S t is the facial expression similarity vector at time t; s t,j is the similarity between the facial expression at time t and the j-th type of facial expression in the standard facial expression library. The calculation of facial expression similarity uses weighted cosine similarity:
[0180]
[0181] In the formula, p i,t is the position vector of the i-th feature point at time t; p i,j is the position vector of the i-th feature point of the j-th type of facial expression in the standard facial expression library; w i is the weight coefficient of the i-th feature point. Statistically analyze the frequency of facial expression changes:
[0182]
[0183] In the formula, F is the frequency of facial expression changes; D(S t , S t+1 ) is the Euclidean distance between adjacent time window facial expression similarity vectors; δ(·) is the indicator function; τ d is the expression change threshold, with a value of 0.3; T is the number of time windows. The Euclidean distance is calculated as follows:
[0184]
[0185] The specific implementation of step S03 is to implement an eye movement tracking and fixation analysis system. High-precision eye movement tracking equipment (sampling rate of 120Hz) is used to collect the eye movement data of students. After preprocessing the eye movement data, the I-VT (Velocity Threshold Identification) algorithm is used to identify the fixation points and saccadic movements in the line of sight trajectory:
[0186]
[0187] In the formula, v t is the eye movement speed at time t; (x t , y t ) is the line of sight position at time t; Δt is the sampling time interval. The identified fixation points form a sequence of line of sight focus coordinates, and the DBSCAN algorithm is used to cluster and analyze the fixation points to identify the area where the eyes are concentrated. A fixation heat map is generated:
[0188]
[0189] In the formula, H(x, y) is the heat value at coordinates (x, y); (x i , y i ) is the coordinate of the i-th fixation point; t i is the duration of the i-th fixation point; σ is the standard deviation of the Gaussian kernel function; N is the total number of fixation points. Calculate the eye concentration index:
[0190]
[0191] In the formula, E is the eye concentration index; Ω is the area in the heat map where the heat value exceeds the threshold; T f is the total time of the fixation state; T is the total observation time.
[0192] The specific implementation of step S04 is the same as that described above and will not be elaborated here.
[0193] The specific implementation of step S05 is to perform specific expression recognition and learning state anomaly detection. Establish a reference model for the individual expression baseline and class average expression state of students. The Mahalanobis distance is used to measure the deviation of the current expression of students from the baseline or class average state:
[0194]
[0195] In the formula, D M (F t , F b ) is the expression feature matrix Ft The Mahalanobis distance from the baseline expression matrix F b ; Σ is the covariance matrix. Calculate the occurrence frequency of specific expressions:
[0196]
[0197] In the formula, f a is the occurrence frequency of specific expressions; τ m is the Mahalanobis distance threshold, with a value of 2.5; T is the time window length. Analyze the types of specific expressions and establish the time correspondence with the classroom content.
[0198] The specific implementation of step S06 is the same as the foregoing, and will not be elaborated here.
[0199] The specific implementation of step S07 is to implement the state transition function and the attention state evaluation system. Construct an attention state transition function based on the hidden Markov model, and define the attention state space S = {S1, S2, S3, S4, S5}. The state transition matrix and the observation probability matrix are calculated as follows:
[0200] P ij = P(S t+1 = j|S t = i);
[0201] B ij = P(O t = j|S t = i);
[0202] In the formula, P ij is an element of the state transition matrix; B ij is an element of the observation probability matrix; S t represents the attention state at time t; O t represents the observation vector at time t. The observation vector contains multiple features:
[0203] O t = [S t , E t , f a,t , A t-1 , d t ;
[0204] In the formula, S t is the expression similarity vector; E t is the eye gaze concentration index; f a,t is the occurrence frequency of specific expressions; A t-1 is the attention state score at the previous moment; d t is the classroom content difficulty coefficient. Calculate the attention state score:
[0205]
[0206] Wherein, A t is the attention state score at time t; α i is the score weight corresponding to state i; P(S t =i|O 1:t ) is the posterior probability that the state at time t is i under the condition of the given observation sequence O 1:t . The calculation formula of the dynamic threshold adjustment mechanism is:
[0207]
[0208] Wherein, τ A (d t ) is the adjusted attention state threshold; is the basic threshold, with a value of 60; β is the adjustment coefficient, with a value of 0.3; d t is the difficulty coefficient of the classroom content.
[0209] The specific implementation manner of step S08 is the same as the foregoing, and will not be elaborated herein.
[0210] The specific implementation manner of step S09 is to carry out the clustering analysis of the class learning state and the construction of the teaching feedback system. The improved K-means++ clustering algorithm is used to analyze the dynamic portrait feature map of the whole class of students, and the number of clusters K is adaptively determined by the silhouette coefficient method. Calculate the attention synchronization index:
[0211]
[0212] Wherein, ASI is the attention synchronization index; A i is the attention state score of the i-th student; is the average attention state score of the class; N is the total number of students. Calculate the emotional distribution balance degree:
[0213]
[0214] Wherein, EDB is the emotional distribution balance degree; p j is the proportion of the j-th type of emotion in the class. The objective function of the clustering analysis is:
[0215]
[0216] Wherein, J is the clustering objective function; x i is the feature vector of the i-th student; μ j is the j-th clustering center; K is the number of clusters. The silhouette coefficient is calculated as:
[0217]
[0218] Where s(i) is the silhouette coefficient of the i-th sample; a(i) is the average distance between sample i and other samples in the same cluster; b(i) is the average distance between sample i and samples in the nearest different cluster. Calculate the distance between the ideal learning state and the cluster center:
[0219]
[0220] Where D(c j , c * ) is the weighted Euclidean distance between the j-th cluster center c j and the ideal learning state c * ; c j,k is the k-th feature of the j-th cluster center; is the k-th feature of the ideal learning state; w k is the weight of the k-th feature. Calculate the evaluation index of classroom interaction effect based on the clustering analysis results, set the threshold of the evaluation index, and construct a teaching suggestion push mechanism.
[0221] Optionally, the Kalman filter algorithm plays a key role in smoothing the facial feature point trajectory. Its state equation s t+1 = As t + w t describes the motion model of the feature points, and the observation equation z t = Hs t + v t relates the actual observation value to the state variable. Through the prediction-update recursive process, the algorithm effectively filters out the noise caused by factors such as light changes and small head movements, ensuring the continuity and stability of the feature point trajectory. The state vector contains both position and velocity information at the same time, enabling the algorithm to predict the future position of the feature points and adapt to the rapid changes of facial expressions.
[0222] Optionally, the expression similarity vector S t = [s t,1 , s t,2 ,..., s t,10 T reflects the similarity between the current expression of the student and various expressions in the standard expression library, and is the basis for emotional state analysis. The weighted cosine similarity formula emphasizes the importance of key facial regions (such as eyes and mouth) in emotional expression by introducing the feature point weight coefficient w i , improving the accuracy of expression recognition. The weight coefficient is determined according to facial anatomy knowledge and emotional expression rules. The weights of the eyes and mouth are usually 1.5 - 2.0, and those of other regions are 0.5 - 1.0.
[0223] Optionally, the expression change frequency Quantifies the degree of fluctuation of the emotional state and is an important indicator of students' cognitive and emotional engagement. Through the Euclidean distance Calculate the difference between the facial expression similarity vectors of adjacent time windows, and use the threshold τ d = 0.3 to determine whether an effective facial expression change is constituted. The indicator function δ(·) converts the continuous distance value into a discrete count for easy statistical analysis.
[0224] Optionally, the speed threshold recognition (I-VT) algorithm in eye movement analysis distinguishes the fixation state and saccade movement by calculating the line-of-sight movement speed The speed threshold τ v = 30° / s is determined based on the physiological characteristics of the human eye. Eye movements below this threshold are classified as fixations, indicating that the student is extracting information and performing cognitive processing in the specified area; those above this threshold are classified as saccades, representing the process of rapid eye movement.
[0225] Optionally, the fixation heatmap Intuitively presents the distribution of line-of-sight attention and is the core tool for analyzing eye gaze concentration. Based on the two-dimensional Gaussian kernel function, each fixation point is spread into a hot spot area, and the fixation time t i is used as the weight coefficient to reflect the fixation intensity; the standard deviation σ corresponds to the pixel value of a viewing angle of 1.5°, usually between 30 and 50 pixels, which determines the spread range of the hot spot. The high-intensity area on the heatmap represents the focus of visual attention, facilitating the identification of the teaching content that students are concerned about.
[0226] Optionally, the eye gaze concentration index Comprehensively considers two dimensions: spatial concentration and temporal stability. The spatial concentration is reflected by the proportion of the high-intensity area Ω in the heatmap, and the temporal stability is represented by the proportion of the fixation state time This index ranges from [0, 1], and the larger the value, the more concentrated the fixation. In practical applications, the eye gaze concentration threshold E0 = 0.6 is used as the basis for judging the state of attention concentration.
[0227] Optionally, the Mahalanobis distance Has unique advantages in specific facial expression recognition. It considers the correlation between features (through the covariance matrix Σ) and can accurately capture abnormal states in the multi-dimensional feature space. Compared with the Euclidean distance, the Mahalanobis distance is insensitive to feature scales, can balance the contributions of different facial regions, and improve the accuracy of anomaly detection. The frequency of specific facial expressions By statistically analyzing the proportion of time when the Mahalanobis distance exceeds the threshold τ m = 2.5, the degree of students' cognitive conflict or emotional fluctuation is quantified.
[0228] Optionally, the Hidden Markov Model is applied in the attention state evaluation, and the probability relationship between the observed features and the potential attention states is established through the state transition matrix P ij =
[0229] P(S t+1 = j|S t = i) and the observation probability matrix B ij = P(O t = j|S t = i). This model takes into account the temporal dependence of the attention states, avoids frequent state jumps caused by instantaneous feature fluctuations, and enhances the stability and reliability of the evaluation results. The observation vector O t = [S t , E t , f a,t , A t-1 , d t integrates multi-dimensional information such as expressions, eye movements, and historical states, comprehensively capturing all aspects of students' cognitive engagement.
[0230] Optionally, the attention state scoring converts the discrete state probability distribution into a continuous scoring value, where α i is the scoring weight corresponding to each state, determined based on educational psychology research. The dynamic threshold adjustment mechanism considers the influence of the classroom content difficulty on the attention judgment criteria. When the content difficulty coefficient d t increases, the judgment threshold is appropriately reduced, reflecting the application of the cognitive load theory.
[0231] Optionally, the parameter design of the multi-head attention mechanism in the learning state evaluation network reflects a deep consideration of students' individual characteristics. The number of heads is related to the expression change frequency F. A higher expression change frequency corresponds to more attention heads, enhancing the model's ability to capture complex expression change patterns; the dimension of each attention head is related to the eye gaze concentration threshold E0, reflecting the importance of eye movement features in attention evaluation; the depth of the attention layer is related to the attention volatility R. A higher volatility corresponds to a deeper hierarchical structure, enhancing the model's ability to handle unstable attention states. This parameter adaptive design enables the model to dynamically adjust the network structure according to students' behavioral characteristics, improving the accuracy and personalization level of the evaluation.
[0232] Optionally, the attention synchronization index Essentially, it is the standard deviation of the attention state scores of students in the class, which quantifies the degree of consistency of the class learning state. A lower ASI value indicates high synchronization, reflecting that the teaching content or activities can attract the attention of most students simultaneously; a higher ASI value indicates low synchronization, which may mean uneven teaching difficulty or some students cannot keep up with the teaching rhythm. This indicator provides a quantitative assessment of the overall engagement of the class for teachers, and the threshold is set at 15. If this value is exceeded, teaching strategies should be considered for adjustment.
[0233] Optionally, the emotional distribution balance Adopts the concept of information entropy to measure the diversity and balance of the class emotional state. p j Represents the proportion of the jth type of emotion in the class. An evenly distributed emotional state corresponds to a high EDB value, and a concentrated distribution corresponds to a low EDB value. The ideal range of this indicator is 0.7 - 0.9. Too low indicates a single emotion (either negative or positive), and too high may mean a chaotic classroom atmosphere and lack of clear emotional direction. Combining the ASI and EDB indicators can comprehensively evaluate the classroom interaction effect.
[0234] Optionally, the objective function of the K-means++ clustering algorithm By minimizing the sum of the squared distances from each sample to the nearest cluster center, it realizes the automatic grouping of students' learning states. Compared with the traditional K-means algorithm, K-means++ improves the selection strategy of the initial center point, avoids local optimal solutions, and improves the clustering quality. The silhouette coefficient As an evaluation indicator of clustering effectiveness, it helps to determine the optimal number of clusters K. An average silhouette coefficient close to 1 indicates good clustering effect, with each cluster being tight and separated from each other; close to 0 indicates that the samples are near the cluster boundaries; negative values indicate possible misassignment.
[0235] Optionally, the distance between the ideal learning state and the cluster center Quantifies the degree of difference between each cluster and the best learning state through the weighted Euclidean distance. The weight coefficient w k Reflects the importance of different feature dimensions. Usually, the weight of the attention state score is the highest (about 0.5), and the weights of the expression change frequency and eye concentration indicators are the second highest (each about 0.25). This distance metric helps to identify the best learning state group and the group to be improved, providing a basis for targeted teaching intervention for teachers.
[0236] Optionally, the key difference between the expression feature matrix extraction technology and the traditional facial expression recognition method is that the traditional methods are mostly based on the analysis of static expression images, while the expression feature matrix constructed in this solution It contains complete timing information and can capture the dynamic change process of expressions. This dynamic expression analysis method is more in line with the true characteristics of human emotion expression, effectively avoiding the ambiguity problem in static expression recognition and improving the accuracy of emotion state judgment.
[0237] Optionally, the weighted cosine similarity formula Compared with the traditional cosine similarity, the feature point weight coefficient w is introduced i . This improvement stems from the research in facial expression science, which shows that different facial regions contribute unequally to emotion expression. The eyes and mouth play a dominant role in most basic emotion expressions, so higher weights (usually 1.5 - 2.0) are given, while regions such as the facial contour have lower weights (usually 0.5 - 1.0). This weighted strategy significantly improves the sensitivity and accuracy of expression recognition, especially for subtle emotion changes in the classroom environment.
[0238] Furthermore, the calculation formula of the eye gaze concentration index innovatively combines two dimensions: spatial concentration and temporal stability. Traditional eye movement analysis mostly focuses on the distribution of fixation points or fixation time, while ignoring the combined effect of these two dimensions. Spatial concentration reflects the size of the visual attention range, and temporal stability reflects the degree of fixation duration. The two multiplied rather than simply added reflects the mutually enhancing relationship between the two dimensions. Only when the line of sight is concentrated in space and stable in time can it be considered a highly focused state.
[0239] Furthermore, the application of the hidden Markov model in attention state assessment reflects an in - depth understanding of the dynamic change law of attention. The attention state has obvious temporal dependence, that is, the characteristic that the current state is affected by the previous state and will not jump frequently in a short time. The state transition matrix P ij = P(S t+1 = j|S t = i) captures this temporal law, making the assessment result smoother and more stable. The observation vector O t = [S t , E t , f a,t , A t-1 , d t integrates multiple heterogeneous features, comprehensively reflecting the cognitive and emotional states of students and avoiding the one - sidedness that may be brought by a single index.
[0240] Furthermore, the dynamic threshold adjustment mechanism Based on the cognitive load theory, the influence of the difficulty of classroom content on attention performance is considered. When the difficulty of learning materials increases, students need to invest more cognitive resources in information processing, and the external characteristics of attention shown may be weakened, such as an increase in blink frequency, nervous facial expressions, etc. By dynamically reducing the judgment threshold, the problem of misjudging the attention state of students due to content difficulty is avoided, making the evaluation results more fair and reasonable.
[0241] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: Researchers applied the interactive classroom student dynamic portrait generation method in the eighth-grade mathematics classroom of a certain middle school and conducted an observation and analysis of 32 students in a class for 4 weeks. The class was equipped with a high-definition camera and an eye movement tracking device. The camera had a frame rate of 30 frames per second and a resolution of 1080P; the sampling rate of the eye movement tracking device was 120Hz. Before the start of the research, the researchers collected the basic data of the students and established an initial expression baseline. During the classroom observation, the system collected the student facial image sequence and eye movement data in real time, constructed an expression feature matrix and a line-of-sight focus coordinate sequence, and then generated a student dynamic portrait feature map.
[0242] In a course on quadratic functions, the researchers detailedly recorded the changes in the attention state of students throughout the teaching process. The course was divided into five stages: review and introduction (10 minutes), new concept explanation (15 minutes), example analysis (15 minutes), group discussion (15 minutes), and summary test (10 minutes). The system constructed an expression feature matrix F by extracting the data of 68 key points on the students' faces. t For a representative student Zhang, the calculation results of the expression similarity vectors at different times are shown in Table 1:
[0243] Table 1 Expression similarity vectors of student Zhang at different teaching stages
[0244]
[0245] By analyzing the expression similarity vector data, the system identified that Zhang showed a relatively high level of confused expressions (similarities were 0.79 and 0.83 respectively) during the new concept explanation and example analysis stages, significantly higher than the threshold of 0.75, indicating that he had difficulties in understanding the concept of quadratic functions. At the same time, the expression change frequency F was recorded, which reached the highest value of 3.6 times per minute during the example analysis stage, much higher than the class average of 2.1 times per minute, further confirming the confused state.
[0246] In terms of eye movement analysis, the system identified Zhang's fixation points and saccade movements through the I-VT algorithm, and the calculated eye gaze concentration index E is shown in Table 2:
[0247] Table 2 Eye movement feature data of student Zhang in different teaching stages
[0248]
[0249]
[0250] Table 2 shows that Zhang's eye concentration indices in the new concept explanation and example analysis stages are 0.52 and 0.48 respectively, both lower than the threshold of 0.6, indicating inattentiveness. Combining the heat map analysis, it is found that his sight frequently jumps between the blackboard and the textbook, and no effective fixation is formed on the key formulas and graphs.
[0251] The researchers calculated the deviation degree between Zhang's facial expression feature matrix and his personal baseline using the Mahalanobis distance formula to identify specific expressions. At the 8th minute of the example analysis stage, when the teacher was explaining the calculation method of the quadratic function's maximum value, the Mahalanobis distance value was detected to be 3.2, exceeding the threshold of 2.5, and the system marked it as a specific expression point. By analyzing the corresponding relationship with the facial expression similarity vector, it was determined to be a composite emotion of "confusion + surprise", indicating that this knowledge point may be Zhang's learning obstacle.
[0252] Based on the state transition function of the hidden Markov model, the system comprehensively combines facial expression features, eye movement data, and historical states to calculate Zhang's attention state score A during the entire course. t As shown in Table 3:
[0253] Table 3 Comparison of the changes in student Zhang's attention state score and the class average level
[0254] Time point (minutes) Attention state score of Zhang Average attention state score of the class Attention state evaluation 5 78 72 High 15 56 68 Medium 25 43 65 Low 35 61 74 Medium 45 75 76 High 55 82 79 Extremely high
[0255] The data in Table 3 show that Zhang's attention state decreased significantly in the new concept explanation and example analysis stages (15 - 25 minutes), and the gap with the class average level widened. Subsequently, it rebounded in the group discussion and summary test stages and finally reached an extremely high level. This is highly consistent with the analysis results of the facial expression similarity vector and the eye concentration index.
[0256] The researchers applied a multi-layer bidirectional transformer network for learning state evaluation. The network parameters were set as follows: the number of multi-head attention heads h = 6 (calculated based on the facial expression change frequency F = 2.8), the dimension d of each attention head = 38 (calculated based on the eye concentration threshold E0 = 0.6), and the depth L of the attention layer = 7 (calculated based on the attention volatility R = 0.18). The model training adopted a two-stage strategy. In the first stage, self-supervised learning was carried out for 50 epochs, and in the second stage, supervised fine-tuning was carried out for 100 epochs. The initial learning rate was 0.001, and the batch size was 32.
[0257] The dynamic portrait feature map of 32 students in the class was analyzed by the K-means++ clustering algorithm, and the optimal number of clusters K = 4 was determined by the silhouette coefficient method. The clustering results and the characteristics of each group are shown in Table 4:
[0258] Table 4 Clustering analysis results of the learning status of class students
[0259]
[0260]
[0261] Zhang was classified into Group C (the group to be improved), and the system provided targeted teaching suggestions for the teacher: (1) Focus on explaining the difficulties in calculating the maximum and minimum values of quadratic functions; (2) Increase intuitive graphics to assist understanding; (3) Arrange students in Group A and Group C to carry out group cooperative learning. The teacher adjusted the teaching according to the suggestions. In the subsequent courses, Zhang's attention state score increased significantly, reaching an average of 72.5 points, the frequency of facial expression changes decreased to 2.0 times per minute, and the eye concentration index increased to 0.68.
[0262] Compared with the traditional methods for evaluating students' attention, the method for generating the dynamic portrait of students in the interactive classroom shown in this embodiment has significant advantages. Traditional methods mainly rely on teachers' subjective observations or simple behavior counting (such as the number of times of raising hands, nodding frequency, etc.), and cannot capture subtle facial expression changes and eye movement characteristics. The evaluation results are highly subjective and lack timeliness. While this method extracts 68 key point data through facial feature point localization technology, combines eye movement tracking and multi-dimensional feature fusion, and realizes objective, accurate, and real-time evaluation of students' attention states. In particular, the use of Mahalanobis distance to identify specific expressions and state transition function to model attention fluctuations successfully solves the problems that traditional methods are difficult to quantify emotional changes and predict attention trends. The K-means++ clustering analysis enables teachers to quickly identify different learning status groups and achieve precise teaching intervention. The implementation results show that the teaching adjustment based on this method has increased the average attention state score of the target students by 25.9%, and the learning effect has been significantly improved, which proves the practical value of this method in improving teaching quality and promoting personalized education.
[0263] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 5 and 6 below.
[0264] Table 5 Variable explanation table (the first part)
[0265]
[0266]
[0267] Table 6 Variable explanation table (the second part)
[0268]
[0269] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A method for generating a dynamic portrait of students in an interactive classroom, characterized in that, Including: Collect the facial image sequence of students in class, extract the key point data of the face to construct an expression feature matrix; Calculate the similarity between the expression feature matrix and the standard expression library to generate an expression similarity vector, record the frequency of expression changes, and establish an expression dynamic change curve; collect the sequence of gaze focus coordinates, calculate the gaze concentration area and retention time, and quantify the gaze concentration index; establish an individual association model; identify specific expressions and mark the abnormal points of the learning state; Analyze the correlation between the attention concentration and the classroom content nodes, and construct an attention fluctuation curve; Model the attention fluctuation to generate an attention state score; Based on the learning state evaluation network model, fuse multi-dimensional features to generate a dynamic portrait feature map as the student's dynamic portrait; among them, the core of the learning state evaluation network model is the multi-head attention mechanism, the number of heads is determined by the frequency of expression changes, the dimension of each attention head is determined by the gaze concentration threshold, and the depth of the attention layer is calculated from the attention volatility.
2. The method for generating a dynamic portrait of students in an interactive classroom according to claim 1, wherein The frequency of expression change is the number of significant changes in the student's expression similarity vector per unit time, which is obtained by calculating the cumulative difference between the expression similarity vectors at adjacent time points, and reflects the activity level of the student's emotional state changes.
3. The method for generating the dynamic portrait of students in the interactive classroom according to claim 2, wherein A specific expression is an expression change in which a student shows a significant statistical difference from their personal expression baseline or the class average expression state in class.
4. The method for generating a dynamic portrait of students in an interactive classroom according to claim 3, wherein The gaze concentration index is the proportion of the time that the student's line of sight stays in the specified area in the total observation time, which is obtained by calculating the spatial aggregation degree and time density of the eye movement trajectory points, and reflects the stability of the student's visual attention.
5. The method for generating a dynamic portrait of students in an interactive classroom according to claim 4, wherein, The gaze concentration threshold is the critical value for judging whether the student's line of sight is in a concentrated state, which is determined by analyzing historical data and is used to convert the continuous gaze concentration index into a discrete attention state judgment.
6. The method for generating a dynamic portrait of students in an interactive classroom according to claim 5, wherein, The gaze concentration area is the spatial range where the student's line of sight frequently stays, which is obtained by analyzing the sequence of gaze focus coordinates through a density clustering algorithm, and is characterized as a set of hot area coordinates and weight distribution on a two-dimensional plane.
7. The method for generating a dynamic portrait of students in an interactive classroom according to claim 6, wherein The attention concentration index is a quantitative index of the student's cognitive investment degree calculated by comprehensively considering the expression feature stability and the gaze concentration index. By weighted fusion of multiple physiological behavior feature parameters, it reflects the student's attention degree to the classroom content and the depth of cognitive processing.
8. The method for generating a dynamic portrait of students in an interactive classroom according to claim 7, characterized in that, The individual association model is a mapping function established between the expression feature matrix and the attention concentration index for each student, which is obtained through deep learning algorithms and is used to adjust the weight of the evaluation parameters according to the individual differences of students to improve the accuracy of attention state judgment.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions, which are used to execute the method for generating a dynamic portrait of students in an interactive classroom according to any one of claims 1-8 when running on a computer.
10. An interactive classroom student dynamic portrait generation system, characterized in that, Including the computer-readable storage medium according to claim 9, the system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.
Citation Information
Patent Citations
Eye movement and facial expression normal form-based student learning state evaluation system and method
CN113486744A
Online classroom learning state analysis method
CN115797829A
Classroom cognitive input identification method and system based on multi-modal data
CN117237766A
Artificial intelligence-based system for analyzing student behavior
DE202024107622U1
Cited By
Campus information intelligent pushing method and system based on all-purpose card, and storage medium
CN120873302A
Campus information intelligent pushing method and system based on a card and storage medium
CN120873302B
Personalized online education system and method based on artificial intelligence
CN120950521A
Classroom teaching interaction method and device, electronic equipment and storage medium
CN121301425A
Interactive teaching methods, devices, electronic equipment and storage media
CN121301425B