Student state real-time analysis method and device based on deep learning

By enhancing and standardizing the image data of students in the classroom, combining multi-scale face detection and feature extraction, and applying heavy-tail recursive neural network for timing modeling, the adaptability and personalization problems of student state analysis in traditional methods are solved, and high-precision dynamic analysis and personalized evaluation are achieved.

CN120375481AActive Publication Date: 2025-07-25FUTURE GENE (BEIJING) ARTIFICIAL INTELLIGENCE RES INST CO LTD

Patent Information

Application Number
CN202510873260.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately capture the subtle changes and long-term trends of student status, and has poor adaptability and lacks personalized analysis capabilities. Traditional convolutional neural networks are poor in processing long-term series data and are not adaptable to complex environments.

Method used

By obtaining student image data in the classroom for image enhancement, lighting compensation and standardization processing, multi-scale face detection and facial feature point extraction, combined with convolutional neural network to extract expression features and posture features, construct timing feature sequences and apply re-tail recursive neural network for timing modeling, and establish evaluation standards for adaptive threshold adjustment.

Benefits of technology

It realizes high-precision dynamic analysis of student status, improves identification accuracy and personalized evaluation ability, can accurately capture subtle changes and long-term trends, and adapt to different environments and individual differences among students.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375481A_ABST
    Figure CN120375481A_ABST
Patent Text Reader

Abstract

The invention discloses a student state real-time analysis method and device based on deep learning, and the method comprises the steps: obtaining and preprocessing image data of students in a classroom, and carrying out the multi-scale face detection and facial feature extraction of the preprocessed image data; expression features are extracted based on standardized facial feature data, fusion is carried out in combination with attention indexes and attitude features, a time sequence feature sequence is constructed, a heavy-tailed recurrent neural network is applied to carry out time sequence modeling, an evaluation standard is established, evaluation parameters are adjusted through a self-adaptive threshold value, and finally a student state evaluation result is obtained. Through the heavy-tailed recurrent neural network and a slow transition mechanism to low-dimensional chaos, subtle changes and long-term trends of student states can be accurately captured, and the technical problems that a traditional student state monitoring method is poor in real-time performance, limited in coverage and insufficient in individuation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of educational informatization technology, and particularly to a method and device for real-time analysis of student status based on deep learning, which is applicable to the monitoring and analysis of student status in smart classrooms, online education platforms, and educational management systems. Background Art

[0002] In the process of educational modernization, the real-time monitoring and analysis of student status have become a key link in improving teaching quality. In the traditional classroom teaching process, it is difficult for teachers to simultaneously pay attention to the status of each student. Especially in the large-class teaching environment, this problem is more prominent.

[0003] Currently, the common methods for monitoring student status mainly include the questionnaire survey method and the manual observation method. The questionnaire survey method understands the student status by regularly collecting self-evaluation data of students, but it lacks real-time performance; the manual observation method relies on the teacher's experience judgment, and has disadvantages such as strong subjectivity and limited coverage. In recent years, the student status analysis technology based on computer vision has begun to be applied in the education field, such as using the OpenCV library for simple facial expression recognition or pose detection.

[0004] The existing student status analysis technology based on computer vision mainly adopts the traditional convolutional neural network structure to process and analyze the collected student images. This type of technology identifies features such as the facial expressions and head postures of students through a pre-trained model, and then infers their attention level and emotional state, and presents the recognition results in a visual form to the teacher for reference to adjust the teaching progress and methods.

[0005] However, the existing technology has obvious deficiencies in practical applications: First, the traditional convolutional neural network has poor performance in processing long-term sequence data and is difficult to accurately capture the subtle changes and long-term trends of student status; second, the existing models have poor adaptability to non-standard postures, complex lighting conditions, and partial occlusion situations, and are prone to misjudgment; in addition, the existing systems generally lack the adaptive ability to individual differences of students and cannot perform personalized analysis based on the benchmark status of different students, which affects the recognition accuracy. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for real-time analysis of student status based on deep learning, aiming to solve the technical problems existing in the prior art, such as difficulty in accurately capturing the subtle changes and long-term trends of student status, poor adaptability to complex environments, and lack of personalized analysis ability.

[0007] To achieve the above object, the present invention provides a method for real-time analysis of student status based on deep learning, including: obtaining student image data in the classroom, performing image enhancement, light compensation and normalization processing on the student image data to obtain preprocessed image data; performing multi-scale face detection on the preprocessed image data, extracting facial feature points and performing normalization processing to obtain normalized facial feature data; based on the normalized facial feature data, extracting expression features through a convolutional neural network, combining attention metrics and pose features, and performing feature fusion processing to obtain a status feature vector; constructing a temporal feature sequence for the status feature vector, performing temporal modeling through a heavy-tailed recurrent neural network, analyzing the status change trend, and obtaining a status temporal analysis result; based on the status temporal analysis result, establishing an evaluation criterion, adaptively adjusting the parameters of the evaluation criterion through an adaptive threshold, and performing status evaluation on students through the evaluation criterion after parameter adjustment to obtain a student status evaluation result.

[0008] Preferably, performing image enhancement, light compensation and normalization processing on the student image data includes: based on the student image data, calculating an image sharpness score, screening out valid images with an image sharpness score higher than a preset threshold; performing adaptive histogram equalization processing on the valid images to obtain images with balanced brightness; applying a bilateral filtering algorithm to the images with balanced brightness to remove image noise and obtain denoised images; performing size adjustment and pixel value normalization on the denoised images to obtain the preprocessed image data.

[0009] Preferably, performing multi-scale face detection on the preprocessed image data, extracting facial feature points and performing normalization processing includes: applying a cascaded convolutional neural network to the preprocessed image data for face region detection to obtain face detection box data; applying a non-maximum suppression algorithm to the face detection box data to eliminate overlapping detection boxes and obtain optimized face position data; extracting facial feature points based on the optimized face position data to obtain facial feature point coordinate data; performing pose angle estimation and geometric correction on the facial feature point coordinate data to obtain the normalized facial feature data.

[0010] Preferably, based on the standardized facial feature data, expression features are extracted through a convolutional neural network, combined with attention metrics and pose features, and feature fusion processing is performed, including: applying a convolutional neural network to the standardized facial feature data to extract expression features in the eye, eyebrow, and mouth regions to obtain an expression feature vector; calculating the eye opening degree and blink frequency based on the standardized facial feature data to obtain an attention metric feature vector; applying a pose estimation network to the preprocessed image data to extract upper body key points to obtain a pose feature vector; and performing weighted fusion on the expression feature vector, attention metric feature vector, and pose feature vector by calculating an attention weight coefficient to obtain the state feature vector.

[0011] Preferably, a temporal feature sequence is constructed for the state feature vector, including: arranging the state feature vector in chronological order to obtain original feature temporal data; setting sliding windows of multiple time scales for the original feature temporal data to obtain segmented temporal feature data; applying a heavy-tailed recurrent neural network to the segmented temporal feature data to obtain state change feature data; identifying inflection points of concentration changes based on the state change feature data and predicting state change trends to obtain the state temporal analysis result.

[0012] Preferably, applying a heavy-tailed recurrent neural network to the segmented temporal feature data includes: introducing an activation function of α-stable distribution to the segmented temporal feature data to obtain feature data with heavy-tailed characteristics; applying an attention mechanism to the feature data with heavy-tailed characteristics for long-term dependence modeling to obtain temporally correlated feature data; dynamically adjusting the network parameter entropy value for the temporally correlated feature data to obtain feature data with introduced randomness; calculating a state transition probability matrix based on the feature data with introduced randomness to obtain state transition feature data; and applying a multi-task learning strategy to train and optimize the state transition feature data to obtain the state change feature data.

[0013] Preferably, based on the state temporal analysis result, an evaluation criterion is established, including: extracting concentration mean and standard deviation features from the state temporal analysis result to obtain statistical feature data; applying a clustering algorithm to the statistical feature data for state pattern division to obtain state distribution data; calculating a normal concentration interval and a fatigue threshold based on the state distribution data to obtain state evaluation criterion parameters; performing multivariate regression analysis on the state evaluation criterion parameters combined with environmental parameters to obtain environmental impact coefficient data; and correcting the state judgment criterion based on the environmental impact coefficient data to obtain the evaluation criterion.

[0014] Preferably, the parameters of the evaluation criteria are adjusted through an adaptive threshold, and the state of the student is evaluated by the evaluation criteria after parameter adjustment to obtain the student state evaluation result, including: applying a decision tree model to the environmental impact coefficient data to quantify the impact weights of various environmental factors to obtain environmental weight data; calculating a state deviation threshold under the current environmental conditions based on the environmental weight data to obtain adaptive threshold data; dynamically adjusting the evaluation criteria by combining the adaptive threshold data with historical state evaluation data; evaluating the real-time state of the student based on the evaluation criteria after adjusting the parameters, calculating the degree of state deviation to obtain state evaluation data; generating a change trend graph of concentration and emotional state for the state evaluation data to obtain the student state evaluation result.

[0015] Preferably, the method further includes: generating class state visualization data according to the student state evaluation result, specifically including: performing statistical clustering processing on the student state evaluation result to obtain class overall state distribution data; generating a concentration heat map based on the class overall state distribution data to obtain spatial distribution visualization data.

[0016] The present invention also provides a real-time analysis device for student state based on deep learning, including: an image acquisition module, configured to acquire student image data in the classroom, perform image enhancement, light compensation and standardization processing on the student image data to obtain preprocessed image data; a feature extraction module, configured to perform multi-scale face detection on the preprocessed image data, extract facial feature points and perform standardization processing to obtain standardized facial feature data; a state recognition module, configured to extract expression features based on the standardized facial feature data through a convolutional neural network, combine attention indicators and pose features, and perform feature fusion processing to obtain a state feature vector; a time series analysis module, configured to construct a time series feature sequence for the state feature vector, perform time series modeling through a heavy-tailed recurrent neural network, analyze the state change trend to obtain a state time series analysis result; an evaluation module, configured to establish an evaluation criteria based on the state time series analysis result, adjust the parameters of the evaluation criteria through an adaptive threshold, and evaluate the state of the student by the evaluation criteria after adjusting the parameters to obtain the student state evaluation result.

[0017] The beneficial effects of the present invention are: 1. By "acquiring student image data in the classroom and performing image enhancement, light compensation and standardization processing", high-quality preprocessing of images under different lighting conditions and environments is realized, and the quality of the basic data for subsequent analysis is improved; 2. By "multi-scale face detection and standardization processing of facial feature points", the adaptability problem of traditional face detection in a complex classroom environment is solved, and the faces of students at different distances and angles can be accurately located, providing accurate regional positioning for feature extraction; 3. Through the "fusion processing of extracting facial expression features by convolutional neural network and combining attention metrics and pose features", the effective integration of multi-modal features is achieved, overcoming the limitations of single features and improving the robustness and accuracy of student state recognition; 4. By "constructing a time series feature sequence and applying a heavy-tailed recurrent neural network for time series modeling", the subtle changes and long-term trends of the student state can be accurately captured, realizing high-precision dynamic analysis of the student state changes and effectively enhancing the prediction ability for abnormal states; 5. By "establishing an evaluation criterion and using an adaptive threshold to adjust parameters", the problem that traditional fixed criteria cannot adapt to individual differences among students is solved, realizing personalized state evaluation for different students and different environments, and greatly improving the accuracy and practicality of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a real-time analysis method for student state based on deep learning provided by an embodiment of the present invention; Figure 2 It is a flowchart of time series data modeling and state evolution analysis provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a real-time analysis device for student state based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0021] As Figure 1 shown, an embodiment of the present invention provides a real-time analysis method for student state based on deep learning, including the following steps: S11. Obtain the student image data in the classroom, perform image enhancement, light compensation, and normalization processing on the student image data to obtain preprocessed image data.

[0022] In this embodiment, this step is achieved by means of a high-definition camera array installed at different positions in the classroom, which can collect the original image data stream containing multiple angles of students, such as the front and side views. At the same time, environmental parameters, such as light intensity and classroom layout, are recorded to provide reference for subsequent processing. The acquired original image data often has problems such as uneven illumination and noise interference. Therefore, image enhancement, illumination compensation, and normalization processing are required. Specifically, first, an image sharpness evaluation algorithm (such as the Laplacian operator and gradient magnitude analysis) is applied to score the quality of each frame of the image, and high-quality valid images are screened out; then, these valid images are applied with the Contrast Limited Adaptive Histogram Equalization (CLAHE) and multi-scale Retinex algorithms for illumination compensation and enhancement to balance the image brightness and contrast; next, bilateral filtering and non-local means filtering algorithms are used to remove image noise while retaining edge details; finally, size adjustment, cropping, and pixel value normalization processing are performed, and data augmentation techniques are used to improve the generalization ability of the model, and finally high-quality preprocessed image data is obtained.

[0023] S12. Perform multi-scale face detection on the preprocessed image data, extract facial feature points and perform normalization processing to obtain normalized facial feature data.

[0024] In this embodiment, performing multi-scale face detection on the preprocessed image data is a key step in realizing the analysis of students' states. An improved MTCNN (Multi-task Cascaded Convolutional Networks) or RetinaFace detection algorithm is used to construct a feature pyramid network to detect the face regions in the image at different scales. This multi-scale detection strategy can effectively process faces at different distances and angles and improve the robustness of detection. The detected face regions usually have overlaps and redundancies. Therefore, the Non-Maximum Suppression (NMS) algorithm is applied to eliminate the overlapping detection boxes, and the false detection results are filtered out by combining the temporal continuity constraint (through inter-frame tracking) to ensure the accurate positioning of the face positions of each student. Subsequently, an improved facial key point localization model (such as the FAN - Face Alignment Network or the shape predictor of the DLib library) is used to extract 68 standard feature points of each face, including the key positions of the eyes, eyebrows, nose, mouth, and facial contour. To process non-frontal faces, the Perspective-n-Point (PnP) pose estimation algorithm is applied based on the feature points to calculate the three-dimensional pose angles of the head and generate an affine transformation matrix for geometric correction. Finally, through eye center point alignment and facial proportion normalization processing, individual differences and the influence of the camera angle are eliminated to obtain normalized facial feature data, laying a foundation for subsequent analysis.

[0025] S13. Based on the standardized facial feature data, extract expression features through a convolutional neural network, combine attention metrics and pose features, and perform feature fusion processing to obtain a state feature vector.

[0026] In this embodiment, based on the standardized facial feature data, the present invention extracts and fuses multi-modal features through a deep learning method to comprehensively analyze the state of students. First, apply a deep convolutional neural network with an improved VGG-Face or ResNet architecture to extract deep expression features in the facial region, paying particular attention to key regions for emotion expression such as the eyes, eyebrows, and mouth. Subtle changes in these regions often reflect the emotions and cognitive states of students. At the same time, calculate a series of attention-related metrics based on facial feature points, including eye aspect ratio (EAR), pupil tracking, blink frequency, and head pose stability. These metrics are important bases for evaluating students' attentiveness. In addition, a human pose estimation network (such as OpenPose or HRNet) can be applied to extract the key points of the upper body of students and analyze pose features such as sitting posture, body tilt angle, and movement frequency. These features can reflect students' engagement and state from the perspective of body language. Finally, integrate the three modal features of expression features, attention metrics, and body pose through a feature fusion network with an attention mechanism, dynamically allocate weights according to data quality and environmental conditions, and apply a multi-task learning framework to simultaneously perform attentiveness level scoring and emotion category judgment, and finally output a state feature vector containing attentiveness score, emotion category, and their confidence levels.

[0027] S14. Construct a time series feature sequence for the state feature vector, perform time series modeling through a heavy-tailed recurrent neural network, analyze the state change trend, and obtain a state time series analysis result.

[0028] In this embodiment, constructing a time-series feature sequence for the state feature vector is the core step in analyzing the dynamic changes of students' states. The present invention first arranges the state feature vectors of each student in chronological order to construct the original feature time-series data. To capture the change patterns at different time scales simultaneously, sliding windows of multiple time scales (such as 30 seconds, 1 minute, 5 minutes, etc.) are set to segment the sequence, obtaining multi-scale time-series feature data. Considering the complexity and uncertainty of students' states, a recurrent neural network structure with heavy-tailed distribution characteristics is designed. By introducing the activation function of α-stable distribution, the network's ability to model abnormal states is enhanced. This network realizes a slow transition mechanism to low-dimensional chaos. By dynamically adjusting the entropy value of the network parameters and introducing controlled randomness, the model can capture subtle state changes while maintaining stability. The model is trained using multi-task learning strategies and regularization techniques (such as Dropout and weight decay), and the hyperparameters are tuned through the validation set to obtain an optimized time-series model of students' states. Based on this model, trend analysis is performed on the state sequences of each student to identify the inflection points of concentration and the patterns of emotional changes, and to predict the state change trends in the short term (such as the next 5 - 10 minutes). Finally, the time-series analysis results of students' states are output.

[0029] S15. Based on the state time-series analysis results, establish an evaluation criterion, adjust the parameters of the evaluation criterion through an adaptive threshold, and evaluate the students' states using the adjusted evaluation criterion to obtain the students' state evaluation results.

[0030] In this embodiment, based on the state time-series analysis results, a student evaluation criterion is established. Specifically, first, combined with historical cumulative data, the state statistical features of each student (such as the mean of concentration, standard deviation, typical emotional distribution, etc.) are extracted to construct a basic state portrait of the student. Then, an adaptive clustering algorithm and a probability distribution fitting method are applied to calculate personalized state benchmark intervals for each student, including parameters such as the normal concentration interval, fatigue threshold, and emotional fluctuation tolerance. At the same time, multivariate regression analysis and decision tree models are used to quantify the influence degree of environmental factors (such as light, course type, time period, etc.) on students' states, generating an environmental influence coefficient matrix. Based on these data, an adaptive threshold adjustment algorithm is designed to dynamically adjust the judgment criterion according to the current environmental conditions and the students' real-time states. Specifically, first, the decision tree model is used to quantify the influence weights of each environmental factor and calculate the state deviation threshold under the current environmental conditions; then, the evaluation criterion is dynamically adjusted in combination with historical state evaluation data; then, based on the adjusted evaluation criterion, the students' real-time states are evaluated to calculate the degree of state deviation; finally, a comprehensive evaluation report including concentration, emotional state and their change trends is generated, and the students' state evaluation results are output.

[0031] Optionally, the real-time analysis method for student status based on deep learning further includes: S16. Generate visual data of class status according to the student status evaluation result.

[0032] In this embodiment, according to the student status evaluation result, the present invention generates intuitive visual data of class status to provide decision-making support for teachers. First, through weighted average and statistical clustering methods, calculate the overall class concentration distribution, emotion hotspots and status abnormal areas to obtain class status statistical data. Then, apply data visualization techniques (such as heat maps, status curves, emotion radar charts, etc.) to design an intuitive teacher interface to realize multi-level display of class overview and student individual details.

[0033] In addition, status abnormal detection rules can be set to monitor and identify key events that require teachers' attention in real time, such as the status where the concentration is continuously lower than the preset threshold and the status where the emotion fluctuation amplitude exceeds the set range, and generate reminder information through priority sorting.

[0034] In an embodiment of the present invention, the image enhancement, light compensation and normalization processing of the student image data include: Based on the student image data, calculate the image sharpness score, and screen out the effective images with the image sharpness score higher than the preset threshold; perform adaptive histogram equalization processing on the effective images to obtain images with balanced brightness; apply a bilateral filtering algorithm to the images with balanced brightness to remove image noise and obtain denoised images; perform size adjustment and pixel value normalization on the denoised images to obtain the preprocessed image data.

[0035] Specifically, first, collect the original image data streams including multiple angles such as the front and side of students through a high-definition camera array installed at different positions in the classroom, and record the environmental parameters (such as light intensity, classroom layout, etc.) at the same time, and output a multi-view original image data set and an environmental parameter set.

[0036] Then, based on the multi-view original image data set, apply an image sharpness evaluation algorithm (such as Laplace operator, gradient magnitude analysis) to score the quality of each frame of image, screen out the effective images that meet the requirements of subsequent processing, and output the effective image set after quality scoring.

[0037] Next, based on the effective image set, combine the environmental parameter set to apply adaptive histogram equalization (CLAHE) and multi-scale Retinex algorithms to enhance and compensate the images under different lighting conditions, balance the image brightness and contrast, and output the enhanced image set after light compensation.

[0038] Subsequently, based on the enhanced image set, bilateral filtering and non-local means (NLM) filtering algorithms are applied to remove image noise while preserving edge details, and possible occluded or blurred areas are repaired, and a set of repaired images after denoising is output.

[0039] Finally, based on the set of repaired images, size adjustment, cropping, and pixel value normalization (normalized to the range [0, 1]) are performed, and the dataset is augmented through operations such as rotation, scaling, and horizontal flipping to improve the generalization ability of the model, and finally a high-quality preprocessed image dataset is output.

[0040] In an embodiment of the present invention, multi-scale face detection is performed on the preprocessed image data, and facial feature points are extracted and normalized, including: applying a cascaded convolutional neural network to the preprocessed image data for face region detection to obtain face detection box data; applying a non-maximum suppression algorithm to the face detection box data to eliminate overlapping detection boxes to obtain optimized face position data; extracting facial feature points based on the optimized face position data to obtain facial feature point coordinate data; performing pose angle estimation and geometric correction on the facial feature point coordinate data to obtain the normalized facial feature data.

[0041] Specifically, first, based on the preprocessed image dataset, an improved MTCNN (Multi-task Cascaded Convolutional Networks) or RetinaFace detection algorithm is applied to construct a feature pyramid network to detect face regions in the image at different scales, and a set of face detection boxes containing position coordinates and confidence levels is output.

[0042] Then, based on the set of face detection boxes, a non-maximum suppression (NMS) algorithm is applied to eliminate overlapping detection boxes, and at the same time, false detection results are filtered out by combining temporal continuity constraints (through inter-frame tracking), and an accurate set of student face positions is output.

[0043] Next, based on the set of student face positions, an improved facial key point localization model (such as the FAN - Face Alignment Network or the shape predictor of the DLib library) is applied to extract 68 standard feature points for each face, including the key positions of eyes, eyebrows, nose, mouth, and facial contour, and an original facial feature point coordinate set is output.

[0044] Subsequently, based on the original facial feature point coordinate set, a PnP (Perspective-n-Point)-based pose estimation algorithm is applied to calculate the pitch angle, yaw angle, and roll angle of the head, and an affine transformation matrix is generated based on the estimated pose angles to perform geometric correction on non-frontal faces, and a set of facial images after pose correction is output.

[0045] Finally, based on the set of face images after pose correction, re-extract and finely adjust the positions of facial feature points. Through the alignment of the center points of the eyes and the normalization of facial proportions, eliminate the influence of individual differences and camera angles, and finally output a standardized facial feature point data set.

[0046] In an embodiment of the present invention, based on the standardized facial feature data, extract expression features through a convolutional neural network, and combine attention metrics and pose features, including: applying a convolutional neural network to the standardized facial feature data to extract expression features in the eye, eyebrow, and mouth regions to obtain an expression feature vector; calculating the eye opening degree and blink frequency based on the standardized facial feature data to obtain an attention metric feature vector; applying a pose estimation network to the preprocessed image data to extract upper body key points to obtain a pose feature vector; and performing weighted fusion on the expression feature vector, attention metric feature vector, and pose feature vector by calculating an attention weight coefficient to obtain the state feature vector.

[0047] Specifically, first, based on the standardized facial feature point data set, apply a deep convolutional neural network with an improved VGG-Face or ResNet architecture to extract deep expression features in the facial region, focus on key regions for emotional expression such as eyes, eyebrows, and mouth, and output a high-dimensional expression feature vector.

[0048] Then, based on the standardized facial feature point data set, calculate attention-related metrics such as eye opening degree (EAR), pupil tracking, blink frequency, and head pose stability, and normalize these metrics to output an attention metric feature vector.

[0049] Next, based on the preprocessed image data set, apply a human pose estimation network (such as OpenPose or HRNet) to extract the upper body key points of the student, analyze pose features such as sitting posture, body tilt angle, and movement frequency, and output a body pose feature vector.

[0050] Finally, based on the expression feature vector, attention metric feature vector, and body pose feature vector, integrate these three features through a feature fusion network with an attention mechanism, assign dynamic weights to different features, output a comprehensive feature representation vector, and apply a multi-task learning framework to simultaneously perform a focus level score (1-5 points) and emotion category (focused, confused, tired, anxious, etc.) judgment, and calculate the confidence of each result, and finally output a state feature vector including the focus score, emotion category, and their confidence.

[0051] Specifically, multi-modal feature fusion is a key link in identifying students' states. For facial expression feature extraction, an improved ResNet-50 architecture is used. This network is pre-trained on a large-scale facial expression dataset and fine-tuned for the task of student state recognition. The network input is an RGB image of 224×224×3. After passing through 5 residual blocks and global average pooling, a 2048-dimensional feature vector is output. To enhance the expressive power of the features, a spatial attention module is added to the network, focusing on key regions for emotion expression such as eyes, eyebrows, and mouth. The specific implementation is to calculate the spatial attention weight matrix of the feature map and weight the original feature map to highlight the feature expression in important regions. The data used for network fine-tuning includes 10,000 student facial images labeled with 7 basic emotions (neutral, focused, confused, tired, bored, excited, anxious). The cross-entropy loss function is adopted, with a learning rate of 0.0001, a batch size of 32, and trained for 50 epochs until convergence.

[0052] In terms of attention metric feature extraction, first, the eye aspect ratio (EAR) is obtained by calculating the ratio of the distance between feature points in the vertical and horizontal directions of the eyes. Among them, the distance between feature points in the vertical direction is represented by the distance between the midpoints of the upper and lower eyelids, and the distance between feature points in the horizontal direction is represented by the distance between the corners of the eyes. This ratio can effectively reflect the opening and closing state of the eyes. Second, the blink frequency is calculated by the number of times the EAR value is lower than the threshold (usually 0.2) within a sliding time window (3 seconds). In addition, the present invention also calculates the standard deviations of the pitch angle, yaw angle, and roll angle of the head through a head pose estimation algorithm as the pose stability metric; calculates the length of the pupil movement trajectory and the degree of gaze dispersion through a pupil position tracking algorithm as the gaze concentration metric. After normalization, these metrics form a 12-dimensional attention metric feature vector.

[0053] For body pose feature extraction, the HRNet network is used, which can accurately locate 17 key points of the human body. Special attention is paid to the upper body pose, including the positions and movement features of the shoulders, arms, and upper torso. By calculating the relative positions and movement features of the key points, features such as sitting posture angle, body tilt, movement frequency, and amplitude are extracted. Through time series analysis of these features, a 20-dimensional pose feature vector reflecting students' engagement and focus states is obtained.

[0054] Multi-modal feature fusion adopts a dynamic weight allocation method based on the attention mechanism. First, the three types of modal features (2048-dimensional expression features, 12-dimensional attention index features, and 20-dimensional pose features) are mapped to the same latent space dimension (128-dimensional) through a fully connected layer. Then, a multi-head attention mechanism is designed to calculate the correlation and importance weights of the three features in different representation spaces. Specifically, each type of modal feature passes through 3 attention heads respectively. Each attention head independently learns the representation of this modality in a specific semantic space, and then calculates the fusion weights of the three modalities in each semantic space through a soft attention mechanism. This design can dynamically adjust the importance of different modalities according to real-time data quality, environmental conditions, and task requirements, improving the robustness and self-adaptability of feature fusion.

[0055] The fused features are processed by a multi-task learning head, and at the same time, the concentration score (a continuous value from 1 to 5) and the emotion category (the probability distribution of 7 basic emotions) are output. The mean squared error loss function is used for the concentration score task, and the cross-entropy loss function is used for the emotion classification task. The losses of the two tasks are weighted and combined in a ratio of 3:2. In actual tests, the recognition accuracy of this fusion method is increased by 18.5% on average compared with single-modal features. Especially in complex environments such as insufficient lighting or partial occlusion, the performance improvement is more significant, reaching more than 27%.

[0056] As Figure 2 shown, in an embodiment of the present invention, a temporal feature sequence is constructed for the state feature vector, including: S21. Arrange the state feature vector in chronological order to obtain the original feature temporal data.

[0057] S22. Set sliding windows of multiple time scales for the original feature temporal data to obtain segmented temporal feature data.

[0058] S23. Process the segmented temporal feature data by applying a heavy-tailed recurrent neural network to obtain state change feature data.

[0059] S24. Identify the inflection points of concentration changes based on the state change feature data, predict the state change trend, and obtain the state temporal analysis result.

[0060] In this embodiment, temporal data modeling and state evolution analysis are key steps in understanding the dynamic changes in students' states. First, the state feature vectors are arranged in chronological order to construct the original feature temporal data. This step arranges the state feature vectors (including information such as attention scores, emotion categories, and confidence levels) obtained by each student at consecutive time points in the order of time stamps, forming a time series data structure. This temporal arrangement preserves the time continuity of students' state changes and provides a basic data framework for subsequent analysis. The original feature temporal data not only records the state features of students at each time point but also implies information about the rate and pattern of state changes, which is crucial for understanding students' learning engagement and cognitive processes.

[0061] To comprehensively capture the state change characteristics at different time scales, multiple time-scale sliding windows are set for the original feature temporal data to obtain segmented temporal feature data. Specifically, short-time windows (such as 30 seconds), medium-time windows (such as 1 minute), and long-time windows (such as 5 minutes) are simultaneously applied to segment the original temporal data. The short-time window can capture the instantaneous changes and minute fluctuations in students' states and is suitable for identifying sudden attention shifts or emotional fluctuations; the medium-time window can smooth out short-term fluctuations and reflect relatively stable state trends, making it suitable for evaluating students' responses to specific teaching content; the long-time window can show long-term state change trends and is suitable for analyzing the learning effects of an entire class or teaching unit. Through multi-scale window segmentation, the state change characteristics at different time granularities can be obtained simultaneously, providing rich input data for subsequent deep temporal modeling. The data within each window not only retains the detailed information of the original features but also achieves a smooth transition in time through the sliding mechanism, avoiding information loss that may be caused by hard segmentation.

[0062] The segmented temporal feature data is processed using an innovatively designed heavy-tailed recurrent neural network to obtain state change feature data. Traditional recurrent neural networks often perform poorly when dealing with temporal data such as student states that have uncertainty and mutation characteristics. Therefore, a recurrent neural network structure with heavy-tailed distribution characteristics is designed. The network first introduces the activation function of the α-stable distribution, enabling the network to have the ability to handle outliers and mutations. Different from the Gaussian distribution, the α-stable distribution has a "heavy-tailed" characteristic, which can better model the common non-normal distribution phenomena in the real world, especially situations such as sudden shifts in student attention or sudden changes in emotions. Secondly, an attention mechanism is incorporated into the network structure to enhance the modeling ability of long-term dependence relationships, enabling it to "remember" important state information at earlier times and utilize it in the current state prediction. In addition, a slow transition mechanism to low-dimensional chaos is also implemented. By dynamically adjusting the entropy value of the network parameters and introducing controlled randomness, the model can sensitively capture subtle changes in the state while maintaining overall stability. This design enables the network to neither overfit noise (such as meaningless actions of students) nor ignore important state change signals (such as micro-expressions when having difficulty understanding). Through the multi-task learning strategy, the network simultaneously predicts multiple targets such as changes in concentration, emotion transitions, and state durations, further improving the generalization ability and prediction accuracy of the model.

[0063] Based on the state change feature data obtained after processing by the heavy-tailed recurrent neural network, the inflection points of concentration changes are further identified, the state change trend is predicted, and finally the state temporal analysis result is obtained. The inflection point of concentration change refers to the time point when the student's attention level undergoes a significant change, usually related to changes in the difficulty of teaching content, transitions in teaching activities, or changes in the student's cognitive load. By analyzing the derivative information (rate of change) and second derivative information (acceleration of change) in the state change feature data, combined with statistical significance test methods, these key inflection points are accurately identified. At the same time, time series prediction techniques are also applied to predict the change trend of the student's state in the short term (such as 5 - 10 minutes) in the future based on historical state change patterns. This prediction adopts a hybrid prediction strategy, combining deterministic components (such as regular changes brought about by the course progress) and stochastic components (such as uncertainties brought about by individual differences), to improve the accuracy and reliability of the prediction. Finally, a comprehensive analysis result including the state change trajectory, key inflection point markers, state stable intervals, and future trend predictions is generated.

[0064] Through this complete process of time series data modeling and analysis, discrete student state observation data can be transformed into continuous state change curves. Compared with traditional methods, the time series analysis method of the present invention has three significant advantages: First, the multi-scale time window design can capture state change characteristics at different time granularities simultaneously, providing a more comprehensive view of the learning process; Second, the innovative design of the heavy-tailed recurrent neural network greatly improves the sensitivity to abnormal states and subtle changes, making the prediction results more accurate and reliable; Finally, the focus inflection point identification and trend prediction functions provide forward-looking decision-making support for teachers, enabling teaching interventions to be more timely and effective.

[0065] Specifically, first, based on the state feature vector, a feature sequence is constructed for each student in chronological order, and a sliding window (multi-scales such as 30 seconds, 1 minute, 5 minutes, etc.) is set to segment the sequence, and a multi-scale time series feature sequence set is output.

[0066] In an embodiment of the present invention, the segmented time series feature data is processed by applying a heavy-tailed recurrent neural network, including: introducing an activation function of the α-stable distribution to the segmented time series feature data to obtain feature data with heavy-tailed characteristics; applying an attention mechanism to the feature data with heavy-tailed characteristics for long-term dependence modeling to obtain time series correlation feature data; dynamically adjusting the entropy value of the network parameters for the time series correlation feature data to obtain feature data with introduced randomness; calculating a state transition probability matrix based on the feature data with introduced randomness to obtain state transition feature data; applying a multi-task learning strategy to train and optimize the state transition feature data to obtain the state change feature data.

[0067] Specifically, first, based on statistical learning theory and dynamic system theory, a recurrent neural network structure with heavy-tailed distribution characteristics is designed. By introducing an activation function of the α-stable distribution and an attention mechanism, the network's modeling ability for abnormal states and long-term dependence is enhanced, and a heavy-tailed RNN model structure is output.

[0068] Then, based on the heavy-tailed RNN model structure, a slow transition mechanism to low-dimensional chaos is realized. By dynamically adjusting the entropy value of the network parameters and introducing controlled randomness, the model can capture subtle state changes while maintaining stability, and an enhanced RNN model with low-dimensional chaos characteristics is output.

[0069] Next, based on the multi-scale time series feature sequence set and the enhanced RNN model, a multi-task learning strategy and regularization techniques (such as Dropout and weight decay) are applied for model training, and the hyperparameters are tuned through a validation set, and an optimized student state time series model is output.

[0070] Finally, based on the optimized student state time series model, perform trend analysis on the student state sequence of each student, identify the focus inflection points and emotional change patterns, predict the state change trend in the short term (such as the next 5 - 10 minutes), and output the time series analysis results of the student state.

[0071] Specifically, the designed heavy-tailed recurrent neural network is a deep learning model customized for the special needs of student state analysis. The heavy-tailed distribution refers to a probability distribution whose tail decays more slowly than the exponential distribution, and this property makes the model more sensitive to abnormal events or extreme values. In this embodiment, the α-stable distribution is used as the basis of the activation function, where the α parameter controls the tail thickness of the distribution, and its value range is (0, 2]. When α = 2, the distribution degenerates into a Gaussian distribution; when α approaches 0, the tail of the distribution becomes thicker. According to the characteristics of the student state data, the α value is selected between 1.5 and 1.8, and this range can improve the ability to capture abnormal states while maintaining the stability of the model.

[0072] The basic structure of the heavy-tailed recurrent neural network includes an input layer, a hidden layer, and an output layer. The input layer receives the state feature vector with a dimension of 32; the hidden layer consists of three layers of bidirectional LSTM units, each layer containing 128 neurons, and an activation layer with α-stable distribution characteristics is added after each layer of LSTM; the output layer is set with the corresponding dimension according to the task requirements. The model is trained using the Adam optimizer, with the initial learning rate set to 0.001, and a learning rate decay strategy is used, decaying to 0.9 times the original value every 50 epochs. To prevent overfitting, the Dropout technique is applied to the model, with a dropout probability of 0.3 set after each layer of LSTM, and an L2 regularization coefficient of 0.0001 is used. The training data comes from the real student state data collected in the classroom, including 5000 samples, 80% of which are used for training and 20% for validation. During the training process, a multi-task loss function is used to optimize both the focus prediction and the emotion classification tasks simultaneously, and the loss weight ratio of the two tasks is 3:2.

[0073] The slow transition mechanism to low-dimensional chaos is an innovation point of the present invention, which is achieved by dynamically adjusting the entropy value of the network parameters. Specifically, in the initial stage of model training, the dimension of the network parameter space is relatively high, and the system behavior is relatively regular; as the training progresses, the dimension of the parameter space is gradually reduced, and at the same time, controlled randomness is introduced, presenting low-dimensional chaos characteristics. The implementation method is to calculate the entropy value of the network parameters after each training epoch and adjust it according to the preset entropy decay curve. When a small change in the student state sequence is detected, the randomness is appropriately increased to enhance the sensitivity to this change; when the state is relatively stable, the randomness is reduced to maintain the prediction stability. This mechanism enables the model to balance stability and sensitivity and adapt to the dynamic change characteristics of the student state.

[0074] In practical applications, the heavy-tailed recurrent neural network has improved the processing accuracy of student status data by about 24% compared with traditional RNNs. Especially in detecting abrupt changes in concentration, the early warning time is advanced by an average of 5.7 minutes, providing a sufficient reaction window for teacher intervention. The model shows good generalization ability in tests under different environmental conditions and different student groups, and its robustness to interference factors (such as light changes and partial occlusion) is also significantly better than the benchmark model.

[0075] In one embodiment of the present invention, based on the state time series analysis results, an evaluation criterion is established, including: extracting the mean and standard deviation features of concentration from the state time series analysis results to obtain statistical feature data; applying a clustering algorithm to the statistical feature data for state pattern division to obtain state distribution data; calculating the normal concentration interval and fatigue threshold based on the state distribution data to obtain state evaluation criterion parameters; performing multivariate regression analysis on the state evaluation criterion parameters combined with environmental parameters to obtain environmental impact coefficient data; and correcting the state judgment criterion based on the environmental impact coefficient data to obtain the evaluation criterion.

[0076] Specifically, first, based on the student state time series analysis results and combined with historical cumulative data, the state statistical features of each student (such as mean concentration, standard deviation, typical emotion distribution, etc.) are extracted to construct a student basic state portrait and output a student individual feature model.

[0077] Then, based on the student individual feature model, an adaptive clustering algorithm and a probability distribution fitting method are applied to calculate personalized state benchmark intervals for each student, including parameters such as normal concentration interval, fatigue threshold, and emotion fluctuation tolerance, and output a personalized state benchmark model.

[0078] Next, based on the environmental parameter set and the student state time series analysis results, multivariate regression analysis and a decision tree model are applied to quantify the influence degree of environmental factors (such as light, course type, time period, etc.) on the student state and output an environmental impact coefficient matrix.

[0079] Finally, based on the personalized state benchmark model and the environmental impact coefficient matrix, an adaptive threshold adjustment algorithm is designed to dynamically adjust the judgment criterion according to the current environmental conditions and the real-time state of the student, improve the accuracy of state judgment, and output an adaptive threshold system.

[0080] Specifically, first, a basic state portrait is established for each student. By observing for a long time (data of at least 5 class hours), the mean, standard deviation, fluctuation period, and emotional distribution characteristics of the student's concentration are calculated. These statistical characteristics are modeled using the Gaussian Mixture Model (GMM), usually with 3 - 5 Gaussian components. The model parameters are estimated through the EM algorithm to obtain a probability model that can describe the state distribution characteristics of this student. After the model training is completed, key parameters such as an individualized normal concentration interval (usually mean ± 1.5 times the standard deviation), fatigue threshold (usually mean - 2 times the standard deviation), and emotional fluctuation tolerance (95% confidence interval based on the historical emotional change range) are automatically determined for each student.

[0081] The impact of environmental factors on students' states cannot be ignored. These impacts are quantified through multivariate regression analysis and decision tree models. The recorded environmental factors include light intensity (lux value), environmental noise level (dB value), classroom temperature (°C), humidity (%), course type (liberal arts and science classification), class time period (morning, afternoon, evening), and course duration, etc. The Gradient Boosting Decision Tree (GBDT) model is used to analyze the relationship between these factors and the deviation of students' states, and an environmental impact coefficient matrix is constructed. This matrix records the impact weights of each environmental factor on concentration and emotional state within different value ranges. For example, it may be found that in a high - temperature environment above 30°C, the average concentration of students drops by 0.7 points, and the emotional fluctuation increases by 25%. These quantification results provide a basis for environmental compensation in subsequent state evaluations.

[0082] The adaptive threshold adjustment algorithm can dynamically adjust the judgment criteria according to the current environmental conditions and the student's historical performance. First, the algorithm calculates the expected state deviation under the current environmental conditions based on the environmental impact coefficient matrix, and then makes personalized adjustments in combination with the student's individual environmental sensitivity coefficient (obtained through learning historical data). For example, for students who are particularly sensitive to noise, the concentration judgment threshold is correspondingly reduced in a noisy environment. In addition, the algorithm also considers time factors, such as the fatigue accumulation effect and the natural fluctuation period of attention, and dynamically adjusts the threshold through a time - weight function. Experiments show that compared with a fixed threshold, the adaptive threshold increases the accuracy of state evaluation by 21.3%, especially in scenarios where environmental conditions change significantly.

[0083] During the student status assessment process, the deviation degree between the student's current status and the personal benchmark is calculated, and its statistical significance is judged. The assessment results include not only the concentration score and emotion category, but also information such as the status change trend, abnormal duration, and severity. The generated assessment report is presented in a visual form, including the status time series curve, current status marker, warning information, and recommended measures, etc., providing intuitive and comprehensive student status information for teachers. During the long-term use process, the basic status portrait of the student and the environmental sensitivity model will be continuously updated to achieve the continuous optimization and self-evolution of the assessment criteria and adapt to the long-term change law of the student status.

[0084] In one embodiment of the present invention, the parameters of the assessment criteria are adjusted through an adaptive threshold, and the student status assessment is performed on the student by using the assessment criteria after the parameters are adjusted, and the student status assessment result is obtained, including: applying a decision tree model to the environmental impact coefficient data to quantify the influence weights of various environmental factors to obtain environmental weight data; calculating a status deviation threshold under the current environmental conditions based on the environmental weight data to obtain adaptive threshold data; dynamically adjusting the assessment criteria by combining the adaptive threshold data with the historical status assessment data; performing an assessment on the real-time status of the student based on the assessment criteria after the parameters are adjusted, calculating the status deviation degree to obtain status assessment data; generating a change trend chart of the concentration and emotion status for the status assessment data to obtain the student status assessment result.

[0085] In one embodiment of the present invention, class status visualization data is generated according to the student status assessment result, including: performing statistical clustering processing on the student status assessment result to obtain class overall status distribution data; generating a concentration heat map based on the class overall status distribution data to obtain spatial distribution visualization data.

[0086] Optionally, generating class status visualization data according to the student status assessment result further includes: performing anomaly detection on the student status change trend chart to identify the status where the concentration is continuously lower than the preset threshold; and the status where the emotion fluctuation amplitude exceeds the set range to obtain key event reminder data.

[0087] Specifically, the generation of class status visualization data adopts a multi-level and multi-dimensional method. First, aggregate analysis is performed on the student status assessment results of the whole class of students to calculate the concentration distribution characteristics of the class as a whole. The students are divided into different status groups such as high concentration, medium concentration, and low concentration by using the K-means clustering algorithm (usually K = 3 or K = 4), and the proportion and spatial distribution characteristics of each group are calculated. The emotional computing model can also be further applied to analyze the overall emotional atmosphere of the class and identify the dominant emotion type and emotion synchronization index.

[0088] The focus heat map is one of the core visualization tools, which correlates the physical locations of students in the classroom with their focus states, intuitively showing the spatial distribution law of focus. The heat map adopts a gradient color scale system, usually using a mapping relationship from cold colors (low focus) to warm colors (high focus), and the color saturation represents the focus intensity. The time dimension can also be overlaid on the heat map, and the evolution process of the class state over time can be shown through the dynamic playback function, helping teachers identify the state contagion effect and key turning points. The generation of the heat map uses the bilinear interpolation algorithm to create a smooth transition effect between discrete student location points, improving the aesthetics and readability of the visualization.

[0089] Key event detection can automatically identify abnormal states that require teachers' attention and intervention. This application sets multiple abnormal state detection rules, including: a continuous low focus state where the focus is continuously lower than the individual benchmark threshold for more than a preset duration (usually 5 minutes); a sharp decline state where the focus drops by more than a preset threshold (usually 1.5 points) within a short period (usually 2 minutes); an emotional fluctuation state where the emotional state changes significantly within a short period (usually 3 minutes); and a group effect state where multiple students show abnormal states simultaneously. These rules are combined with statistical significance tests to ensure that the detected abnormalities have practical teaching significance rather than random fluctuations. The detected abnormal events are sorted according to importance and urgency to generate a priority list, and visual cues (such as flashing marks, color coding) are used to draw teachers' attention.

[0090] As Figure 3 shown, the present invention also provides a real-time student state analysis device based on deep learning, including: An image acquisition module 301, configured to acquire student image data in the classroom, perform image enhancement, light compensation, and normalization processing on the student image data to obtain preprocessed image data; A feature extraction module 302, configured to perform multi-scale face detection on the preprocessed image data, extract facial feature points and perform normalization processing to obtain normalized facial feature data; A state recognition module 303, configured to extract expression features based on the normalized facial feature data, combine attention metrics and pose features, and perform feature fusion processing to obtain a state feature vector; A time series analysis module 304, configured to construct a time series feature sequence for the state feature vector, perform time series modeling through a heavy-tailed recurrent neural network, analyze the state change trend, and obtain a state time series analysis result; An evaluation module 305, configured to establish an evaluation criterion based on the state time series analysis result, adjust the parameters of the evaluation criterion through an adaptive threshold, and obtain a student state evaluation result.

[0091] Through the heavy-tailed recurrent neural network and the slow transition mechanism to low-dimensional chaos, the present invention can accurately capture the subtle changes and long-term trends of students' states, improving the sensitivity and prediction accuracy for abnormal states; adopting the multi-modal feature fusion and dynamic weight allocation technology to integrate three modal features of facial expression features, attention indicators and body postures, improving the robustness and adaptability of students' state recognition; establishing an evaluation criterion and an adaptive threshold adjustment mechanism to make adaptive adjustments according to students' individual differences, achieving a more accurate evaluation of students' states; through multi-scale time series feature analysis, analyzing short-term fluctuations and long-term trends simultaneously, identifying the inflection points of concentration and emotional change patterns, improving the timeliness and accuracy of state prediction; realizing the quantification and compensation of environmental influencing factors, reducing the interference of external factors on the recognition results, and improving the adaptability in complex environments.

[0092] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A real-time analysis method for student status based on deep learning, characterized in that, Including: Obtain student image data in the classroom, perform image enhancement, light compensation, and normalization processing on the student image data to obtain preprocessed image data; Perform multi-scale face detection on the preprocessed image data, extract facial feature points and perform normalization processing to obtain normalized facial feature data; Based on the normalized facial feature data, extract expression features through a convolutional neural network, combine attention metrics and pose features, and perform feature fusion processing to obtain a state feature vector; Construct a temporal feature sequence for the state feature vector, perform temporal modeling through a heavy-tailed recurrent neural network, analyze the state change trend, and obtain a state temporal analysis result; Based on the state temporal analysis result, establish an evaluation criterion, adaptively adjust the parameters of the evaluation criterion through an adaptive threshold, and perform state evaluation on students through the adjusted evaluation criterion to obtain a student state evaluation result.

2. The method according to claim 1, characterized in that Performing image enhancement, light compensation, and normalization processing on the student image data includes: Based on the student image data, calculate an image sharpness score, and screen out valid images with an image sharpness score higher than a preset threshold; Perform adaptive histogram equalization processing on the valid images to obtain images with balanced brightness; Apply a bilateral filtering algorithm to the images with balanced brightness to remove image noise and obtain denoised images; Perform size adjustment and pixel value normalization on the denoised images to obtain the preprocessed image data.

3. The method according to claim 2, wherein Performing multi-scale face detection on the preprocessed image data, extracting facial feature points and performing normalization processing includes: Apply a cascaded convolutional neural network to the preprocessed image data for face region detection to obtain face detection box data; Apply a non-maximum suppression algorithm to the face detection box data to eliminate overlapping detection boxes and obtain optimized face position data; Extract facial feature points based on the optimized face position data to obtain facial feature point coordinate data; Perform pose angle estimation and geometric correction on the facial feature point coordinate data to obtain the normalized facial feature data.

4. The method according to claim 3, wherein Based on the normalized facial feature data, extracting expression features through a convolutional neural network, combining attention metrics and pose features, and performing feature fusion processing includes: Apply a convolutional neural network to the normalized facial feature data to extract expression features of the eye, eyebrow, and mouth regions to obtain an expression feature vector; Calculate the eye opening degree and blink frequency based on the normalized facial feature data to obtain an attention metric feature vector; Apply a pose estimation network to the preprocessed image data to extract upper body key points to obtain a pose feature vector; Perform weighted fusion on the expression feature vector, attention metric feature vector, and pose feature vector by calculating attention weight coefficients to obtain the state feature vector.

5. The method according to claim 4, wherein Constructing a temporal feature sequence for the state feature vector includes: Arrange the state feature vector in chronological order to obtain original feature temporal data; Set sliding windows of multiple time scales for the original feature temporal data to obtain segmented temporal feature data; Process the segmented time-series feature data using a heavy-tailed recurrent neural network to obtain state change feature data; Based on the state change feature data, identify the inflection points of the attention change, predict the state change trend, and obtain the state time-series analysis result.

6. The method according to claim 5, wherein Processing the segmented time-series feature data using a heavy-tailed recurrent neural network includes: Introduce an activation function of the α-stable distribution to the segmented time-series feature data to obtain feature data with heavy-tailed characteristics; Apply an attention mechanism to the feature data with heavy-tailed characteristics for long-term dependence modeling to obtain time-series correlation feature data; Dynamically adjust the network parameter entropy value of the time-series correlation feature data to obtain feature data with introduced randomness; Calculate the state transition probability matrix based on the feature data with introduced randomness to obtain state transition feature data; Apply a multi-task learning strategy to train and optimize the state transition feature data to obtain the state change feature data.

7. The method according to claim 1, wherein Based on the state time-series analysis result, establish an evaluation criterion, including: Extract the attention mean and standard deviation features from the state time-series analysis result to obtain statistical feature data; Apply a clustering algorithm to the statistical feature data for state pattern division to obtain state distribution data; Calculate the normal attention interval and fatigue threshold based on the state distribution data to obtain state evaluation criterion parameters; Conduct a multivariate regression analysis on the state evaluation criterion parameters combined with environmental parameters to obtain environmental impact coefficient data; Modify the state judgment criterion based on the environmental impact coefficient data to obtain the evaluation criterion.

8. The method according to claim 7, wherein Adjust the parameters of the evaluation criterion through an adaptive threshold, and evaluate the state of students through the adjusted evaluation criterion to obtain the student state evaluation result, including: Apply a decision tree model to the environmental impact coefficient data to quantify the influence weights of various environmental factors to obtain environmental weight data; Calculate the state deviation threshold under the current environmental conditions based on the environmental weight data to obtain adaptive threshold data; Dynamically adjust the evaluation criterion based on the adaptive threshold data combined with historical state evaluation data; Evaluate the real-time state of students based on the adjusted evaluation criterion, calculate the state deviation degree, and obtain state evaluation data; Generate a change trend graph of attention and emotional state for the state evaluation data to obtain the student state evaluation result.

9. The method according to claim 1, characterized in that, It also includes: Generate class state visualization data according to the student state evaluation result, specifically including: Conduct statistical clustering processing on the student state evaluation result to obtain the overall class state distribution data; Generate an attention heat map based on the overall class state distribution data to obtain spatial distribution visualization data.

10. A real-time analysis device for students' states based on deep learning, characterized in that, It includes: An image acquisition module for acquiring student image data in the classroom, performing image enhancement, light compensation, and normalization processing on the student image data to obtain preprocessed image data; A feature extraction module for performing multi-scale face detection on the preprocessed image data, extracting facial feature points and performing normalization processing to obtain normalized facial feature data; A state recognition module, which is used to extract expression features through a convolutional neural network based on the standardized facial feature data, combine attention metrics and pose features, and perform feature fusion processing to obtain a state feature vector; A time series analysis module, which is used to construct a time series feature sequence for the state feature vector, perform time series modeling through a heavy-tailed recurrent neural network, analyze the state change trend, and obtain a state time series analysis result; An evaluation module, which is used to establish an evaluation criterion based on the state time series analysis result, adaptively adjust the parameters of the evaluation criterion through an adaptive threshold, and perform state evaluation on students through the evaluation criterion with adjusted parameters to obtain a student state evaluation result.

Citation Information

Patent Citations

  • Multi-mode information fusion-based classroom learning state monitoring method and system

    CN108805009A

  • Classroom concentration degree detection method and device

    CN111931585A

  • Classroom student state analysis method and system based on face monitoring

    CN117496575A

  • Video image processing method based on ambient light standardization processing, medium and equipment

    CN118135162A

  • Character action recognition analysis method and system based on infrared laser and deep learning

    CN118747911A

Cited By

  • Facial close-up generation method and device based on improved GAN, equipment and medium

    CN120543716A

  • LIMS laboratory process maintenance method based on adaptive learning model

    CN121482712A