A Classroom Student Attention Assessment Method and System Based on Multi-Source Feature Fusion

By adopting a multi-source feature fusion method in the classroom environment, combining student expression recognition neural network model and extreme learning machine model, students' attention is evaluated, and the problems of large errors and insufficient accuracy in the existing technology are solved, and more accurate and personalized attention assessment is achieved.

CN119559703BActive Publication Date: 2025-06-10ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510099263.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-10
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing technology has many limitations when evaluating the attention of students in classrooms, including the difficulty in monitoring student goals with small scales, low pixels, and multiple occlusions, the singularity of evaluation methods, and the lack of effective joint models, resulting in the evaluation results relying on a certain standard, large errors, insufficient accuracy of expression recognition technology, resulting in bias in attention assessment.

Method used

The attention evaluation method of classroom students based on multi-source feature fusion is adopted, students' video frames are obtained through the camera, face detection and expression recognition are carried out, and students' expression recognition neural network model SFaceNet is constructed. Combined with the extreme learning machine model ELM, students' different emotional states, fatigue degrees and head postures are fused to perform attention evaluation with multi-dimensional cross-coordinated optimization.

Benefits of technology

It improves the robustness and generalization ability of complex classroom environments, optimizes attention analysis ability, reduces evaluation errors, enhances the accuracy and consistency of evaluation results, supports the implementation of personalized teaching strategies, and achieves real-time response in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559703B_ABST
    Figure CN119559703B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for evaluating classroom student attention based on multi-source feature fusion. In the present invention, face recognition is performed through the face detection unit of the system, the occurrence frequency of different emotional states of students is obtained through the student expression recognition neural network model including four expression discrimination blocks completed by training in the expression recognition unit, the number of times of students' fatigue is obtained through the fatigue detection unit, the number of times of students' abnormal head postures is obtained through the head pose estimation unit, and finally the student attention evaluation parameters are obtained through the extreme learning machine model of the attention evaluation unit, so as to evaluate the student attention state in the classroom and display it. The present invention can provide a more efficient dynamic student expression recognition method, provide real-time and accurate classroom feedback for teachers, support the optimization and implementation of personalized teaching strategies, facilitate the longitudinal extension of teachers' teaching, the horizontal expansion of students' learning methods, and expand the new paradigm for cultivating new-quality talents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an attention evaluation method, belonging to the technical fields of artificial intelligence and education informatization, and specifically to a method and system for evaluating students' classroom attention based on multi-source feature fusion. Background Art

[0002] Currently, in the teaching scenario, there are still several limitations in the methods for evaluating students' classroom attention. First, in most teaching scenarios, due to the one-to-many teaching mode, the student targets with small scale, low pixel, and multiple occlusions greatly increase the difficulty of monitoring students' attention and behavior, thus affecting the generalization ability and practical application effect of existing models. Second, the singularity and low dimension of the evaluation means limit the comprehensive consideration of numerous factors affecting students' attention. Third, there is no effective joint model between different evaluation criteria, resulting in the evaluation results being overly dependent on a certain criterion, and thus the evaluation error cannot be effectively reduced. Finally, the lack of accuracy in the student expression recognition technology in existing evaluation models also causes deviations in the evaluation of students' attention. Therefore, there is an urgent need to construct a comprehensive evaluation paradigm system with multi-dimensional intersection and collaborative optimization to improve the accurate judgment of students' attention states, and thus effectively support the deep embedding of personalized teaching strategies. Summary of the Invention

[0003] To solve the problems in the background art, the present invention provides a method and system for evaluating students' classroom attention based on multi-source feature fusion. The method of the present invention provides a more efficient Dynamic Student Expression Recognition (DSER) method, provides real-time and accurate classroom feedback for teachers, supports the optimization and implementation of personalized teaching strategies, is beneficial to the longitudinal extension of teachers' teaching and the horizontal expansion of students' learning methods, further opens up a new track for AI education, and expands a new paradigm for cultivating new-quality talents.

[0004] The technical solution adopted by the present invention is as follows:

[0005] I. A method for evaluating students' classroom attention based on multi-source feature fusion:

[0006] Step S1: Obtain several videos of students' classroom lectures through a camera, extract the initial video frames, perform face detection preprocessing on each initial video frame, and then identify different emotional states of the face through face key point positioning and add expression labels, so as to obtain several face expression images and construct a face expression image dataset.

[0007] Step S2: Construct a student facial expression recognition neural network model SFaceNet that includes four facial expression discrimination blocks, and train the student facial expression recognition neural network model SFaceNet with a facial expression image dataset.

[0008] Step S3: Obtain the video of students' classroom lectures to be recognized through a camera and extract the video frames to be recognized. Then, perform face recognition, and further recognize the occurrence frequencies of different emotional states of each student through the trained student facial expression recognition neural network model SFaceNet. The fatigue degree and the number of abnormal head postures of each student are recognized through face key point localization.

[0009] Step S4: For each student, input the occurrence frequencies of different emotional states, the fatigue degree, and the number of abnormal head postures of the student into a pre-trained extreme learning machine model ELM (Extreme Learning Machine). After processing, output the attention evaluation parameters of the student, and evaluate and display the attention state of the students in the classroom according to the attention evaluation parameters of each student.

[0010] In the described step S1, each initial video frame is subjected to face recognition through the face detection model RetinaFace to obtain an image with a detection frame with the facial features of each student, and then data augmentation processing is performed, including random cropping, resizing to a unified size, random horizontal flipping, and image normalization processing, so as to obtain pre-processed video frames.

[0011] In the described step S1, each pre-processed video frame is used to recognize four emotional states of the human face, including rejecting expression, listening expression, confused expression, and excited expression, as typical classroom emotional expressions through the 468 face key point localization method based on machine learning Mediapipe and the Russell emotion model. For the human face of the student in each detection frame of the pre-processed video frame, the four emotional states of the human face are as follows:

[0012] a) Rejecting expression:

[0013] Judge the tilt degree of the mouth. When the angle formed by the line connecting the 62nd key point and the 292nd key point in the human face and the 17th key point and the horizontal line is negative, indicating that it is downward relative to the horizontal line direction, and the tilt degree of the mouth of the human face is that the corners of both sides of the mouth move downward, then it is determined that the emotional state of the current human face is a rejecting expression, and there are no obvious features in the eyes.

[0014] b) Listening expression:

[0015] Judge the tilt degree of the mouth. When the line connecting the 62nd key point and the 292nd key point in the human face and the 17th key point is approximately parallel to the horizontal line, the tilt degree of the mouth of the human face is not obvious.

[0016] Determine the degree of eye opening. Obtain the eye height of the left eye in the eye features by subtracting the vertical coordinates of the 470th key point and the 472nd key point in the face. Obtain the eye width of the left eye in the eye features by subtracting the horizontal coordinates of the 33rd key point and the 133rd key point in the face. Obtain the eye height of the right eye in the eye features by subtracting the vertical coordinates of the 475th key point and the 477th key point in the face. Obtain the eye width of the right eye in the eye features by subtracting the horizontal coordinates of the 362nd key point and the 263rd key point in the face. The ratios of the eye height to the eye width of the left eye and the right eye are both greater than the ratio threshold, and the degree of eye opening of the face is natural opening.

[0017] Determine the degree of eyebrow bending. The angles formed by the lines connecting the 70th key point and the 65th key point in the face to the 105th key point respectively are greater than the angle threshold. The angles formed by the lines connecting the 300th key point and the 295th key point in the face to the 334th key point respectively are greater than the angle threshold. The degree of eyebrow bending of the face is natural stretching.

[0018] When the degree of mouth tilt of the face is not obvious, the degree of eye opening of the face is natural opening, and the degree of eyebrow bending of the face is natural stretching, then it is determined that the emotional state of the current face is a listening expression;

[0019] c) Confused expression:

[0020] Determine the degree of eyebrow bending. The angles formed by the lines connecting the 70th key point and the 65th key point in the face to the 105th key point respectively are less than the angle threshold. The angles formed by the lines connecting the 300th key point and the 295th key point in the face to the 334th key point respectively are less than the angle threshold. The degree of eyebrow bending of the face is frowning;

[0021] When the degree of mouth tilt of the face is that the corners of both sides of the mouth move downward, and the degree of eyebrow bending of the face is frowning, then it is determined that the emotional state of the current face is a confused expression.

[0022] d) Excited expression:

[0023] Determine the degree of mouth tilt. When the angles formed by the lines connecting the 62nd key point and the 292nd key point in the face to the 17th key point and the horizontal line are positive, indicating upward relative to the horizontal line direction, the degree of mouth tilt of the face is that the corners of both sides of the mouth turn up.

[0024] When the degree of mouth tilt of the face is that the corners of both sides of the mouth turn up, the degree of eye opening of the face is natural opening, the degree of eyebrow bending of the face is natural stretching, and the Euclidean distance between the 12th key point and the 14th key point of the mouth of the face is greater than the preset distance threshold. Specifically, when implemented, it is set to 2 pixels, then it is determined that the emotional state of the current face is a listening expression.

[0025] An expression label of an emotional state is added to the face of each student in each detection box.

[0026] In the said step S2, the student facial expression recognition neural network model SFaceNet includes an input layer, a feature extraction and fusion layer, and an output layer connected in sequence. The input layer includes a first convolutional layer, a first batch normalization layer (Batch Norm), a first activation function layer (ReLu), and a max pooling layer (Maxpooling) connected in sequence. The feature extraction and fusion layer includes three emotion discrimination blocks one (EmotionBlock V1), four emotion discrimination blocks two (EmotionBlock V2), six emotion discrimination blocks three (EmotionBlock V3), three emotion discrimination blocks four (EmotionBlock V4), and an attention mechanism module (CBAM). It also includes a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. The output of the max pooling layer (Maxpooling) is input into the first emotion discrimination block one (EmotionBlock V1). The output of the max pooling layer (Maxpooling) and the output of the first emotion discrimination block one (EmotionBlock V1) are added to obtain a first addition result. The first addition result is input into the second emotion discrimination block one (EmotionBlock V1). The first addition result and the output of the second emotion discrimination block one (EmotionBlock V1) are added to obtain a second addition result. The second addition result is input into the third emotion discrimination block one (EmotionBlock V1). The second addition result and the output of the third emotion discrimination block one (EmotionBlock V1) are added to obtain a third addition result. The third addition result is input into the first emotion discrimination block two (EmotionBlock V2). The third addition result is processed by the second convolutional layer and added to the output of the first emotion discrimination block two (EmotionBlock V2) to obtain a fourth addition result. The fourth addition result is input into the second emotion discrimination block two (EmotionBlock V2). The fourth addition result and the output of the second emotion discrimination block two (EmotionBlock V2) are added to obtain a fifth addition result. The fifth addition result is input into the third emotion discrimination block two (EmotionBlock V2). The fifth addition result and the output of the third emotion discrimination block two (EmotionBlock V2) are added to obtain a sixth addition result. The sixth addition result is input into the fourth emotion discrimination block two (EmotionBlock V2). The sixth addition result and the output of the fourth emotion discrimination block two (EmotionBlock V2) are added to obtain a seventh addition result. The seventh addition result is input into the first emotion discrimination block three (EmotionBlock V3). The seventh addition result is processed by the third convolutional layer and added to the output of the first emotion discrimination block three (EmotionBlock V3) to obtain an eighth addition result. The eighth addition result is input into the second emotion discrimination block three (EmotionBlock V3).The result of the eighth addition is added to the output of the second Emotion Discrimination Block III (EmotionBlock V3) to obtain the result of the ninth addition. The result of the ninth addition is input into the third Emotion Discrimination Block III (EmotionBlock V3). The result of the ninth addition is added to the output of the third Emotion Discrimination Block III (EmotionBlock V3) to obtain the result of the tenth addition. The result of the tenth addition is input into the fourth Emotion Discrimination Block III (EmotionBlock V3). The result of the tenth addition is added to the output of the fourth Emotion Discrimination Block III (EmotionBlock V3) to obtain the result of the eleventh addition. The result of the eleventh addition is input into the fifth Emotion Discrimination Block III (EmotionBlock V3). The result of the eleventh addition is added to the output of the fifth Emotion Discrimination Block III (EmotionBlock V3) to obtain the result of the twelfth addition. The result of the twelfth addition is input into the sixth Emotion Discrimination Block III (EmotionBlock V3). The result of the twelfth addition is added to the output of the sixth Emotion Discrimination Block III (EmotionBlock V3) to obtain the result of the thirteenth addition. The result of the thirteenth addition is processed by the fourth convolutional layer and added to the output of the first Emotion Discrimination Block IV (EmotionBlock V4) to obtain the result of the fourteenth addition. The result of the fourteenth addition is input into the second Emotion Discrimination Block IV (EmotionBlock V4). The result of the fourteenth addition is added to the output of the second Emotion Discrimination Block IV (EmotionBlock V4) to obtain the result of the fifteenth addition. The result of the fifteenth addition is input into the third Emotion Discrimination Block IV (EmotionBlock V4). The result of the fifteenth addition is added to the output of the third Emotion Discrimination Block IV (EmotionBlock V4) and then input into the Convolutional Block Attention Module (CBAM); The output layer includes an average pooling layer (Avgpooling) and a fully connected layer (FC, Fully Connected) connected in sequence.

[0027] The first Emotion Discrimination Block I (EmotionBlock V1), the second Emotion Discrimination Block II (EmotionBlock V2), the third Emotion Discrimination Block III (EmotionBlock V3), and the fourth Emotion Discrimination Block IV (EmotionBlock V4) include a second batch normalization layer (Batch Norm), a second activation function layer (ReLu), a fifth convolutional layer, a third batch normalization layer (Batch Norm), a third activation function layer (ReLu), a sixth convolutional layer, a Contextual Transformer (CoT) module, a fourth batch normalization layer (Batch Norm), a fourth activation function layer (ReLu), and a seventh convolutional layer connected in sequence.

[0028] In the said step S3, for each video frame to be recognized, each video frame to be recognized is subjected to face recognition through the face detection model RetinaFace to obtain an image of a detection frame with the facial features of each student. Then, the image of the detection frame with the facial features of each student is input into the trained student expression recognition neural network model SFaceNet for processing to obtain the emotional state of each student in the current video frame. Finally, according to the images of the detection frames with the facial features of each student, the occurrence frequency of different emotional states of each student during the shooting period of the student classroom lecture video to be recognized is identified.

[0029] In the said step S3, for each image of a detection frame with the facial features of each student, first, it is processed through the 468-face key point localization method based on machine learning Mediapipe to obtain 468 key point information of the face of each student in the graph. For each student, 6 key points are taken from each of the three parts of the left eye, right eye, and mouth of the student's face, a total of 18 key points. According to the 6 key points of the left eye and right eye of the student's face, the eye aspect ratio (EAR) is obtained to further judge the closed-eye state of the student. According to the 6 key points of the mouth of the student's face, the mouth aspect ratio (MAR) is obtained to further judge the yawning state of the student. Finally, according to the images of the detection frames with the facial features of each student, the number of times each student closes their eyes and yawns during the shooting period of the student classroom lecture video to be recognized is identified as the number of times of fatigue of each student.

[0030] In step S3, for each video frame to be recognized and for each image of the detection box with the facial features of each student to be recognized, first, it is processed by the 468 facial key point localization method based on the machine learning Mediapipe to obtain the two-dimensional coordinates of 468 key points of each student's face in the graph. Then, normalization processing of the two-dimensional coordinates is performed to avoid errors caused by differences in the sizes of different input images, making subsequent calculations more stable and general. Then, the two-dimensional coordinates of each normalized key point are mapped into the three-dimensional world coordinate system. For each student, the pose estimation algorithm PnP (Perspective-n-Points) is used to match the three-dimensional coordinates and two-dimensional coordinates of each key point of the student's face in the three-dimensional world coordinate system, so as to obtain the rotation and displacement of the camera, and finally obtain the three-degree-of-freedom information of the student's head, including the rotation Yaw around the vertical axis (z-axis), the rotation Pitch around the horizontal axis (y-axis), and the rotation Roll around the longitudinal axis (x-axis). According to the three-degree-of-freedom information of the student's head, the bowing state of the student is judged. Finally, based on each image of the detection box with the facial features of each student to be recognized, the number of times each student bows their head during the shooting period of the video of the student's classroom lecture to be recognized is identified as the number of times of the abnormal head posture of each student.

[0031] In step S4 described above, the obtained seven-dimensional eigenvalue is used as a 1×7 input vector ( x 1 : number of times of bowing the head; x 2 : number of times of yawning; x 3 : number of times of closing eyes; x 4 : number of times of the appearance of the listening expression; x 5 : number of times of the appearance of the confused expression; x 6 : number of times of the appearance of the resistant expression; x 7 : number of times of the appearance of the excited expression) and input into the pre-trained extreme learning machine model ELM model for the attention evaluation of the student. Finally, an attention evaluation parameter between 0 and 1 is obtained. Then, the average value of the attention evaluation parameters of each student at each moment is used as the average attention evaluation parameter of the classroom at the current moment. Introducing the time dimension, time series analysis is carried out. The preset time interval is 1 min, and the video records once every 15 s to obtain an average attention evaluation parameter of the classroom at a moment. The average value of the average attention evaluation parameters of the classroom at several moments within each preset time interval is used as the attention evaluation parameter of the classroom unit cycle, so as to judge the attention state of the students in the classroom as follows:

[0032] When 0≤Focus t When it is less than 0.4, the students’ attention in the class is very poor, that is, the teaching effect is inefficient. t It is the attention assessment parameter of the class unit cycle at the current preset time interval;

[0033] When 0.4≤Focus t When it is less than 0.6, the students’ attention in the class is poor, which means the teaching effect is not good.

[0034] When 0.6≤Focus t When it is <0.8, the students’ attention level in the class is average, which means the teaching effect is acceptable.

[0035] When 0.8≤Focus t When it is less than 1.0, the students in the class are paying very good attention, which means the class is effective.

[0036] 2. A classroom student attention assessment system based on multi-source feature fusion:

[0037] The face detection unit uses the face detection model RetinaFace to perform face recognition on the video frames of the student's classroom lecture video obtained through the camera, and then outputs an image of the detection frame with each student's facial features as the face detection result. It can realize accurate detection of the student's face area and ensure that small-scale, low-pixel, and multi-occluded student faces can be effectively identified in complex classroom scenes.

[0038] The expression recognition unit constructs and trains the student expression recognition neural network model SFaceNet, which contains four expression discrimination blocks. The trained student expression recognition neural network model SFaceNet is used to identify the different emotional states of each student as the expression recognition result, and then the occurrence frequency of each student's different emotional states is obtained.

[0039] The fatigue detection unit locates the key points of the face and identifies the number of times each student closes his eyes and yawns as the number of fatigue levels.

[0040] The head posture estimation unit identifies the number of times each student lowers his head as the number of abnormal head postures by locating facial key points.

[0041] The attention assessment unit inputs the frequency of occurrence of students' different emotional states, fatigue levels, and the number of abnormal head postures into the pre-trained extreme learning machine model ELM. After processing, it outputs the students' attention assessment parameters. Based on the attention assessment parameters of each student, the attention status of the students in the class is evaluated and displayed on the monitor.

[0042] The beneficial effects of the present invention are:

[0043] 1. The present invention improves the robustness and generalization ability in complex classroom environments: In teaching scenarios, the difficulty of obtaining facial information of students increases significantly due to problems such as position, occlusion, and low resolution. Based on the multi-task optimization scheme of the RetinaFace neural network for face detection, the present invention significantly enhances the recognition accuracy and stability of the system under small-scale, low-resolution, and complex occlusion conditions, ensuring the accurate capture of students' facial features.

[0044] 2. The present invention performs multi-source feature fusion and optimizes attention parsing: By deeply extracting multi-source features such as students' expressions, fatigue levels, and head postures, and comprehensively utilizing facial geometric key points and dynamic expression trajectories, the attention parsing ability of the system is improved in a multi-dimensional cross-fusion manner, thereby overcoming the analysis limitations brought by traditional single-dimensional features and ensuring the fine-grained and three-dimensional attention assessment.

[0045] 3. The present invention realizes an efficient collaborative evaluation paradigm: By inputting multi-source feature data into the Extreme Learning Machine model ELM, a jointly evaluated architecture with deep collaborative optimization is constructed. Through the multi-dimensional feature space mapping and weight optimization of this model, the deviation caused by a single evaluation criterion is significantly reduced, thereby enhancing the accuracy and consistency of the evaluation results and providing an efficient paradigm for the objective quantitative evaluation of classroom performance.

[0046] 4. The present invention improves the accuracy and recognition rate of expression recognition: The SFaceNet neural network model for student expression recognition built by self-making a high-quality expression dataset and introducing the Convolutional Block Attention Module (CBAM) and the Context Transformer Module (CoT) significantly improves the accuracy and recognition rate of the present invention in student expression recognition. The efficient emotional state discrimination and dynamic expression extraction enable the system to perform more detailed emotional analysis of students' classroom performance, improving the accuracy of the attention model in different emotional states.

[0047] 5. The present invention has refined fatigue and head posture analysis: Based on the Mediapipe key point detection and the Perspective-n-Point (PnP) algorithm for pose estimation, the system realizes the three-degree-of-freedom estimation of students' head postures by calculating the pose matrix, accurately monitors behavior indicators such as the number of times students lower their heads, and precisely quantifies fatigue behaviors such as closing eyes and yawning, further improving the analysis accuracy of classroom performance.

[0048] 6. The present invention supports the deep embedding of personalized teaching feedback: Through the generation of high-precision attention time-series evaluation parameters, the system can provide detailed classroom feedback data for teachers, supporting the quantitative deployment of personalized teaching strategies. The intelligent evaluation driven by multi-source data enables teachers to grasp students' cognitive states in real time, providing an objective support for precise teaching and differentiated teaching methods.

[0049] 7. The present invention meets the real-time response requirements in a dynamic and complex environment: Due to the introduction of the efficient Extreme Learning Machine model ELM, the data processing and feedback response speed of this system have been optimized, enabling real-time monitoring and evaluation in a dynamic classroom environment, ensuring the efficient feedback ability of the attention state, and thus supporting the real-time adjustment and optimization of the teaching process. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the flowchart of the method in the specific implementation of the present invention;

[0051] Figure 2 is the flowchart of the classification and recognition of the expression dataset of the present invention;

[0052] Figure 3 is the network structure diagram of the student expression recognition network model SFaceNet of the present invention;

[0053] Figure 4 is the network structure diagram of the first expression discrimination block EmotionBlock V1, the second expression discrimination block EmotionBlock V2, the third expression discrimination block EmotionBlock V3, and the fourth expression discrimination block EmotionBlock V4 of the present invention;

[0054] Figure 5 is the flowchart of the fatigue detection of the present invention;

[0055] Figure 6 is the flowchart of the student attention evaluation of the present invention;

[0056] Figure 7 is the parameter curve graph of the classroom student attention evaluation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following is a detailed elaboration of the specific embodiments of the present invention, with reference to the accompanying illustrations. It should be noted that the embodiments shown here are only a part of the numerous possible implementation manners, rather than all. These descriptions should not be construed as limiting the protection scope of the present invention. According to the spirit of the present invention, any other implementation manners that do not involve creative labor and can be easily conceived by those skilled in the art based on this embodiment shall be regarded as within the protection scope of the present invention.

[0058] As Figure 1 shown, the method for evaluating classroom student attention based on multi-source feature fusion of the present invention is specifically as follows:

[0059] In the specific implementation of the present invention, first, the facial expression recognition dataset FER2013 (Facial Emotion Recognition 2013) is obtained. Each facial expression image in the facial expression recognition dataset FER2013 is recalibrated, and the original label of each facial expression image is removed. Then, different emotional states of the human face are identified through face key point localization and expression labels are added. That is, the facial expression recognition dataset FER2013 is used to identify four emotional states of the human face, including rejection, listening, confusion, and excitement, as typical classroom emotional expressions, through the 468 face key point localization method based on machine learning Mediapipe and the Russell emotion model. Then, manual recalibration is performed, and finally, the recalibrated facial expression recognition dataset FER2013 is obtained. As shown in Table 1, when the 468-point coordinate information of a human face is obtained using the 468 face key point localization method based on machine learning Mediapipe, relevant parameter calculations are performed according to the theoretical features given in Table 1, and the expression labels of the facial expression recognition dataset FER2013 are redefined for subsequent neural network training, such as Figure 2 shown; for the face of the student in each detection box of the preprocessed video frame, the four emotional states of the face are as follows:

[0060] a) Rejection expression:

[0061] Judge the tilt degree of the mouth. When the angle formed between the line connecting the 62nd key point and the 292nd key point in the human face and the 17th key point and the horizontal line is negative, indicating that it is downward relative to the horizontal line direction, and the tilt degree of the mouth of the human face is that the corners of the mouth move downward on both sides, then it is determined that the emotional state of the current human face is a rejection expression, and there are no obvious features in the eyes.

[0062] b) Listening expression:

[0063] Judge the tilt degree of the mouth. When the line connecting the 62nd key point and the 292nd key point in the human face and the 17th key point is approximately parallel to the horizontal line, the tilt degree of the mouth of the human face is not obvious.

[0064] Judge the degree of eye opening. The eye height of the left eye in the eye features is obtained by subtracting the ordinate of the 470th key point and the 472nd key point in the human face. The eye width of the left eye in the eye features is obtained by subtracting the abscissa of the 33rd key point and the 133rd key point in the human face. The eye height of the right eye in the eye features is obtained by subtracting the ordinate of the 475th key point and the 477th key point in the human face. The eye width of the right eye in the eye features is obtained by subtracting the abscissa of the 362nd key point and the 263rd key point in the human face. The ratios of the eye height to the eye width of the left eye and the eye height to the eye width of the right eye are both greater than the ratio threshold, and the degree of eye opening of the human face is naturally open. In the specific implementation, the ratio threshold is set to 0.25.

[0065] Determine the degree of eyebrow curvature. The angles formed by the lines connecting the 70th key point and the 65th key point in the human face to the 105th key point respectively are greater than the angle threshold, and the angles formed by the lines connecting the 300th key point and the 295th key point in the human face to the 334th key point respectively are greater than the angle threshold. The degree of eyebrow curvature of the human face is natural and relaxed. Specifically, when implemented, the angle threshold is 120°.

[0066] When the degree of tilt of the mouth of the human face is not obvious, the degree of eye opening of the human face is natural, and the degree of eyebrow curvature of the human face is natural and relaxed, then it is determined that the emotional state of the current human face is a listening expression;

[0067] c) Confused expression:

[0068] Determine the degree of eyebrow curvature. The angles formed by the lines connecting the 70th key point and the 65th key point in the human face to the 105th key point respectively are less than the angle threshold, and the angles formed by the lines connecting the 300th key point and the 295th key point in the human face to the 334th key point respectively are less than the angle threshold. The degree of eyebrow curvature of the human face is frowning.

[0069] When the degree of tilt of the mouth of the human face is that the corners of both sides of the mouth move downward, and the degree of eyebrow curvature of the human face is frowning, then it is determined that the emotional state of the current human face is a confused expression.

[0070] d) Excited expression:

[0071] Determine the degree of tilt of the mouth. When the angles formed by the lines connecting the 62nd key point and the 292nd key point in the human face to the 17th key point and the horizontal line are positive, it indicates that it is upward relative to the horizontal line direction. The degree of tilt of the mouth of the human face is that the corners of both sides of the mouth turn up.

[0072] When the degree of tilt of the mouth of the human face is that the corners of both sides of the mouth turn up, the degree of eye opening of the human face is natural, the degree of eyebrow curvature of the human face is natural and relaxed, and the Euclidean distance between the 12th key point and the 14th key point of the mouth of the human face is greater than the preset distance threshold. Specifically, when implemented, it is set to 2 pixels, then it is determined that the emotional state of the current human face is a listening expression.

[0073] Table 1 Definition of expression parameters

[0074]

[0075] Then, the lecture videos of 15 students are obtained through a camera, and the initial video frames are extracted. Each of the 15 students watches a 4-minute online class in front of the camera, making four expressions of listening, confusion, rejection, and excitement every minute. Thus, a number of initial video frames with one of the four expressions are obtained, and each initial video frame is used to construct a self-made expression dataset. The self-made expression dataset is preprocessed by face detection. Each initial video frame is subjected to face recognition through the face detection model RetinaFace. After obtaining the class video stream, the face detection model RetinaFace processes the video images frame by frame and finally returns the detected faces in the video and their key point information, which is stored in the array b

[14] . The array b

[14] is an array (vector) with a length of 15, corresponding to the information of a detected face. Specifically, as follows: b[0] represents the abscissa of the upper left corner of the face box in the original image, b[1] represents the ordinate of the upper left corner of the face box in the original image, b[2] represents the abscissa of the lower right corner of the face box in the original image, b[3] represents the ordinate of the lower right corner of the face box in the original image, b[4] represents the confidence score indicating whether the predicted box is a face, and the higher the value, the greater the probability of the face. b[5] represents the abscissa of the left eye in the face key points, b[6] represents the ordinate of the left eye in the face key points, b[7] represents the abscissa of the right eye in the face key points, b[8] represents the ordinate of the right eye in the face key points, b[9] represents the abscissa of the nose in the face key points, b

[10] represents the ordinate of the nose in the face key points, b

[11] represents the abscissa of the left corner of the mouth in the face key points, b

[12] represents the ordinate of the left corner of the mouth in the face key points, b

[13] represents the abscissa of the right corner of the mouth in the face key points, and b

[14] represents the ordinate of the right corner of the mouth in the face key points. Finally, the facial features of the 15 students in the detection box are added with expression labels according to the four expressions of the present invention, and finally a face expression image dataset is constructed.

[0076] The face detection model RetinaFace uses the ResNet50 as the backbone network to detect the faces of students in each frame of the teaching video to ensure the accurate recognition of students' facial features. After testing, the face detection model RetinaFace with the ResNet50 as the backbone network has shown strong adaptability and efficient face detection and recognition capabilities in the face of various challenges and complex situations existing in the actual classroom. In many situations in the actual classroom, such as frequent student activities, uneven light, variable camera angles, small-scale recognition targets, low pixels, and multiple occlusions, the face detection model RetinaFace can still quickly and accurately recognize the facial features of each student.

[0077] Then, the re-calibrated facial expression recognition dataset FER2013 and the facial expression image dataset are mixed in a ratio of 7:3 to construct an initial dataset for training the expression recognition neural network. In the initial dataset, there are 8,774 excited data, 5,740 confused data, 6,628 listening data, and 5,509 rejecting data. The pixel size of the images of these four types of data is 48×48, and the number of channels is 3.

[0078] Then, a student expression recognition neural network model SFaceNet containing four expression discrimination blocks is constructed, as Figure 3 shown. The student expression recognition neural network model SFaceNet includes an input layer, a feature extraction and fusion layer, and an output layer connected in sequence. The input layer includes a first convolutional layer, a first batch normalization layer Batch Norm, a first activation function layer ReLu, and a max pooling layer Maxpooling connected in sequence. It is responsible for receiving the 3-channel, 224×224 face image data input from the facial expression image dataset, and changing the image size to 64×112×112 through the first convolutional layer with a size of 7×7, a stride of 2, and a padding of 3 to reduce the image size and extract preliminary features. Finally, the image size is changed to 64×56×56 through the max pooling layer Maxpooling with a size of 3×3, a stride of 2, and a padding of 1.

[0079] The feature extraction and fusion layer includes three consecutive Emotion Discrimination Blocks 1 - EmotionBlock V1, four Emotion Discrimination Blocks 2 - EmotionBlock V2, six Emotion Discrimination Blocks 3 - EmotionBlock V3, three Emotion Discrimination Blocks 4 - EmotionBlock V4, and an attention mechanism module CBAM. It also includes a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. The output of the max - pooling layer Maxpooling is input into the first Emotion Discrimination Block 1 - EmotionBlock V1. The output of the max - pooling layer Maxpooling and the output of the first Emotion Discrimination Block 1 - EmotionBlock V1 are added together to obtain a first addition result. The first addition result is input into the second Emotion Discrimination Block 1 - EmotionBlock V1. The first addition result and the output of the second Emotion Discrimination Block 1 - EmotionBlock V1 are added together to obtain a second addition result. The second addition result is input into the third Emotion Discrimination Block 1 - EmotionBlock V1. The second addition result and the output of the third Emotion Discrimination Block 1 - EmotionBlock V1 are added together to obtain a third addition result. The third addition result is input into the first Emotion Discrimination Block 2 - EmotionBlock V2. The third addition result is processed by the second convolutional layer and then added to the output of the first Emotion Discrimination Block 2 - EmotionBlock V2 to obtain a fourth addition result. The fourth addition result is input into the second Emotion Discrimination Block 2 - EmotionBlock V2. The fourth addition result and the output of the second Emotion Discrimination Block 2 - EmotionBlock V2 are added together to obtain a fifth addition result. The fifth addition result is input into the third Emotion Discrimination Block 2 - EmotionBlock V2. The fifth addition result and the output of the third Emotion Discrimination Block 2 - EmotionBlock V2 are added together to obtain a sixth addition result. The sixth addition result is input into the fourth Emotion Discrimination Block 2 - EmotionBlock V2. The sixth addition result and the output of the fourth Emotion Discrimination Block 2 - EmotionBlock V2 are added together to obtain a seventh addition result. The seventh addition result is input into the first Emotion Discrimination Block 3 - EmotionBlock V3. The seventh addition result is processed by the third convolutional layer and then added to the output of the first Emotion Discrimination Block 3 - EmotionBlock V3 to obtain an eighth addition result. The eighth addition result is input into the second Emotion Discrimination Block 3 - EmotionBlock V3. The eighth addition result and the output of the second Emotion Discrimination Block 3 - EmotionBlock V3 are added together to obtain a ninth addition result. The ninth addition result is input into the third Emotion Discrimination Block 3 - EmotionBlock V3. The ninth addition result and the output of the third Emotion Discrimination Block 3 - EmotionBlock V3 are added together to obtain a tenth addition result.The tenth addition result is input into the fourth Emotion Discrimination Block III (EmotionBlock V3). After adding the tenth addition result and the output of the fourth Emotion Discrimination Block III (EmotionBlock V3), the eleventh addition result is obtained. The eleventh addition result is input into the fifth Emotion Discrimination Block III (EmotionBlock V3). After adding the eleventh addition result and the output of the fifth Emotion Discrimination Block III (EmotionBlock V3), the twelfth addition result is obtained. The twelfth addition result is input into the sixth Emotion Discrimination Block III (EmotionBlock V3). After adding the twelfth addition result and the output of the sixth Emotion Discrimination Block III (EmotionBlock V3), the thirteenth addition result is obtained. The thirteenth addition result is processed by the fourth convolutional layer and then added to the output of the first Emotion Discrimination Block IV (EmotionBlock V4) to obtain the fourteenth addition result. The fourteenth addition result is input into the second Emotion Discrimination Block IV (EmotionBlock V4). After adding the fourteenth addition result and the output of the second Emotion Discrimination Block IV (EmotionBlock V4), the fifteenth addition result is obtained. The fifteenth addition result is input into the third Emotion Discrimination Block IV (EmotionBlock V4). After adding the fifteenth addition result and the output of the third Emotion Discrimination Block IV (EmotionBlock V4), it is then input into the Convolutional Block Attention Module (CBAM); The output layer includes an average pooling layer (Avgpooling) and a fully connected layer (FC) connected in sequence, which is responsible for converting the above-mentioned feature map into the final four expression prediction results. These layers are responsible for extracting expression features and information from the input face images. The Convolutional Block Attention Module (CBAM) is placed at the end to enhance the neural network's attention to important feature channels and key regions of expressions; and between modules, the input features are added to the output of the convolutional layer through the introduction of a residual structure for feature fusion.,

[0080] The EmotionBlock is responsible for extracting emotion features and information from the input face images. The CBAM attention mechanism module is placed at the end to enhance the neural network's attention to important feature channels and key regions of expressions, while suppressing unimportant features. Moreover, between modules, the input features are added to the output of the convolutional layer through the introduction of a residual structure for feature fusion. The output layer is responsible for converting the above-mentioned feature maps into a feature vector by passing them through the Avgpooling average pooling layer, which compresses the spatial information. This vector is then fed into a fully connected layer, which outputs a vector with the dimension of the number of emotion categories, thereby obtaining the final prediction results for the four expressions of students. The student emotion recognition neural network model SFaceNet is optimized and improved for its specific scenario requirements to recognize the four emotions of students. The student emotion recognition neural network model SFaceNet can accurately classify the expressions of the faces recognized by the RetinaFace neural network face detection model in the class.

[0081] As Figure 4 shown, the EmotionBlock V1, EmotionBlock V2, EmotionBlock V3, and EmotionBlock V4 of the expression discriminant block include a second Batch Norm layer, a second ReLu activation function layer, a fifth convolutional layer, a third Batch Norm layer, a third ReLu activation function layer, a sixth convolutional layer, a Contextual Transformer (CoT) module, a fourth Batch Norm layer, a fourth ReLu activation function layer, and a seventh convolutional layer, which are connected in sequence.

[0082] The Emotionblock of the expression discriminant block refines features and strengthens expression-related information. The size of the fifth convolutional layer of the EmotionBlock V1 is 1×1, the stride is 1, and the padding value is 1. After processing, the image size becomes 64×56×56; the size of the sixth convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size is 64×56×56; the size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 64×56×56.

[0083] The output of the max pooling layer Maxpooling will pass through three identical emotion discrimination blocks - EmotionBlock V1 in total, and the final output will be fed to four emotion discrimination blocks - EmotionBlock V2 in sequence. Among them, the size of the fifth convolutional layer of the first emotion discrimination block - EmotionBlock V2 is 1×1, the stride is 1, and the padding value is 1. After processing, the image size remains 64×56×56. The size of the sixth convolutional layer is 3×3, the stride is 2, and the padding value is 1. After processing, the image size becomes 128×28×28. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 128×28×28. For the remaining second, third, and fourth emotion discrimination blocks - EmotionBlock V2, the size of the fifth convolutional layer is 1×1, the stride is 1, and the padding value is 1. The size of the sixth convolutional layer is 3×3, the stride is 1, and the padding value is 1. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 128×28×28; then the result is input to six emotion discrimination blocks - EmotionBlock V3. Among them, the size of the fifth convolutional layer of the first emotion discrimination block - EmotionBlock V3 is 1×1, the stride is 1, and the padding value is 1. After processing, the image size remains 128×28×28. The size of the sixth convolutional layer is 3×3, the stride is 2, and the padding value is 1. After processing, the image size becomes 256×14×14. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 256×14×14. For the remaining second, third, fourth, fifth, and sixth discrimination blocks - EmotionBlock V3, the size of the fifth convolutional layer is 1×1, the stride is 1, and the padding value is 1. The size of the sixth convolutional layer is 3×3, the stride is 1, and the padding value is 1. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 256×14×14; then the result is input to the last four emotion discrimination blocks - EmotionBlock V4. Among them, the size of the fifth convolutional layer of the first emotion discrimination block - EmotionBlock V4 is 1×1, the stride is 1, and the padding value is 1. After processing, the image size remains 256×14×14. The size of the sixth convolutional layer is 3×3, the stride is 2, and the padding value is 1. After processing, the image size becomes 512×7×7. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 512×7×7. For the remaining second and third emotion discrimination blocks - EmotionBlock V4, the size of the fifth convolutional layer is 1×1, the stride is 1, and the padding value is 1. The size of the sixth convolutional layer is 3×3, the stride is 1, and the padding value is 1. The size of the seventh convolutional layer is 3×3, the stride is 1, and the padding value is 1. After processing, the image size remains 512×7×7.Finally, the 512×7×7 image output is successively passed through an attention mechanism module CBAM and an average pooling layer Avgpooling, adjusted to 512×1×1, and finally passed through a fully connected layer FC to obtain a 4-dimensional output, which represents the predicted probabilities of four expressions respectively.

[0084] Then, data augmentation is performed on the initial dataset. First, the function RandomResizedCrop(224) is used to randomly crop the image and resize it to a unified size of 224×224 to increase data diversity and enable the model to better generalize to different image sizes and scenarios. Then, the function RandomHorizontalFlip() is used to randomly flip the image horizontally to increase data symmetry and help the model learn expression features that are invariant to left-right flipping. At the same time, the function Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]) is used to standardize the image, thereby accelerating the convergence of the model and improving the stability of training. Finally, the initial dataset after data augmentation is obtained, and the student expression recognition neural network model SFaceNet is trained.

[0085] When training the student expression recognition neural network model SFaceNet, the loss function used is the CrossEntropyLoss cross-entropy loss function to minimize the difference between the class probabilities output by the model and the true labels, thereby optimizing the performance of the model. The optimization method uses the Adam optimizer with a learning rate set to 0.0001 to handle sparse gradients or different update speeds of different parameters. The number of iterations Epoch is 300, and the batch size Batch_size is 16. The finally obtained SFaceNet model shows good convergence during training and achieves an ideal classification accuracy.

[0086] Then, obtain the 40-minute classroom lecture videos of 80 students to be recognized through a camera, extract the video frames to be recognized, and then perform face recognition. Furthermore, use the trained student emotion recognition neural network model SFaceNet to recognize the occurrence frequencies of different emotion states of each student. For each video frame to be recognized, perform face recognition on each video frame to be recognized through the face detection model RetinaFace, obtain the image of the detection box with the facial features of each student to be recognized, and then input the image of the detection box with the facial features of each student to be recognized into the trained student emotion recognition neural network model SFaceNet for processing to obtain the emotion state of each student in the current video frame. Finally, based on the images of the detection boxes with the facial features of each student to be recognized, recognize the occurrence frequencies of different emotion states of each student during the shooting period of the student classroom lecture video to be recognized. After testing, in the 2 minutes before class, 58% of the 80 students had an expression of "listening", 36% of the students had an expression of "excited", only 4% of the students had an expression of "confused", and 2% of the students had an expression of "rejected".

[0087] At the same time, identify the number of times of fatigue of each student through face key point localization. For each image of the detection box with the facial features of each student to be recognized, first process it through the 468-face key point localization method based on machine learning Mediapipe to obtain the 468 key point information of the face of each student in the graph. For each student, take 6 key points from each of the three parts of the left eye, right eye, and mouth of the student's face, a total of 18 key points. Obtain the eye aspect ratio EAR (Eye Aspect Ratio) based on the 6 key points of the left and right eyes of the student's face, and then judge the closed-eye state of the student; obtain the mouth opening degree MAR (Mouth Aspect Ratio) based on the 6 key points of the student's face's mouth, and then judge the yawning state of the student. Finally, based on the images of the detection boxes with the facial features of each student to be recognized, recognize the number of times of closing eyes and yawning of each student during the shooting period of the student classroom lecture video to be recognized as the number of times of fatigue of each student.

[0088] Perform threshold judgment on the two parameters of the eye aspect ratio EAR and the mouth opening degree MAR, and then judge the eye and mouth states of the student, specifically as follows:

[0089] Closed eyes: 0 ≤ EAR ≤ 0.15

[0090] Normal eye opening: 0.15 < EAR

[0091] Normal mouth state: 0 ≤ MAR < 0.25

[0092] Mouth slightly open: 0.25 ≤ MAR < 0.35

[0093] Yawning: 0.35 ≤ MAR

[0094] As Figure 5 shown, the present invention uses two indicators for characterization: the number of times a student closes their eyes and the number of times a student yawns during a class. By using the RetinaFace algorithm of the face detection model to detect the face, and combining the face key point information obtained by the 468-face key point localization method based on machine learning Mediapipe, the eye aspect ratio EAR value reflecting the eye state and the mouth opening degree MAR value reflecting the mouth state are calculated. The calculation of the eye aspect ratio EAR value is based on several key points of the eye, and the opening and closing state of the eye is evaluated by measuring the ratio of the eye width to the eye height. When the eye aspect ratio EAR value is lower than the set threshold, it can be determined that the student closes their eyes. Similarly, the calculation of the mouth opening degree MAR value is based on several key points of the mouth, and the opening and closing state of the mouth is evaluated by measuring the ratio of the mouth width to the mouth height. When the mouth opening degree MAR value exceeds the set threshold, it can be determined that the student yawns.

[0095] In order to accurately evaluate the fatigue state of students, the present invention measures the different class states of students, determines the thresholds of the eye aspect ratio EAR and the mouth opening degree MAR to distinguish different eye and mouth states. As shown in Table 2 below, the threshold of the eye aspect ratio EAR value is set to distinguish the normal open-eye and closed-eye states.

[0096] Table 2 Relationship between eye state and eye aspect ratio EAR value

[0097]

[0098] Similarly, the threshold of the mouth opening degree MAR value in the present invention is used to distinguish the "normal mouth state", "mouth slightly open" and "yawning" states, as shown in Table 3 below.

[0099] Table 3 Relationship between mouth state and mouth opening degree MAR value

[0100]

[0101] In summary, the present invention can understand the fatigue state of students in real time by monitoring the eye closing and yawning situations of students in class and combining the set thresholds for state classification. Such an evaluation can not only help teachers detect the fatigue of students in time, but also provide a scientific basis for adjusting teaching strategies, thereby better promoting the learning effect and physical and mental health of students. The fatigue detection effect is that 2 minutes before the class, only 0.5% of the 80 students "yawned" and "closed their eyes", and the rest of the students remained focused on listening to the class.

[0102] Meanwhile, the number of times of abnormal head postures of each student is recognized through face key point positioning. For each video frame to be recognized and for each image of a detection box with the facial features of each student to be recognized, first, it is processed by the 468-face key point positioning method based on the machine learning Mediapipe to obtain the two-dimensional coordinates of 468 key points of each student's face in the graph. Then, normalization processing of the two-dimensional coordinates is performed to avoid errors caused by differences in the sizes of different input images, making subsequent calculations more stable and general. Then, the two-dimensional coordinates of each normalized key point are mapped into the three-dimensional world coordinate system. For each student, the pose estimation algorithm PnP is used to match the three-dimensional coordinates and two-dimensional coordinates of each key point of the student's face in the three-dimensional world coordinate system, so as to obtain the rotation and displacement of the camera, and finally obtain the three-degree-of-freedom information of the student's head, including the rotation Yaw around the vertical axis (z-axis), the rotation Pitch around the horizontal axis (y-axis), and the rotation Roll around the longitudinal axis (x-axis). Among them, Pitch judges looking up and looking down. When looking up, Pitch > 0, and when looking down, Pitch < 0; Yaw judges turning left and right. When turning the head to the right, Yaw > 0, and when turning the head to the left, Yaw < 0; Roll judges tilting the head left and right. When tilting the head to the right, roll > 0, and when tilting the head to the left, roll < 0. Through the combination of these three Euler angles, the present invention can describe any posture of the head in three-dimensional space, and further provide accurate human body posture information for applications such as teaching interaction and virtual reality, which is used for the establishment of a subsequent student attention evaluation system. The bowing state of the student is judged according to the three-degree-of-freedom information of the student's head. Finally, the number of times each student bows his head during the shooting period of the video of the student's classroom lecture to be recognized is recognized as the number of times of abnormal head postures of each student according to the image of each detection box with the facial features of each student to be recognized.

[0103] Judge whether the three values of the rotation Yaw around the vertical axis (z-axis), the rotation Pitch around the horizontal axis (y-axis), and the rotation Roll around the longitudinal axis (x-axis) in the three-degree-of-freedom information of the student's head are within the set interval, and then judge whether the student bows his head, specifically as follows:

[0104] Bowing state: Pitch < -5, -15 < Roll < 15

[0105] After testing, it can be obtained from the data that 2 minutes before class, only 2% of the 80 students bowed their heads, and the number of times a single student in this class bowed his head did not exceed 2 times.

[0106] Such as Figure 6As shown, for each student, the occurrence frequency of the student's different emotional states, fatigue level, and the number of abnormal head postures are input into the pre-trained extreme learning machine model ELM, that is, the obtained seven-dimensional eigenvalue is used as a 1×7 input vector ( x 1 : the number of times of lowering the head; x 2 : the number of times of yawning; x 3 : the number of times of closing the eyes; x 4 : the number of times of listening expressions appear; x 5 : the number of times of confused expressions appear; x 6 : the number of times of resistant expressions appear; x 7 : the number of times of excited expressions appear) are input into the pre-trained extreme learning machine model ELM for the attention assessment of students. Finally, an attention assessment parameter between 0 and 1 is obtained. This value reflects different intervals and represents different attention states and classroom effects of students during the class. Then, the average value of the attention assessment parameters of each student at each moment is used as the average attention assessment parameter of the class at the current moment. Introducing the time dimension for time series analysis, the preset time interval is 1 min, and the video records once every 15 s to obtain the average attention assessment parameter of the class at a moment. The average value of the average attention assessment parameters of the class at several moments within each preset time interval is used as the attention assessment parameter of the class unit cycle to judge the attention state of the students in the class. The specific mapping relationship is shown in Table 4.

[0107] Table 4 Relationship table between attention assessment parameter Focus and attention state and classroom effect

[0108]

[0109] The present invention also introduces the concept of month-on-month increment ▽t to analyze more carefully the changes in students' attention in the class, helping educators identify the fluctuation trend of students' attention from the time dimension, and then providing more scientific data support for teaching intervention. By paying attention to the changes in the class in adjacent time periods, the interference of single-moment data can be avoided, so as to more accurately evaluate and improve the classroom teaching effect, as follows:

[0110] ▽t =(Focus t / Focus t-1 -1)×100%, t =2, 3,..., N

[0111] Focus0 =0

[0112] Among them, Focus t and Focus t-1 respectively represent t minutes and t the classroom unit cycle attention evaluation parameters at -1 minute. N represents the total duration of the shooting time period of the classroom video of the student to be recognized.

[0113] Finally, the change of the classroom student's attention in the time series state can be calculated. Taking a 40 - minute university course participated by 80 students as an example, when evaluating the students' attention, in order to ensure the simplicity and intuitiveness of the interface, the visualization window on the computer side only shows the face frame of each recognized student and five key facial feature points. Other attention - related behaviors such as the number of times of lowering the head, yawning, and closing eyes are all recorded and processed in real - time through an automated process in the system background. Through these data, the system can provide real - time feedback on the attention changes of each student in the classroom, providing decision - making support for teachers. As Figure 7 shown, it shows the change trend of the attention of 80 students within 40 minutes and the corresponding month - on - month increment, used to analyze and measure the difference in the attention change of the class students between two consecutive time periods; at the same time, it can be obtained that the student attention evaluation parameters obtained by the present invention generally conform to the 10 - minute rule, that is, the students' attention state is relatively high at the beginning of the class, but usually can only be maintained for about 10 - 15 minutes, and then it will start to gradually decline. This phenomenon indicates that the students' attention is most concentrated at the beginning of the class, probably due to the curiosity and thirst for knowledge about the new course content. However, as time goes by, the students' attention begins to gradually disperse, which may be due to fatigue caused by long - term concentration or loss of interest in the classroom content. When the class is about 30 minutes in progress, the students' attention evaluation parameters reach the lowest point. And when approaching the end of the class, the students' attention will return to the class again, and the facial expressions are also relatively positive, so the attention evaluation parameters will increase again.

[0114] In order to further verify the accuracy of the established multi - feature student attention evaluation model, the present invention uses the method of manually supervised evaluation parameters to score the classroom effect, that is, three students are invited to watch different time - period segments of the same classroom video, and score them between 0 and 1 according to the criteria in Table 4, and then compare them with the evaluation parameters of the model in this article. The comparison results show that the accuracy of this model is relatively high and basically coincides with the manual evaluation, indicating that the method of the present invention has high accuracy and reliability.

[0115] The classroom student attention assessment of the present invention includes a face detection unit, an expression recognition unit, a fatigue detection unit, a head pose estimation unit, and an attention assessment unit. The face detection unit processes the video frames of the student classroom lecture videos to be recognized obtained through a camera, uses the face detection model RetinaFace for face recognition, and then outputs an image with detection frames of the facial features of each student as the face detection result, which can achieve accurate detection of the student face area and ensure effective recognition of small-scale, low-pixel, and multi-occluded student faces in complex classroom scenarios; the expression recognition unit constructs and trains a student expression recognition neural network model SFaceNet containing four expression discrimination blocks, so as to use the trained student expression recognition neural network model SFaceNet to recognize the different emotional states of each student as the expression recognition result, and then obtain the occurrence frequency of the different emotional states of each student; the fatigue detection unit identifies the number of times each student closes their eyes and yawns as the number of times of fatigue through face key point positioning; the head pose estimation unit identifies the number of times each student lowers their head as the number of times of abnormal head pose through face key point positioning; the attention assessment unit inputs the occurrence frequency of the different emotional states of the students, the fatigue level, and the number of times of abnormal head pose into the pre-trained extreme learning machine model ELM. After processing, it outputs the attention assessment parameters of the students, evaluates the attention state of the students in the classroom according to the attention assessment parameters of each student, and displays it on the monitor.

[0116] It should be noted that the parts not elaborated in detail in the embodiments of the present invention belong to the category of common knowledge or publicly disclosed prior art for those skilled in the art. These contents are regarded as general knowledge in the industry.

[0117] In addition, for those of ordinary skill in the art, the steps in the embodiments described in the present invention, whether all or part of them, can be implemented by programming instructions for corresponding hardware devices. The corresponding control programs can be stored in various computer-readable media, such as but not limited to read-only memory, hard disk, or optical disc, etc.

[0118] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes prior art known to those skilled in the art.

Claims

1. A method for evaluating students' attention in classroom based on multi-source feature fusion, characterized in that: include: Step S1: Obtain several videos of students listening to lectures in class through a camera and extract initial video frames, perform face detection preprocessing on each initial video frame, and then identify different emotional states of the face by locating facial key points and adding expression labels, thereby obtaining several facial expression images and constructing a facial expression image dataset; Step S2: construct a student expression recognition neural network model SFaceNet containing four expression discrimination blocks, and train the student expression recognition neural network model SFaceNet using a facial expression image dataset; Step S3: Obtain the classroom video of the student to be identified through the camera and extract the video frame to be identified, then perform face recognition, and then identify the different emotional states of each student through the trained student expression recognition neural network model SFaceNet, obtain the frequency of occurrence of different emotional states of each student, and identify the fatigue level and abnormal head posture of each student through facial key point positioning; Step S4: for each student, the occurrence frequency of different emotional states, fatigue level and number of abnormal head postures of the student are input into the pre-trained extreme learning machine model ELM, and after processing, the student's attention evaluation parameter is output. The attention state of the students in the class is evaluated and displayed according to the attention evaluation parameter of each student; In the step S2, the student expression recognition neural network model SFaceNet includes an input layer, a feature extraction and fusion layer, and an output layer connected in sequence, the input layer includes a first convolutional layer, a first batch normalization layer BatchNorm, a first activation function layer ReLu, and a maximum pooling layer Maxpooling connected in sequence, the feature extraction and fusion layer includes three expression discrimination blocks EmotionBlock V1, four expression discrimination blocks EmotionBlock V2, six expression discrimination blocks EmotionBlock V3, three expression discrimination blocks EmotionBlock V4, and an attention mechanism module CBAM connected in sequence, and also includes a second convolutional layer, a third convolutional layer, and a fourth convolutional layer, the output of the maximum pooling layer Maxpooling is then input into the first expression discrimination block EmotionBlock V1, the maximum pooling layer Maxpooling and the output of the first expression discrimination block EmotionBlock V1 are added to obtain a first addition result, the first addition result is input into the second expression discrimination block EmotionBlock V1, the first addition result and the second expression discrimination block EmotionBlock V1 to obtain a second addition result, and the second addition result is input into the third expression discrimination block one EmotionBlock V1. The second addition result and the output of the third expression discrimination block one EmotionBlock V1 are added to obtain a third addition result, and the third addition result is input into the first expression discrimination block two EmotionBlock V2. The third addition result is processed by the second convolutional layer and added to the output of the first expression discrimination block two EmotionBlock V2 to obtain a fourth addition result. The fourth addition result is input into the second expression discrimination block two EmotionBlock V2. The fourth addition result and the output of the second expression discrimination block two EmotionBlock V2 are added to obtain a fifth addition result. The fifth addition result is input into the third expression discrimination block two EmotionBlock V2. The fifth addition result and the output of the third expression discrimination block two EmotionBlock V2 are added to obtain a sixth addition result. The sixth addition result is input into the fourth expression discrimination block two EmotionBlock V2. The sixth addition result and the fourth expression discrimination block two EmotionBlock V2 are added to obtain a sixth addition result. The outputs of V2 are added to obtain a seventh addition result, which is input into the first expression discrimination block three EmotionBlock V3. The seventh addition result is processed by the third convolutional layer and added to the output of the first expression discrimination block three EmotionBlock V3 to obtain an eighth addition result, which is input into the second expression discrimination block three EmotionBlock V3.The eighth addition result is added to the output of the second expression discrimination block three EmotionBlock V3 to obtain a ninth addition result, and the ninth addition result is input into the third expression discrimination block three EmotionBlock V3. The ninth addition result is added to the output of the third expression discrimination block three EmotionBlock V3 to obtain a tenth addition result, and the tenth addition result is input into the fourth expression discrimination block three EmotionBlock V3. The tenth addition result is added to the output of the fourth expression discrimination block three EmotionBlock V3 to obtain an eleventh addition result, and the eleventh addition result is input into the fifth expression discrimination block three EmotionBlock V3. The eleventh addition result is added to the output of the fifth expression discrimination block three EmotionBlock V3 to obtain a twelfth addition result, and the twelfth addition result is input into the sixth expression discrimination block three EmotionBlock V3. The twelfth addition result and the sixth expression discrimination block three EmotionBlock V3 are added together. The output of V3 is added to obtain the thirteenth addition result, the thirteenth addition result is processed by the fourth convolutional layer and added to the output of the first expression discrimination block four EmotionBlock V4 to obtain the fourteenth addition result, the fourteenth addition result is input to the second expression discrimination block four EmotionBlock V4, the fourteenth addition result and the output of the second expression discrimination block four EmotionBlock V4 are added to obtain the fifteenth addition result, the fifteenth addition result is input to the third expression discrimination block four EmotionBlock V4, the fifteenth addition result and the output of the third expression discrimination block four EmotionBlock V4 are added and then input to the attention mechanism module CBAM; the output layer includes the average pooling layer Avgpooling and the fully connected layer FC connected in sequence. , 2. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 1, characterized in that: In the step S1, each initial video frame is subjected to face recognition through the face detection model RetinaFace to obtain an image of a detection frame with facial features of each student, and then data enhancement processing is performed, including random cropping, adjustment to a uniform size, random horizontal flipping and image normalization, so as to obtain a preprocessed video frame.

3. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 2 is characterized in that: In the step S1, each pre-processed video frame is identified by the 468 facial key point positioning method based on machine learning Mediapipe and the Russell emotion model to identify four emotional states of the face including rejection expression, listening expression, confusion expression and excitement expression. For the face of the student in each detection frame of the pre-processed video frame, the four emotional states of the face are as follows: a) Rejection expression: Determine the degree of mouth tilt. When the angle between the line connecting the 62nd key point, the 292nd key point and the 17th key point and the horizontal line is negative, the degree of mouth tilt is that the corners of the mouth on both sides move downward, and the current emotional state of the face is determined to be a rejection expression. b) Listen to the expression: To judge the degree of mouth tilt, when the line connecting the 62nd key point, the 292nd key point and the 17th key point in the face is parallel to the horizontal line, the degree of mouth tilt is not obvious; Determine the degree of eye openness, obtain the eye height of the left eye in the eye feature by subtracting the ordinates of the 470th key point and the 472nd key point in the face, obtain the eye width of the left eye in the eye feature by subtracting the abscissas of the 33rd key point and the 133rd key point in the face, obtain the eye height of the right eye in the eye feature by subtracting the ordinates of the 475th key point and the 477th key point in the face, obtain the eye width of the right eye in the eye feature by subtracting the abscissas of the 362nd key point and the 263rd key point in the face, and if the ratio of the eye height to the eye width of the left eye and the ratio of the eye height to the eye width of the right eye are both greater than the ratio threshold, the degree of eye openness of the face is open; To judge the degree of eyebrow curvature, if the angles between the 70th key point and the 65th key point and the 105th key point in the face are greater than the angle threshold, and the angles between the 300th key point and the 295th key point and the 334th key point in the face are greater than the angle threshold, the degree of eyebrow curvature of the face is stretched; When the inclination of the mouth of the human face is not obvious, the opening degree of the eyes of the human face is open, and the curvature degree of the eyebrows of the human face is stretched, it is determined that the current emotional state of the human face is a listening expression; c) Confused expression: To judge the degree of eyebrow curvature, if the angles between the 70th key point and the 65th key point and the 105th key point in the face are less than the angle threshold, and the angles between the 300th key point and the 295th key point and the 334th key point in the face are less than the angle threshold, the degree of eyebrow curvature of the face is frowning; When the inclination of the mouth of the face is that the corners of the mouth move downward, and the curvature of the eyebrows of the face is tightly wrinkled, it is determined that the current emotional state of the face is a confused expression; d) Excited expression: Determine the degree of mouth tilt. When the angle between the line connecting the 62nd key point, the 292nd key point and the 17th key point and the horizontal line is a positive number, the degree of mouth tilt is that the corners of the mouth on both sides are upturned. When the inclination degree of the mouth of the human face is that the corners of the mouth on both sides are upturned, the opening degree of the eyes of the human face is open, the curvature degree of the eyebrows of the human face is stretched, and the Euclidean distance between the 12th key point and the 14th key point of the mouth of the human face is greater than the preset distance threshold, it is determined that the current emotional state of the human face is an excited expression; An expression label of an emotional state is added to the face of each student in the detection frame.

4. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 1, characterized in that: The expression discrimination block EmotionBlock V1, the expression discrimination block EmotionBlock V2, the expression discrimination block EmotionBlock V3 and the expression discrimination block EmotionBlock V4 include a second batch normalization layer Batch Norm, a second activation function layer ReLu, a fifth convolutional layer, a third batch normalization layer Batch Norm, a third activation function layer ReLu, a sixth convolutional layer, a context converter module CoT, a fourth batch normalization layer Batch Norm, a fourth activation function layer ReLu and a seventh convolutional layer connected in sequence.

5. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 1, characterized in that: In the step S3, for each video frame to be identified, each video frame to be identified is subjected to face recognition through the face detection model RetinaFace to obtain an image of the detection frame with the facial features of each student to be identified, and then the image of the detection frame with the facial features of each student to be identified is input into the trained student expression recognition neural network model SFaceNet for processing to obtain the emotional state of each student in the current video frame, and finally the frequency of occurrence of different emotional states of each student during the shooting time period of the classroom lecture video of the student to be identified is identified according to the images of the detection frame with the facial features of each student to be identified.

6. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 5, characterized in that: In the step S3, for each image with a detection frame of each student's facial features to be identified, it is first processed by a 468 facial key point positioning method based on machine learning Mediapipe to obtain 468 key point information of each student's face in the graphic. For each student, 6 key points are taken from each of the three parts of the student's face, namely the left eye, right eye and mouth, for a total of 18 key points. The eye aspect ratio EAR is obtained based on the 6 key points of the left eye and the right eye of the student's face, and then the student's eye closing state is judged; the mouth opening degree MAR is obtained based on the 6 key points of the student's mouth, and then the student's yawning state is judged. Finally, the number of times each student closes his eyes and yawns during the shooting period of the classroom lecture video of the student to be identified is identified according to each image with a detection frame of each student's facial features to be identified as the number of times each student is tired.

7. The method for evaluating students' attention in class based on multi-source feature fusion according to claim 5, characterized in that: In the step S3, for each video frame to be identified and for each image of a detection frame with facial features of each student to be identified, firstly, the image is processed by the 468 facial key point positioning method based on machine learning Mediapipe to obtain the two-dimensional coordinates of 468 key points of the face of each student in the figure, and then the two-dimensional coordinates are normalized, and then the two-dimensional coordinates of each normalized key point are mapped to the three-dimensional world coordinate system. For each student, the three-dimensional coordinates and the two-dimensional coordinates of each key point of the student's face in the three-dimensional world coordinate system are matched by the pose estimation algorithm PnP, so as to obtain the rotation and displacement of the camera, and finally obtain the three-degree-of-freedom information of the student's head, judge the student's head-lowering state according to the three-degree-of-freedom information of the student's head, and finally identify the number of times each student lowers his head during the shooting time period of the classroom listening video of the student to be identified according to each image of the detection frame with facial features of each student to be identified as the number of abnormal head postures of each student.

8. The method for evaluating classroom student attention based on multi-source feature fusion according to claim 3 is characterized by: In step S4, the average value of the attention evaluation parameters of each student at each moment is used as the average attention evaluation parameter of the class at the current moment, and the average value of the average attention evaluation parameters of the class at several moments within each preset time interval is used as the class unit period attention evaluation parameter, so as to judge the attention state of the students in the class, as follows: When 0≤Focus t When it is less than 0.4, the students in the class have a very poor attention state. t It is the attention assessment parameter of the class unit cycle at the current preset time interval; When 0.4≤Focus t When it is less than 0.6, the students in the class have poor attention status; When 0.6≤Focus t When it is less than 0.8, the students’ attention in the class is average; When 0.8≤Focus t When it is less than 1.0, the students in the class have excellent attention.

9. A classroom student attention assessment system applicable to the method according to any one of claims 1 to 8, characterized in that: include: The face detection unit uses the face detection model RetinaFace to perform face recognition on the video frames of the classroom lecture video of the students to be identified obtained through the camera, and then outputs an image of the detection frame with the facial features of each student as the face detection result; The expression recognition unit constructs and trains the student expression recognition neural network model SFaceNet, which contains four expression discrimination blocks. The trained student expression recognition neural network model SFaceNet is used to identify the different emotional states of each student as the expression recognition result, and then the occurrence frequency of each student's different emotional states is obtained. The fatigue detection unit locates the key points of the face and identifies the number of times each student closes his eyes and yawns as the number of times of fatigue; The head posture estimation unit identifies the number of times each student lowers his head as the number of abnormal head postures by locating facial key points; The attention assessment unit inputs the frequency of occurrence of students' different emotional states, fatigue levels, and the number of abnormal head postures into the pre-trained extreme learning machine model ELM. After processing, it outputs the students' attention assessment parameters. Based on the attention assessment parameters of each student, the attention status of the students in the class is evaluated and displayed on the monitor.

Citation Information

Patent Citations

  • Teaching assessment method

    CN109978732A

  • Behavior action and expression multi-modal learning state evaluation method and system

    CN117315746A

  • Student attention assessment method and system based on big data analysis

    CN119314100A