Student classroom interest intelligent evaluation system and method based on multi-visual perception
By using multi-visual perception technology to identify students' behavior, attention, and emotions, and combining this with the analytic hierarchy process to generate interest scores, this approach solves the problems of high cost and interference in existing technologies, achieving non-intrusive automated classroom interest assessment that is suitable for large-scale classroom teaching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-03-13
AI Technical Summary
Existing automated methods for assessing student classroom interest are costly, may be disruptive to students, and face challenges in data fusion and interest modeling, lacking effective analysis and measurement methods suitable for classroom teaching.
A student classroom interest intelligent assessment system based on multi-visual perception is adopted, including a behavior detection module, an attention estimation module, an emotion recognition module, and an identity matching module. The system identifies students' behavior, attention, and emotions through YOLOv11, 6DRepNet, EmoFAN, and SphereFace models, and generates interest scores by combining the analytic hierarchy process (AHP).
It enables non-intrusive, contactless student classroom interest assessment, provides automated, real-time interest assessment support, improves the accuracy and applicability of the assessment, and is suitable for large-scale classroom environments.
Smart Images

Figure CN121661669A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to, but is not limited to, the field of smart education technology, and particularly relates to a student classroom interest intelligent assessment system and method based on multi-visual perception. Background Technology
[0002] Interest is a crucial factor influencing students' learning motivation, participation, and learning outcomes. Its importance has prompted researchers to delve into its nature and impact. In psychological and educational research, learning interest is often described as a psychological response closely related to emotions, reflecting students' preferences, enthusiasm, and satisfaction with learning content. For example, some researchers believe that interest is a learning experience containing emotional components, often accompanied by feelings of pleasure and satisfaction. Schiefele argues that learning interest involves strong emotional experiences, especially positive emotions related to the learning content. His research found that the stronger the emotional connection between students and the learning content, the better their learning outcomes and long-term memory performance. In the four-stage model of interest development, each stage is described as an emotional process, from initial situational interest to personal interest, with the influence of emotion always present.
[0003] It is worth noting that interest is not simply an emotional experience, but the result of multiple factors working together. Therefore, modern research tends to define interest as a multi-layered psychological state, encompassing not only emotional components but also behavioral and engagement-related manifestations. Renninger et al. argue that interest plays a crucial role in learning relationships, significantly influencing learning motivation and participation behaviors, such as exploratory behavior and practical activities. Kahu et al.'s research shows a positive correlation between student interest and learning participation, a relationship particularly evident in the learning environment of new students. Students' interest in course content significantly enhances their classroom participation and learning behavior. For example, when students are interested in a subject, they are more likely to proactively complete assignments, participate in class discussions, and engage in extracurricular learning. Learning interest can also reduce negative learning behaviors, such as laziness, lack of initiative, and even dropping out. Compared to superficial learning, interest often drives students to engage in deeper levels of learning, such as collaborative learning and problem-solving. Even when faced with formulaic and abstract mathematical problems, interest helps students maintain high learning motivation, keeping them in an active learning state. In the field of science education, students driven by interest are often more willing to conduct experimental investigations, collect and analyze data, and try to apply knowledge to practical problems.
[0004] Interest can also manifest as a special relationship between an individual and the learning content, often reflected in the concentration of attention. When students are interested in a particular topic, they naturally focus more attention on it, ignoring other irrelevant things. Some research suggests that interest involves multiple dimensions in the learning process, including motivation and cognition. These dimensions help students process information more efficiently, reduce their learning burden, and promote deep learning.
[0005] In the classroom setting, students' interest can be manifested in various external ways, such as classroom behavior, attention span, and facial expressions. These external cues transcend internal psychological states, providing educators with richer and more easily interpreted dimensions for assessing student engagement.
[0006] Students' learning interest can be measured in various ways, broadly categorized into traditional methods and data-driven methods. Currently, traditional methods, such as self-assessment and teacher observation, still dominate research on learning interest assessment. Self-assessment is a report-based approach from the student's own perspective, using questionnaires to evaluate students' genuine interest. Many studies consider interest as a positive emotion; therefore, questionnaires include numerous questions inquiring about students' positive emotional experiences, such as "how much I like the task" and "the pleasure I feel during the task." Mazer developed the "Student Interest Scale" to assess university students' interest in a course. This questionnaire investigates two dimensions: emotional interest and cognitive interest, each with specific questions, such as "This course excites me" and "I can remember the course content." Some questionnaires also investigate students' perceived value of the course, their willingness to participate again, or further explore aspects such as attention and activity participation. Observational methods, through pre-defined observation indicators, allow teachers to systematically observe students' classroom performance and interpret their interest levels by recording the results. Tan et al. obtained student participation, attention, and interest performance through classroom observation, using these as indicators for assessing student interest. Traditional methods often rely on experience, have high human and material costs and low automation, making them difficult to apply widely in classroom teaching.
[0007] With the development of artificial intelligence and data analysis technologies, some studies have attempted to use computer vision, deep learning, and natural language processing to assess students' learning interests. For example, interest is considered a type of learning emotion, encompassing types such as happiness, fear, sadness, calmness, and depression. These methods extract students' emotional characteristics and then use machine learning models such as multilayer perceptrons or support vector machines (SVMs) to measure students' interest levels. Luo et al. collected multidimensional data, including students' learning emotions, cognitive attention, and mental activity, to assess their learning interests. However, existing methods still face many challenges in practice. On the one hand, most methods rely on single-dimensional data, making it difficult to comprehensively and accurately assess students' interest levels. On the other hand, some studies have attempted to use wearable sensing devices such as eye trackers, EEG headbands, ECG monitors, and smart bracelets to collect multimodal data for interest analysis. However, these devices are usually expensive, and their contact-based data collection methods may restrict students and interfere with normal classroom activities, limiting their applicability in actual teaching environments.
[0008] Based on the above analysis, the urgent technical problems that need to be solved in the existing technology are:
[0009] Current automation methods have significant limitations, such as the high cost of wearable devices, potential disruption to students, and challenges in data fusion and interest modeling across different technologies. Although student learning interest has received widespread attention in education, effective analytical and measurement methods applicable to classroom teaching still lack due to differing research perspectives. Summary of the Invention
[0010] To address the problems existing in the prior art, this invention provides a student classroom interest intelligent assessment system and method based on multi-visual perception.
[0011] This invention is implemented as follows: a student classroom interest intelligent assessment system based on multi-visual perception, characterized in that the method of the student classroom interest intelligent assessment system based on multi-visual perception specifically includes:
[0012] The behavior detection module uses the YOLOv11 model to identify key behaviors such as sitting upright, sitting sideways, leaning on a table, and standing.
[0013] The attention estimation module uses head posture estimation and other technologies to determine whether students are focused on the blackboard and classifies attention states into three types: focused, distracted and detached.
[0014] The emotion recognition module assesses students' emotional valence and arousal levels based on facial expressions.
[0015] The identity matching module matches each student's facial image with the source data collected before class;
[0016] The quantitative assessment module converts the extracted behavioral, attentional, and emotional information into numerical scores, and integrates these scores through a weighted summation to generate the final interest score.
[0017] Furthermore, the behavior detection module categorizes student behavior into four types: sitting upright, standing, turning to the side, and lying on the desk.
[0018] Furthermore, the attention estimation module operates as follows:
[0019] (1) Input the cropped single student image into the module. The module extracts the coordinates of the center point of the student image and converts them into spatial three-dimensional (3D) coordinates in the camera coordinate system. At the same time, the 6DRepNet model is used to estimate the student's head pose angle.
[0020] (2) Based on the above three-dimensional coordinates, the module determines the range of head posture angles related to the blackboard of interest;
[0021] (3) Compare the estimated head posture angle with the reference range to infer the student's attention state.
[0022] Furthermore, the process of extracting the center point coordinates of the student image and converting them into three-dimensional (3D) coordinates involves calculating the three-dimensional position of each student in the camera coordinate system using the center point of the student image, the shoulder width in pixels, the actual shoulder width, and camera intrinsic parameters. The specific steps are as follows:
[0023] (1) Obtain the camera intrinsic parameter K through camera calibration, which is usually expressed as:
[0024]
[0025] Among them, f x ,f y Let (u0, v0) represent the focal length of the camera along the X and Y axes (in pixels), respectively. (u0, v0) is the position of the optical center in the image coordinate system. Using the shoulder width of the human body as a reference, the coordinates of the shoulder key points are extracted from the bounding box of each student after cropping, and the corresponding pixel values are calculated. The formula for calculating shoulder width in pixels is as follows:
[0026] ω pixels =|x R -x L |
[0027] Where, ω pixels The shoulder width is represented in pixels, x L and x RThese represent the horizontal pixel positions of the left and right shoulder keypoints in the image, respectively.
[0028] (2) Utilizing the actual shoulder width ω of the human body real Shoulder width ω (in pixels) pixels The depth value z (distance along the Z-axis) between the student and the camera is calculated based on the principle of triangulation. The formula for calculating the depth value z is as follows:
[0029]
[0030] Where, ω real The actual shoulder width of the student;
[0031] (3) By eliminating camera intrinsic parameters, the two-dimensional image coordinates (u,v) are converted into normalized image coordinates (xnorm,ynorm). The conversion formula is as follows:
[0032]
[0033] (4) The three-dimensional coordinates (x, y, z) in the camera coordinate system are obtained using the following formula:
[0034] x = x norm ·z,y=y norm ·z,z=z
[0035] Furthermore, the estimation of student head pose is performed using the deep convolutional neural network model 6DRepNet. This model uses ResNet-50 as the backbone network and predicts the six-dimensional representation of the rotation matrix through regression. It is then converted into a complete rotation matrix, and three Euler angles are calculated from it to accurately describe the student's head pose.
[0036] Furthermore, the emotion recognition module employs the EmoFAN facial expression recognition model, which is based on a facial alignment network and can simultaneously predict discrete emotion categories, continuous emotion dimensions, and facial key points. During the feature mapping stage, the model focuses on key facial regions through an attention mechanism. By jointly estimating discrete and continuous emotions, it performs fine-grained tracking of students' emotional dynamics in classroom scenarios.
[0037] Furthermore, the identity matching module employs the SphereFace model for identity matching. This model introduces the A-Softmax loss function and explicitly applies angular interval constraints on the hyperspherical manifold, enabling the convolutional neural network to learn highly discriminative angular feature representations. Before class, all students must face the camera to ensure that the camera can capture clear facial images of each student, and these images are used as source data for identity matching. After the class officially begins, the framework matches the newly acquired student images with the source data and saves the matched student information, thereby achieving continuous tracking of students.
[0038] Furthermore, the quantitative evaluation module incorporates three representative dimensions of visual cues: behavior, attention, and emotion. The fusion strategy for these dimensions is as follows:
[0039] (1) Behavioral score
[0040] Students were assigned different scores for four behavioral types (sitting upright (US), standing (SD), leaning on the desk (LD), and turning sideways (SS)). Bi Students are awarded the following scores: 60 points for maintaining proper posture; 80 points for standing up to answer questions or interact with the teacher; 40 points for turning to the side; and 20 points for slouching on the desk and showing no interest. The quantitative formula for student behavior is as follows:
[0041] S Beh =S Bi
[0042] (2) Attention Score
[0043] Students' attention is categorized into three types: focused, distracted, and detached, with corresponding score ranges of 7-10 points, 4-7 points, and 0-4 points, respectively. Each student's specific attention score is determined based on their attention category and the deviation of their head posture from the average head posture of all students in the class. The specific calculation process is as follows:
[0044]
[0045] Among them, S Att θs and ψs represent the current student's pitch and yaw angles, respectively; θc and ψc represent the average pitch and yaw angles of all students, respectively; d pose d represents the difference between the student's head posture and the class average posture. max and d min S1 and S2 represent the maximum and minimum differences obtained after calculating the differences between the head postures of all students and the average posture of the class, respectively; S1 and S2 represent the lower and upper limits of the score range for the current category, respectively.
[0046] (3) Emotional score
[0047] Emotion scores are primarily represented by arousal and valence, both ranging from -1 to 1. These represent continuous ranges from low to high arousal and from negative to positive emotions, respectively. The emotion scoring function is defined as follows:
[0048] S Emo = (A+1)×V
[0049] Here, A and V represent arousal value and efficacy value, respectively. The calculation process of the emotion score SEmo is as follows: First, the arousal value A is converted from its original range [-1,1] to [0,2]; then the adjusted arousal value is multiplied by the efficacy value V.
[0050] (4) Overall Interest Score
[0051] The behavioral score has been converted to a 100-point scale, and the attention and emotion scores have also been converted to the same scoring scale (100 points), which are S respectively. Att_perc S Emo_perc The weights obtained through the Analytic Hierarchy Process (AHP) are applied to these three dimensions to calculate the final classroom interest score (S). Int_perc The relevant formulas are as follows:
[0052] S Att_perc =S Att ×10
[0053]
[0054] S Int_perc =w1×S Beh
[0055] +w2×S Att_perc
[0056] +w3×S Emo_perc
[0057] Among them, S min and S max These represent the minimum and maximum values of the pre-standardized emotion score, respectively, with values of -2 and 2; w i To correspond to the weights of the three dimensions, when facial information is unavailable, the interest score is estimated solely based on the student's behavior and attention. The relevant formula is as follows:
[0058]
[0059] Another objective of this invention is to provide an intelligent assessment method for student classroom interest based on multi-visual perception, the method specifically comprising:
[0060] S1: At the start of class, students are required to look up and face the camera. The facial images of each student are collected and stored as source data for subsequent identity matching.
[0061] S2: Use the YOLOv11 model to recognize key behaviors such as sitting sideways, leaning on a table, and standing.
[0062] S3: Use head posture estimation and other technologies to determine whether students are focused on the blackboard and classify attentional states into three types: focused, distracted and detached.
[0063] S4: The EmoFAN facial expression recognition model can perform fine-grained tracking of students' emotional dynamics in classroom scenarios.
[0064] S5: Uses the SphereFace model for identity matching;
[0065] S6: Use the Analytic Hierarchy Process (AHP) to determine the relative importance of the three key dimensions of behavior, attention, and emotion in assessing students’ classroom interest;
[0066] S7: The results obtained from the three visual dimensions are converted into numerical scores to provide a calculable basis for assessing classroom interest. These dimensions are then integrated to comprehensively evaluate students' classroom interest.
[0067] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0068] This invention extracts behavioral, attentional, and emotional features from visual data in a non-invasive and contactless manner, quantifies these features to form interest scores, and finally uses face matching to ensure the consistency of each student's identity information, thus achieving automated assessment of students' classroom interest levels. In the attention estimation module, methods such as coordinate transformation, head pose estimation, and geometric derivation are used to estimate students' head poses and the range of head poses for students in different positions when focusing on the blackboard, thereby determining whether a student is paying attention to the blackboard. Experimental results on a self-built classroom dataset verify the effectiveness of this method, providing valuable technical support for intelligent classroom interest assessment.
[0069] The expected benefits and commercial value of the technical solution of this invention after transformation are: to provide automated assessment for classroom teaching;
[0070] The technical solution of this invention fills a technological gap in the industry both domestically and internationally: by extracting students' behavioral, attentional, and emotional characteristics from visual data, it enables the automated assessment of students' classroom interest. In offline classrooms, this invention proposes a novel method for estimating student attention. Attached Figure Description
[0071] Figure 1 This is a block diagram of the intelligent assessment system for student classroom interest based on multi-visual perception provided in an embodiment of the present invention;
[0072] Figure 2 This is the workflow of the student attention estimation module provided in this embodiment of the invention;
[0073] Figure 3 This is a schematic diagram of the camera coordinate system and the typical head posture range of a student facing the blackboard, provided in an embodiment of the present invention.
[0074] Figure 4 This is a schematic diagram of head posture angles provided in an embodiment of the present invention: yaw angle, pitch angle, and roll angle;
[0075] Figure 5 This refers to the allowable range of yaw and pitch angles when viewing the blackboard, as provided in the embodiments of the present invention.
[0076] Figure 6 These are examples of classroom behaviors provided in embodiments of the present invention;
[0077] Figure 7 This is a visualization of the head pose estimation results provided in the embodiments of the present invention;
[0078] Figure 8 This is the visualization of the estimated attention state provided in the embodiments of the present invention;
[0079] Figure 9 This is the student emotion visualization provided in the embodiments of the present invention;
[0080] Figure 10 This is the distribution of attention scores for all sample instances of 42 students provided in this embodiment of the invention;
[0081] Figure 11 This is the distribution of interest levels of all sample instances of 42 students provided in this embodiment of the invention;
[0082] Figure 12 These are example student images corresponding to Table 9 provided in this embodiment of the invention;
[0083] Figure 13 This is a trend chart of interest scores for classes and individual students during the course period provided in an embodiment of the present invention;
[0084] Figure 14 This is a flowchart of the intelligent assessment method for student classroom interest based on multi-visual perception provided in an embodiment of the present invention. Detailed Implementation
[0085] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0086] In existing classroom teaching processes, it is difficult to objectively and quantitatively assess students' learning interest and concentration in real time. Traditional teaching management relies heavily on teachers' experience and subjective observation, which often has significant limitations in large-class or multi-task teaching scenarios. For example, it is difficult for teachers to accurately grasp the behavior, attention status, and emotional changes of dozens of students simultaneously, leading to delayed classroom feedback and untimely interventions, thus affecting teaching quality and individualized student development. The core industry application problem that this system aims to solve is to provide a scalable and real-time means of detecting classroom interest, providing technical support for smart education and personalized teaching.
[0087] In practical applications, the behavior detection module undertakes the task of basic perception. By recognizing students' body postures in classroom scenarios through a deep learning detection framework, different states such as sitting upright, standing, turning to the side, and leaning on the desk can be transformed into standardized behavior categories. This process avoids the instability of relying on manual observation, making the dynamic changes in student classroom participation calculable and traceable. As an important explicit indicator of learning interest, the quantification of behavior helps to construct the first dimension of interest assessment.
[0088] The attention estimation module plays a crucial role in the system, its principle based on the combination of computer vision and 3D geometric modeling. By analyzing the proportional relationship between the camera and the actual size of the human body, the system can calculate the student's spatial position in the camera coordinate system and further use a deep convolutional neural network to regress the head pose angle. This angle is compared with the preset blackboard attention range to determine whether the student is maintaining visual focus on the teaching content. This achieves cross-modal conversion from 2D images to 3D spatial attention inference, solving the problem of traditional classrooms' inability to accurately measure students' attention levels.
[0089] In the emotion recognition stage, the system introduces an expression recognition network based on facial alignment and attention mechanisms. This network can not only classify basic emotions but also quantify them on two continuous dimensions: valence and arousal, constructing a more nuanced emotional profile than discrete labels. This mechanism allows teachers to not only understand the types of students' emotions but also grasp their intensity and tendency, forming a real-time perception of the classroom atmosphere. Through the temporal tracking of emotional dynamics, the system can reveal students' emotional fluctuations in different teaching stages, thereby helping to optimize the pace of the lesson.
[0090] The introduction of the identity matching module ensures the continuity and individualization of interest assessment. In large-scale classrooms, without effective identity identification, the system struggles to distinguish an individual's long-term interest status. Through a highly discriminative facial feature learning and matching mechanism, the system binds student identity with interest indicators, ensuring data integrity in the vertical dimension. This not only solves the problem of multi-target tracking in multi-person classrooms but also provides reliable support for the subsequent establishment of personalized learning profiles.
[0091] The quantitative assessment module transforms information from three dimensions—behavior, attention, and emotion—into a unified interest score. The analytic hierarchy process (AHP) is used to determine the weights, ensuring the evaluation system is both theoretically sound and flexibly adjustable to different teaching scenarios. Through this fusion of multi-source visual information, the system achieves a comprehensive characterization of students' classroom interest, providing teachers with data-driven decision support. This working principle realizes a closed-loop process from perception to cognition to quantification, filling the gap in existing educational technology for the objective assessment of classroom interest and possessing clear industrial application value.
[0092] like Figure 1 As shown, this embodiment of the invention provides a student classroom interest intelligent assessment system based on multi-visual perception. The system specifically includes:
[0093] The behavior detection module uses the YOLOv11 model to identify key behaviors such as sitting upright, turning to the side, leaning on the table, and standing.
[0094] The attention estimation module uses head posture estimation and other technologies to determine whether students are focused on the blackboard and classifies attention states into three types: focused, distracted and detached.
[0095] The emotion recognition module assesses students' emotional valence and arousal levels based on facial expressions.
[0096] The identity matching module matches each student's facial image with the source data collected before class;
[0097] The quantitative assessment module converts the extracted behavioral, attentional, and emotional information into numerical scores, and integrates these scores through a weighted summation to generate the final interest score.
[0098] The behavior detection module categorizes student behavior into four types: sitting upright, standing, turning sideways, and leaning on the desk. These categories form the basis of behavior recognition, which is achieved through an object detection model. However, in a standard classroom environment, factors such as occlusion, lighting variations, and differences in image resolution often pose challenges to accurate behavior detection. To address these issues, this invention employs YOLOv11, a relatively lightweight, single-stage object detection algorithm renowned for its fast and efficient performance, making it particularly suitable for deployment in scenarios with limited computing resources. Furthermore, YOLOv11 exhibits strong robustness to various environmental conditions, including lighting variations, occlusion, and viewing angle differences. These characteristics enable it to effectively adapt to various classroom scenarios and meet the basic requirements for student behavior detection.
[0099] Students' attention is closely related to their interest in class, as sustained attention often reflects a higher level of participation and intrinsic motivation. The attention estimation module categorizes students' attention into three types based on their head posture: focused, distracted, and detached. Focused refers to students actively paying attention to the lecture content or the blackboard; distracted refers to students not focusing on the blackboard; and detached refers to students neither paying attention to the blackboard nor participating in classroom activities, indicating that their attention is completely detached from the lesson.
[0100] The workflow of the attention estimation module is as follows: First, a cropped single student image is input into the module. The module extracts the coordinates of the student image's center point and converts them into three-dimensional (3D) coordinates. Simultaneously, a 6DRepNet model is used to estimate the student's head pose angle. Second, based on the aforementioned 3D coordinates, the module determines the range of head pose angles related to the student's attention on the blackboard. Finally, the estimated head pose angles are compared with this reference range to infer the student's attention state. The workflow of this module is as follows: Figure 2 As shown.
[0101] For each student, this invention uses the center point of the student's image, the shoulder width in pixels, the actual shoulder width, and camera intrinsic parameters to calculate the student's three-dimensional position in the camera coordinate system. Figure 3 The diagram illustrates the camera coordinate system and the range of head poses when a student is facing the blackboard.
[0102] First, the camera intrinsic parameter K is obtained through camera calibration, which is usually expressed as:
[0103]
[0104] Among them, f x ,f y These represent the focal lengths of the camera along the X and Y axes (in pixels), respectively, and (u0, v0) is the position of the optical center in the image coordinate system.
[0105] This invention uses human shoulder width as a reference, extracts the coordinates of key shoulder points from the bounding box of each student after cropping, and calculates the corresponding pixel values. The formula for calculating shoulder width in pixels is as follows:
[0106] ω pixels =|x R -x L |
[0107] Where, ω pixels The shoulder width is represented in pixels, x L and x R These represent the horizontal pixel positions of the left and right shoulder keypoints in the image, respectively.
[0108] Secondly, this invention utilizes the actual shoulder width ω of the human body. real Shoulder width ω (in pixels) pixels The depth value z (distance along the Z-axis) between the student and the camera is calculated based on the principle of triangulation. The formula for calculating the depth value z is as follows:
[0109]
[0110] Where, ω real The actual shoulder width of students is used. Based on actual research, this invention sets the student's shoulder width to 48 centimeters.
[0111] Next, by eliminating camera intrinsics, the two-dimensional image coordinates (u,v) are converted into normalized image coordinates (x,v). norm ,y norm The conversion formula is as follows:
[0112]
[0113] Finally, the three-dimensional coordinates (x, y, z) in the camera coordinate system are obtained using the following formula:
[0114] x = x norm ·z,y=y norm ·z,z=z
[0115] Estimating a student's head posture provides information about their head orientation, which can be used to analyze their attention state. Head posture is typically represented by three angles: yaw, pitch, and roll. Figure 4 As shown, this invention uses the deep convolutional neural network model 6DRepNet for head pose estimation. This model uses ResNet-50 as the backbone network and predicts the six-dimensional representation of the rotation matrix through regression. Then, it is converted into a complete rotation matrix, and three Euler angles are calculated from it to accurately describe the student's head pose.
[0116] Typically, changes in a student's yaw and pitch angles are closely related to their visual attention in the classroom. Specifically, when these two angles are within a certain range, it often indicates that the student's vision is directed towards the blackboard, signifying their participation in the lesson; conversely, significant deviations in yaw or pitch angles may reflect a decline in the student's visual attention to the teaching content. Compared to yaw and pitch angles, roll angle has a smaller impact on attention state estimation; therefore, this invention uses yaw and pitch angles as the primary evaluation indicators.
[0117] To determine whether a student is focused, this invention uses the blackboard as a reference area for estimating gaze: if a student's gaze falls within the blackboard area, they are considered focused. The horizontal and vertical range of the gaze projection is determined by the corresponding range of the student's head posture angle in the horizontal and vertical directions. Only when the yaw and pitch angles of the head are within a reasonable range can the student's gaze point cover the blackboard area, indicating that their attention is focused on the classroom content. Therefore, this invention first needs to determine the range of the yaw angle (denoted as ψ) and the pitch angle (denoted as θ). Specifically, given a student (denoted as A) located at coordinates (x, y, z) in three-dimensional space, and assuming the height and width of the blackboard are w and h respectively, then the ranges of ψ and θ ([ψ1, ψ2][θ1, θ2]) can be derived from geometric relationships (e.g., ...). Figure 5 As shown), the corresponding formula is as follows:
[0118]
[0119] Among them, R ψ and R θ These represent the range of values for the yaw angle ψ and the pitch angle θ, respectively.
[0120] As described above, this invention categorizes students' attention states into three types: focused, distracted, and detached, and clearly defines the criteria for determining the focused state. The remaining two states will be defined below: Specifically, if a student's head yaw angle exceeds the effective range but the pitch angle is within the range, or the yaw angle is within the range but the pitch angle exceeds the upper limit, then the student is determined to be in a distracted state; conversely, if both the yaw angle and pitch angle exceed their respective effective ranges, or the yaw angle is within the range but the pitch angle is below the lower limit, then the student is determined to be in a detached state. In summary, the student's attention state (denoted as As) can be formally defined as follows:
[0121]
[0122] In modern teaching, understanding students' emotional states is crucial for improving teaching effectiveness. Traditional methods rely on teachers observing students' emotions, but this becomes inefficient as class sizes increase. By automatically recognizing students' facial expressions and analyzing their emotional responses in real time, deep learning technology provides teachers with precise tools to help optimize teaching strategies and improve learning outcomes.
[0123] The emotion recognition module employs the EmoFAN facial expression recognition model proposed by Toisoul et al. to detect students' emotions in a classroom environment. This model, based on a facial alignment network, can simultaneously predict discrete emotion categories, continuous emotion dimensions (i.e., valence and arousal), and facial key points. During the feature mapping stage, the model focuses on key facial regions through an attention mechanism, improving the accuracy of emotion recognition in occluded or complex environments. By jointly estimating discrete and continuous emotions, EmoFAN can perform fine-grained tracking of students' emotional dynamics in classroom scenarios.
[0124] To achieve continuous tracking of individual students throughout the lesson, cross-temporal identity matching is required between consecutive video frames to establish identity associations at different points in time. The identity matching module employs the SphereFace model. This model introduces the A-Softmax loss function and explicitly applies angular interval constraints on the hyperspherical manifold, enabling the convolutional neural network to learn highly discriminative angular feature representations. After training on a large-scale face dataset, this model effectively handles various challenges in real-world classroom environments (such as complex lighting, occlusion, and pose changes) while maintaining stable and high-precision identity matching performance.
[0125] Before class, all students must face the camera to ensure it captures a clear facial image of each student, which is then used as source data for identity matching. Once the class begins, the framework matches the newly acquired student images with the source data and saves the matched student information, thus enabling continuous student tracking.
[0126] Traditional methods for estimating student classroom interest based on visual information typically rely on a single type of visual cue, such as attention or facial emotion. However, assessing classroom interest from only a single perspective often introduces bias and has inherent limitations. To improve the comprehensiveness and accuracy of the assessment, this quantitative assessment module incorporates three representative dimensions of visual cues: behavior, attention, and emotion, to reflect students' classroom interest from multiple perspectives. These dimensions capture different levels of student classroom participation, helping to build a more reliable and interpretable interest assessment model. The strategy for integrating these dimensions is as follows:
[0127] (1) Hierarchical structure construction
[0128] The Analytic Hierarchy Process (AHP) was used to determine the relative importance of the three key dimensions—behavior, attention, and emotion—in assessing students' classroom interest, as shown in Table 1. The AHP constructs a judgment matrix through pairwise comparisons and calculates the weights of each dimension, providing a structured and interpretable basis for interest assessment. The following section will explain the construction process of the judgment matrix and the calculation method for the corresponding weight coefficients of these three dimensions.
[0129] Table 1: Selected Criteria for Assessing Student Classroom Interest
[0130]
[0131] (2) Judgment matrix construction and weight calculation
[0132] To construct pairwise comparison matrices using the Analytic Hierarchy Process (AHP), this invention designed a "Student Classroom Interest Assessment Questionnaire." The survey included six participants, comprising graduate students from different years of computer science and teachers from related disciplines. Prior to the survey, researchers explained the principles of the AHP to the participants in detail, including the principles of hierarchical structure construction, the process of constructing pairwise comparison matrices, and the application scenarios and assessment objectives of this invention. Subsequently, participants were guided to use the 1-9 scale proposed by Saaty (as shown in Table 2) to conduct pairwise comparisons of the three dimensions of classroom interest (behavior, attention, and emotion), assigning relative importance values based on the provided instructions and their own professional judgment.
[0133] Table 2: Saaty pairwise comparisons using the 1-9 scale
[0134] Scale meaning 1 This indicates that factors A and B are equally important. 3 This indicates that factor A is slightly more important than factor B. 5 This indicates that factor A is significantly more important than factor B. 7 This indicates that factor A is significantly more important than factor B. 9 This indicates that factor A is extremely important than factor B. 2,4,6,8 This represents the intermediate value of the above adjacent judgments. reciprocal Used for reverse comparison
[0135] To ensure logical consistency in the judgments, a consistency check (e.g., calculating the consistency ratio) is required for each individual judgment matrix. Only matrices that meet the consistency threshold are included in the aggregation process. The final aggregated pairwise comparison matrix is obtained by calculating the geometric mean of the corresponding elements of all valid individual matrices. Subsequently, the aggregated matrix also needs to undergo a consistency check to ensure its validity. Once the consistency requirement is met, the matrix is used to calculate the priority weights of the three evaluation factors: behavior, attention, and emotion. Let the aggregated matrix be A∈Rn×n, and its largest eigenvalue be denoted as λ. max In this invention, the number of evaluation factors is 3, i.e., n = 3. Therefore, the consistency test index (CI) is defined as follows:
[0136]
[0137] To determine whether the CI value is within an acceptable range, the random consistency test ratio CR is introduced, which is defined as follows:
[0138]
[0139] Wherein, RI is the random consistency index, and its value is determined according to Table 3. When the consistency ratio CR < 0.1, the matrix is considered to have reasonable consistency. Once the matrix meets the consistency condition, it is used to calculate the weights of each evaluation factor. The specific calculation process is as follows:
[0140]
[0141] Table 3: Random Consistency Index (CR) values corresponding to different criteria
[0142] n 1 2 3 4 5 RI 0 0 0.52 0.90 1.12
[0143] Among them, a ij Let w be the element in the i-th row and j-th column of matrix A. i Let w1, w2, and w3 represent the weights of the i-th factor. Here, w1, w2, and w3 correspond to the weights of behavior, attention, and emotion, respectively. Table 4 provides detailed calculation results. As shown in the table, the CR value of this invention is less than 0.1, indicating that the judgment matrix meets the consistency requirement, and the obtained weights are reliable and effective.
[0144] Table 4: Judgment Matrix PA
[0145]
[0146] (3) Quantification and Evaluation
[0147] The results obtained from three visual dimensions using computer vision technology are converted into numerical scores, providing a calculable basis for assessing classroom interest. These dimensions are then integrated to comprehensively evaluate students' classroom interest.
[0148] (3.1) Behavioral score
[0149] This invention assigns different scores to four student behavior types (sitting upright (US), standing (SD), leaning on the desk (LD), and turning sideways (SS)), as shown in Table 5. This scoring design is based on the following considerations: upright sitting posture is generally considered a sign of strong classroom interest, and is awarded 60 points; students who stand up to answer questions or interact with the teacher are awarded the highest score of 80 points; when a student turns sideways, they are highly likely to be talking to other students (although there is also the possibility of discussing academic issues), and this behavior is awarded 40 points; when a student leans on the desk and shows a lack of interest, this behavior is awarded 20 points.
[0150] Table 5: Scores for Each Behavior
[0151] category US SD SS LD Fraction (SBi) 60 80 40 20
[0152] The formula for quantifying student behavior is as follows:
[0153] S Beh =SBi
[0154] (3.2) Attention Score
[0155] Students' attention is categorized into three types: focused, distracted, and detached, with corresponding score ranges of 7-10 points, 4-7 points, and 0-4 points, as shown in Table 6. Each student's specific attention score is determined based on their attention category and the deviation of their head posture from the average head posture of all students in the class. The closer a student's head posture is to the overall class average, the higher their score. The specific calculation process is as follows:
[0156]
[0157] Among them, S Att For attention score; θ s and ψ s These are the pitch and yaw angles of the current student, respectively; θ c and ψ c These are the average pitch angle and average yaw angle for all students, respectively; d pose d represents the difference between the student's head posture and the class average posture. max and d min S1 and S2 represent the maximum and minimum differences obtained after calculating the differences between the head postures of all students and the average posture of the class, respectively; S1 and S2 represent the lower and upper limits of the score range for the current category, respectively.
[0158] Table 6: Score Range for Each Attention State
[0159] category Focus Distraction break away Fraction range (7-10] (4,7] [0-4]
[0160] (3.3) Emotional Score
[0161] Emotional scores are primarily represented by arousal and valence. Arousal reflects the intensity of emotional activation, while valence reflects the positive or negative nature of the emotion. Both arousal and valence range from -1 to 1, representing a continuous range from low to high arousal and from negative to positive emotions, respectively. Based on extensive teacher observations and teaching experience, this invention finds that when students are in a positive emotional state, the higher their emotional intensity, the higher their interest in the classroom; conversely, when students are in a negative emotional state, the higher their emotional intensity, the lower their interest in the classroom. Based on this, this invention defines the emotion scoring function as follows:
[0162] S Emo = (A+1)×V
[0163] Here, A and V represent arousal value and efficacy value, respectively. The calculation process of the emotion score SEmo is as follows: First, the arousal value A is transformed from its original range [-1,1] to [0,2] to ensure that the non-negativity required to emphasize the intensity of emotional activation is met; then, the adjusted arousal value is multiplied by the efficacy value V, thereby comprehensively reflecting the combined effect of the activation intensity and directional tendency of the student's emotional state.
[0164] (3.4) Overall Interest Score
[0165] Behavioral scores have been assigned on a 100-point scale. To ensure consistency, attention and emotion scores have also been converted to the same scoring scale (100-point scale). Subsequently, weights obtained through the Analytic Hierarchy Process (AHP) are applied to these three dimensions to calculate the final classroom interest score. The relevant formula is as follows:
[0166] S Att_perc =S Att ×10
[0167]
[0168] S Int_perc =w1×S Beh +w2×S Att_perc +w3×S Emo_perc
[0169] Among them, S min and S max These represent the minimum and maximum values of the pre-standardized emotion score, respectively, with values of -2 and 2; w i These are the weights corresponding to the three dimensions.
[0170] When facial information is unavailable, interest scores are estimated based solely on students' behavior and attention, using the following formula:
[0171]
[0172] like Figure 14 As shown in the figure, an intelligent assessment method for student classroom interest based on multi-visual perception provided by an embodiment of the present invention specifically includes:
[0173] S1: At the start of class, students are required to look up and face the camera. The facial images of each student are collected and stored as source data for subsequent identity matching.
[0174] S2: Use the YOLOv11 model to recognize key behaviors such as sitting sideways, leaning on a table, and standing.
[0175] S3: Use head posture estimation and other technologies to determine whether students are focused on the blackboard and classify attentional states into three types: focused, distracted and detached.
[0176] S4: The EmoFAN facial expression recognition model can perform fine-grained tracking of students' emotional dynamics in classroom scenarios.
[0177] S5: Uses the SphereFace model for identity matching;
[0178] S6: Use the Analytic Hierarchy Process (AHP) to determine the relative importance of the three key dimensions of behavior, attention, and emotion in assessing students’ classroom interest;
[0179] S7: The results obtained from the three visual dimensions are converted into numerical scores to provide a calculable basis for assessing classroom interest. These dimensions are then integrated to comprehensively evaluate students' classroom interest.
[0180] This invention relates to specific application areas or related products. It can be applied to the field of smart education; intelligent auxiliary teaching products such as smart classroom systems and smart classroom apps.
[0181] This invention first presents experimental and visualization results for three dimensions of interest, and then provides experimental results related to classroom interest assessment.
[0182] 1. Experimental results for each dimension of interest
[0183] (1.1) Behavioral Dataset
[0184] Unlike head pose estimation and facial emotion recognition, which can utilize existing pre-trained models, classroom behavior classification requires customized datasets to train models. Therefore, this invention collects a manually labeled dataset of student classroom behaviors. Data was collected in a standard classroom, with a camera installed directly above the center of the blackboard, recording 400 minutes of real classroom teaching video. To improve data quality and training efficiency, the original video was appropriately cropped, and representative frames were selected, resulting in a dataset containing 2200 images. The sample counts for each behavior type in this dataset are as follows: sitting upright (78857), standing (1171), leaning on the desk (530), and turning sideways (1009), for a total of 81567 samples. This invention divides the dataset into training, validation, and test sets in a 7:2:1 ratio. Sample images of the four behaviors are shown below. Figure 6 As shown.
[0185] (1.2) Experimental Details
[0186] The experiments of this invention were conducted on a device equipped with an NVIDIA RTX 4090 GPU using the PyTorch framework. In the behavior detection task, the YOLOv11x model was trained for 80 epochs on a self-built dataset; in the attention estimation task, the 6DRepNet model was used and trained for 80 epochs on the 300W-LP dataset; in the emotion recognition task, the EmoFAN facial expression recognition model was used.
[0187] (1.3) Behavioral detection results
[0188] To verify the effectiveness of the proposed model in student behavior detection tasks, this invention conducted experimental evaluations using a self-built classroom behavior dataset. In this evaluation, P represents precision, R represents recall, F1 is the harmonic mean of precision and recall, and mAP50 represents the mean average precision across all categories when the Intersection over Union (IoU) threshold is set to 50%. The experimental results are shown in Table 7. The overall mAP50 of the model reached 86.9%, with particularly outstanding detection accuracy for standing (SD) and lying on the table (LD) behaviors, at 99.4% and 98.1%, respectively. These results demonstrate that the YOLOv11 model can effectively recognize student behaviors and meet the needs of student behavior analysis in classroom scenarios.
[0189] Table 7: Evaluation Indicators for Student Behavior Monitoring
[0190]
[0191]
[0192] (1.4) Attention estimation results and visualization
[0193] Figure 7 The visualization results of student head pose estimation are shown, where blue lines represent facial orientation, red lines represent lateral orientation, and green lines represent downward orientation. It is evident that the model's estimation of student head pose effectively reflects students' attention levels, thus largely conforming to the definition of student attention.
[0194] The attention estimation module estimates student attention by fusing head posture with the relative positions of the student and the blackboard. Figure 8 The classification results of the attention estimation module are shown, demonstrating that the module can effectively classify student attention.
[0195] (1.5) Emotion Recognition Results and Visualization
[0196] The emotion recognition module estimates valence (V) and arousal (A) values by analyzing students' facial images, specifically as follows: Figure 9 As shown, this module outputs different valence and arousal values based on changes in facial expressions, effectively reflecting students' emotional states in the classroom. Compared to traditional discrete emotion classification methods, this continuous-dimensional approach is more flexible in describing emotional changes and can better capture the dynamic characteristics of students' emotions in real classroom scenarios.
[0197] 2. Results of the experiment on assessing students' classroom interest
[0198] (2.1) Experimental Design and Implementation
[0199] This invention presents a verification experiment designed and conducted in a real classroom teaching scenario. The experiment was carried out in a standard classroom with 42 students, and the total class time was 45 minutes. By processing the recorded video footage, one frame was extracted every 10 seconds, resulting in 240 classroom images containing 10,080 sample instances.
[0200] This invention systematically selects 60 images at fixed intervals and invites four experts to evaluate students' classroom engagement levels. Each image is evaluated jointly by four experts; if more than two experts give different scores, a re-evaluation process is initiated. Experts use a 5-point Likert scale (where 5 represents the highest level and 1 represents the lowest) to score students' classroom engagement based on their professional experience and observed classroom behavior. Considering the inherent subjectivity of student engagement evaluation, detailed evaluation criteria are shown in Table 8.
[0201] Table 8: Student Interest Labeling and Scoring Criteria
[0202]
[0203] (2.2) Experimental Results and Analysis
[0204] During the entire lesson, the behavior detection module detected 9,921 instances of "sitting upright," 9 instances of "turning to the side," and 150 instances of "leaning on the desk." This indicates that behaviors such as "leaning on the desk," "standing," and "turning to the side" were relatively rare, with most students maintaining a "sitting upright" posture.
[0205] Figure 10 The distribution of attention scores is shown, with attention levels quantified on a scale of 0-10. It is evident that most students are either focused or unengaged, while only a small number are distracted.
[0206] Figure 11The distribution of engagement scores for all student samples in the recorded courses is presented. The results show that most students' scores are concentrated in the 45-50 and 60-65 range. Notably, this distribution pattern closely matches the distribution pattern of student attention scores, indicating a statistically significant correlation between these two indicators. These findings provide empirical evidence for the weighting schemes used for different assessment dimensions.
[0207] Table 9 shows the scores of six students at a certain moment in the three dimensions of behavior, attention, and emotion, as well as their comprehensive interest score. The corresponding visualization results are as follows: Figure 12 As shown, students 1-3 achieved relatively high interest scores due to their high scores in behavior and attention, as well as strong emotional engagement, indicating a high level of classroom interest. Conversely, students 4 and 5 had significantly lower attention scores, reflecting distracted behavior; although their emotional scores were at a moderate level, their overall interest level remained low. This highlights the significant impact of attention on the overall interest score. Student 6 lacked facial information, indicating that this student was not participating in class; the assessment of this student was based solely on the behavior and attention dimensions, resulting in the lowest overall score. This result demonstrates that the absence of facial cues has a significant impact on interest estimation.
[0208] Table 9: Student Scores in Each Dimension
[0209]
[0210] To evaluate the impact of the weighting scheme proposed in this invention on student interest assessment, it is compared with an equal-weight scheme (0.33:0.33:0.33). During the analysis, students are divided into high, medium, and low interest groups, and the average interest score for each group is provided. For ease of comparison, the interest scores provided by experts are mapped to a 0-100 score range, and the difference between the model's predicted values and the expert scores is presented, as shown in Table 10.
[0211] The results show that under the weighting scheme proposed in this invention, the score differences between the high interest group and the medium interest group, and between the medium interest group and the low interest group, are 6.69 and 6.35, respectively; while under the equal weighting scheme, the score differences are 9.81 and 7.20, respectively. Although the equal weighting scheme exhibits greater inter-group discrimination, it also deviates more from the expert annotation results.
[0212] Further analysis revealed that the standard deviation of the attention dimension was 32.21, significantly higher than that of the behavior dimension (3.95) and the emotion dimension (4.19). While this high variability in attention scores may seem beneficial for differentiating students, over-reliance on this single dimension could amplify random fluctuations rather than reflecting true differences in interest levels, thus affecting the robustness of the assessment. In contrast, the weighting scheme of this invention, based on expert knowledge, maintains effective differentiation of interest levels while demonstrating higher consistency with actual annotation results.
[0213] Table 10: Average scores of the three interest groups under different weights
[0214]
[0215] This invention uses student identity matching technology to track the interest levels of all students throughout the lesson and divides them into high, medium, and low interest groups based on equal intervals between the highest and lowest interest scores. The high interest group has 27 students, the medium interest group has 6 students, and the low interest group has 9 students. Table 11 presents the average scores of the three groups across each dimension on a 100-point scale. Due to the equal-interval grouping strategy, the differences between groups are relatively gradual, especially in the behavioral and emotional dimensions, where only minor changes were observed. Conversely, the group differentiation in the attention dimension is more significant: the attention scores of the high interest group students are much higher than those of the other two groups, indicating that they are more focused in class.
[0216] Table 11: Average scores for each dimension of the three interest groups
[0217]
[0218] This invention provides a visual analysis of the class average interest score for the entire lesson, as well as the individual interest trends of students 2, 3, and 4, as detailed below. Figure 13 As shown in the image, the green curve represents the average class interest level, which is calculated by averaging the interest scores of all students. It can be observed that the class average score fluctuates around 55 points, indicating a good overall level of classroom interest.
[0219] Students 2 and 3 showed higher than the class average for most of the class time. However, Student 2's interest declined at the beginning of the class, and both Students 2 and 3 showed a decline in interest at the end of the class. Conversely, Student 4's interest remained below the class average throughout the class, indicating that this student had low interest in the course content.
[0220] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0221] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A student classroom interest intelligent assessment system based on multi-visual perception, characterized in that, include: The behavior detection module is used to identify students' posture and behavior in the classroom; The attention estimation module is used to determine the student's attention state based on their head posture. The emotion recognition module is used to identify the valence and arousal of students' emotions based on facial expressions. The identity matching module is used to match the facial images collected in the classroom with pre-stored facial source data; The quantitative assessment module is used to convert the behavioral, attentional, and emotional information into numerical scores, and generate interest scores through weighted fusion.
2. The system as described in claim 1, characterized in that, The behavior detection module categorizes student behavior into four types: sitting upright, standing, turning to the side, and lying on the desk.
3. A method for estimating student attention in the classroom, characterized in that, Includes the following steps: Convert the center point coordinates of the student image into three-dimensional spatial coordinates; Calculate the spatial depth between the student and the camera based on camera parameters; Determine whether a student is focused on the blackboard based on three-dimensional coordinates and head posture angle.
4. The attention estimation method as described in claim 3, characterized in that, The head pose angles are estimated by using a convolutional neural network to estimate the rotation matrix and then converted into Euler angles to represent the student's head pose.
5. A method for recognizing student emotions in the classroom, characterized in that, Includes the following steps: Collect student facial images and align feature points; Predicting students' emotion categories and continuous emotion dimensions based on facial expression recognition models; The efficacy value and arousal value are combined to quantify students' emotions.
6. The emotion recognition method as described in claim 5, characterized in that, The emotion quantification is obtained by mapping the arousal value to a non-negative interval and then multiplying it by the effect value.
7. A student identity matching method, characterized in that, Includes the following steps: Collect student facial images and store them as source data; Extracting angular features from student images using a face recognition network; The system compares real-time images captured in the classroom with the source data and outputs the matching results.
8. The identity matching method as described in claim 7, characterized in that, The face recognition network achieves high-discrimination angular feature learning by introducing angular interval constraints on the hyperspherical manifold.
9. A method for quantitatively assessing classroom interest, characterized in that, Includes the following steps: Obtain numerical scores for student behavior, attention, and emotions; Normalize the fractions to the same dimension; The weights of each dimension are calculated based on the analytic hierarchy process (AHP). The final interest score is obtained through weighted fusion.
10. The interest quantification assessment method as described in claim 9, characterized in that, In the absence of access to students' facial information, interest scores are estimated solely based on behavioral and attentional dimensions.