A teaching process intelligent evaluation method and system based on learning behavior analysis

By extracting specular and dark kernel proxy masks from classroom video streams, calculating the symmetrical pupil specular arc index, and performing weighted average calculations in a knowledge graph, the problem of glasses artifacts caused by polarized screen reflections is solved, improving the accuracy of student behavior recognition and the interpretability of teaching assessment.

CN122336628APending Publication Date: 2026-07-03SICHUAN TONGLI CO CREATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN TONGLI CO CREATION TECH CO LTD
Filing Date
2026-03-31
Publication Date
2026-07-03

Smart Images

  • Figure CN122336628A_ABST
    Figure CN122336628A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent evaluation method and system for the teaching process based on learning behavior analysis, belonging to the field of intelligent evaluation technology for the teaching process. The method includes: acquiring classroom videos to extract student eye region images and establishing a mapping between time blocks and teaching segments; extracting specular and dark nucleus proxy masks and calculating the symmetrical pupil-occluded specular arc index; calculating the evidence reliability weight and eye region visibility evidence value accordingly; constructing a classroom knowledge graph containing evidence and artifact nodes, and writing the above parameters into the graph; retrieving the evidence set of the target segment, using the reliability weight to weight and aggregate the visibility evidence value to obtain an attention score and generate a traceability link. This invention effectively suppresses polarization screen reflection interference by quantifying physical artifacts and utilizing graph aggregation, achieving interpretability and traceability of the evaluation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent assessment technology for the teaching process, and in particular to an intelligent assessment method and system for the teaching process based on learning behavior analysis. Background Technology

[0002] Existing smart classroom environments are typically equipped with multiple high-definition cameras to capture real-time video streams of teachers and students during classroom teaching. Computer vision technology is then used to intelligently analyze classroom focus, interaction quality, and the effectiveness of teaching activities. In this scenario, accurately acquiring students' eye state, such as gaze direction and eye opening / closing status, is the core basis for quantitatively assessing students' listening focus and cognitive state. However, in actual classrooms, interactive large screens or LCD displays are commonly installed at the front. The light emitted from these screens has polarization characteristics. When students wearing glasses face the large screen, the polarized light is reflected by the glasses lenses, easily forming symmetrical and significantly bright arcs or patches on the left and right lenses. This physical reflection phenomenon can cover key areas of the eye, such as the pupil or palpebral fissure, causing local contrast saturation or inversion in the eye area image. This severely interferes with behavior recognition algorithms based on eye area images, making it difficult for the system to accurately capture students' true gaze and focus state.

[0003] Current technologies for handling such complex lighting interference often rely on general image enhancement or simple threshold filtering methods to attempt to eliminate the effects of highlights during the preprocessing stage, or to directly remove sample data containing reflections. These methods have significant limitations when dealing with specific lens artifacts caused by polarizing screen reflections. On the one hand, simple image processing struggles to distinguish between the highlight arc on the lens and the natural highlights of the eye, easily resulting in the loss of valuable information and key eye features. On the other hand, directly removing data affected by reflections significantly reduces the sample size, compromising the continuity and integrity of the teaching process assessment. It becomes impossible to distinguish whether low scores are due to a genuine decrease in student attention or physical reflection interference, leading to a lack of interpretability and traceability in the assessment results, and failing to provide accurate attribution evidence for teaching improvement. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies, such as the inability to effectively identify and suppress the artifacts of the symmetrical pupillary highlight arc induced by polarized screen reflection in eyeglasses, which leads to a high misjudgment rate in identifying students' eye zone behavior, distorted results in the assessment of attention during the teaching process, and difficulty in accurately tracing the causes of low scores. Therefore, this invention proposes an intelligent assessment method and system for the teaching process based on learning behavior analysis.

[0005] To address the problems existing in the prior art, the present invention adopts the following technical solution:

[0006] A method for intelligent assessment of the teaching process based on learning behavior analysis, comprising:

[0007] S1. Collect video streams from classroom cameras and perform face detection and tracking, extract images of students' eye areas, and construct a collection time block from several consecutive frames to establish a mapping between the collection time block and the teaching segment.

[0008] S2. Extract the highlight mask and dark nucleus proxy mask from the student's eye region image, and calculate the symmetrical pupil-masking highlight arc index;

[0009] S3. Calculate the evidence reliability weight based on the symmetrical pupil-masking highlight arc index, and calculate the eye region visibility evidence value based on the overlap relationship between the highlight mask and the dark nucleus proxy mask.

[0010] S4. Establish a classroom knowledge graph, which includes evidence nodes, artifact nodes, student nodes and teaching segment nodes. Write the evidence reliability weight and the eye zone visibility evidence value into the attribute field of the evidence node, and associate the symmetrical pupil occlusion highlight arc index with the artifact node.

[0011] S5. Retrieve the set of evidence nodes pointing to the target teaching segment from the classroom knowledge graph. Perform a weighted average calculation on the visual visibility evidence value based on the evidence reliability weight of each evidence node in the evidence node set to obtain the segment attention score and generate traceability link information.

[0012] Compared with the prior art, the beneficial effects of the present invention are:

[0013] 1. This invention extracts the highlight mask and dark kernel proxy mask from the student's eye area image, calculates the symmetrical pupil-occlusion highlight arc index, and generates evidence reliability weights and eye area visibility evidence values ​​based on this index. Using a knowledge graph, the evidence reliability weights, eye area visibility evidence values, and the symmetrical pupil-occlusion highlight arc index are associated with corresponding evidence nodes and artifact nodes. This enables explicit modeling and quantification of the symmetrical pupil-occlusion highlight arc artifact induced by polarized screen reflection in eyeglasses. It effectively distinguishes between physical reflection artifacts and real eye behavior, automatically weakens evidence weights heavily contaminated by artifacts during the evaluation process, suppresses the negative impact of eye area feature loss due to polarized light reflection on attention assessment, and improves the robustness and accuracy of student behavior state recognition under complex lighting conditions.

[0014] 2. This invention constructs an interconnected network in a classroom knowledge graph, including evidence nodes, artifact nodes, student nodes, and teaching segment nodes. It retrieves a set of evidence nodes pointing to the target teaching segment from the graph and calculates a segment attention score by performing a weighted average of the visual visibility evidence values ​​based on evidence reliability weights. Simultaneously, it generates traceability link information including camera identifiers, student identifiers, and symmetrical pupil occlusion highlight arc indexes associated with the evidence nodes. This not only outputs a teaching process score resistant to artifact interference but also accurately identifies the specific reasons for low scores, clearly distinguishing between a genuine decrease in student attention and interference from physical reflections from specific camera positions or students. This achieves interpretability and traceability of the teaching process evaluation results, providing precise attribution basis for teaching improvement. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0016] Figure 1 This is a flowchart illustrating an intelligent assessment method for the teaching process based on learning behavior analysis, provided as an embodiment of the present invention.

[0017] Figure 2 This is a functional block diagram of an intelligent evaluation system for the teaching process based on learning behavior analysis, provided as an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] Example: This example provides an intelligent assessment method for the teaching process based on learning behavior analysis. See [link to example]. Figure 1 Specifically, including:

[0020] S1. Collect video streams from classroom cameras and perform face detection and tracking, extract images of students' eye areas, and construct a collection time block from several consecutive frames to establish a mapping between the collection time block and the teaching segment.

[0021] S2. Extract the highlight mask and dark nucleus proxy mask from the student's eye region image, and calculate the symmetrical pupil-masking highlight arc index;

[0022] S3. Calculate the evidence reliability weight based on the symmetrical pupil-masking highlight arc index, and calculate the eye region visibility evidence value based on the overlap relationship between the highlight mask and the dark nucleus proxy mask.

[0023] S4. Establish a classroom knowledge graph, which includes evidence nodes, artifact nodes, student nodes and teaching segment nodes. Write the evidence reliability weight and the eye zone visibility evidence value into the attribute field of the evidence node, and associate the symmetrical pupil occlusion highlight arc index with the artifact node.

[0024] S5. Retrieve the set of evidence nodes pointing to the target teaching segment from the classroom knowledge graph. Perform a weighted average calculation on the visual visibility evidence value based on the evidence reliability weight of each evidence node in the evidence node set to obtain the segment attention score and generate traceability link information.

[0025] In an embodiment of the present invention, video streams from classroom cameras are acquired and face detection and tracking are performed. Student eye region images are extracted, and several consecutive frames are combined to form a time block. A mapping between the time block and the teaching segment is established, including:

[0026] In this embodiment, the multi-camera array used to collect classroom video streams includes at least three high-definition network cameras with student perspectives. Two cameras are symmetrically mounted on the walls on either side of the interactive screen at the front of the classroom, with their lenses tilted horizontally inward at a 15-degree angle towards the student seating area. The third camera is mounted in the center of the ceiling at the back of the classroom, with its lens tilted downward at a 30-degree angle towards the entire student seating area. All three cameras use a resolution of 1920×1080 pixels and a frame rate of 30 frames per second. Automatic white balance and automatic exposure are enabled to adapt to changes in lighting conditions at different times in the classroom. During the acquisition process, the cameras transmit real-time video streams to the edge computing device via a wired Ethernet link. The transmission protocol uses the RTSP protocol to ensure the real-time performance and integrity of the video stream. After receiving the multiple video streams, the edge computing device first performs frame-by-frame decoding on each video stream. The decoding format uses H.265 to balance decoding efficiency and storage usage. After decoding, a continuous color frame image corresponding to each video stream is obtained.

[0027] Subsequently, the edge computing device preprocesses each color frame image. The preprocessing process includes denoising the frame image using a Gaussian filtering algorithm. The convolution kernel size of the Gaussian filter is set to 3×3, and the standard deviation is set to 0.8 to remove slight noise interference generated during video acquisition. Then, the frame image is contrast-enhanced using a histogram equalization algorithm to improve the grayscale difference between the face region and the background region, which facilitates the accurate execution of subsequent face detection. After preprocessing, a face detection algorithm based on a convolutional neural network is used to detect the face region in each frame image. This convolutional neural network model is trained based on the MobileNetV3 lightweight network architecture and can quickly identify face regions with different poses and distances in the image. During the detection process, the frame image is first subjected to multi-scale scaling processing, with scaling scales of 0.8x, 1.0x, 1.2x, and 1.5x, respectively, to adapt to face regions of different sizes.

[0028] The model inference outputs the bounding box coordinates of all face regions in each frame of the image. The bounding boxes are rectangular, and the coordinates are based on the top-left corner of the frame image, corresponding to the x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner, respectively. For each detected face region, the edge computing device uses a cross-frame tracking method combining Kalman filtering and the Hungarian algorithm for continuous tracking. Kalman filtering is used to predict the position coordinates of each face region in the current frame. The state equation in the prediction process includes two parameters: position and velocity. The process noise variance is set to 0.01, and the observation noise variance is set to 0.1. The Hungarian algorithm is used to associate detected face regions in two adjacent frames. The association process uses the intersection-union ratio (IU / U) of the bounding boxes as a similarity evaluation index. The IU threshold is set to 0.5. When the IU of face regions in two adjacent frames is greater than or equal to 0.5, they are determined to be the same face and a unique tracking identifier is assigned. This tracking identifier remains unchanged throughout the entire classroom video acquisition cycle. If a new face region is detected in a frame and there is no matching historical tracking identifier, a new unique tracking identifier is assigned to it. If the face region corresponding to a certain tracking identifier is not detected for 5 consecutive frames, it is determined that the student has left the classroom and the tracking of that tracking identifier is terminated.

[0029] After completing cross-frame face tracking, the edge computing device further locates and extracts eye region images based on the bounding box coordinates of each face region in each frame image. Eye candidate region segmentation: Based on the face bounding box output by face detection, it is divided into upper, middle, and lower regions proportionally along the height direction. The upper region accounts for 30% of the total height of the face bounding box; this region is the eye candidate region. This proportion setting is based on the well-known common sense of facial feature proportions and can adapt to eye region positioning for different face shapes. Accurate left and right eye positioning: Gray-scale projection analysis is performed on the eye candidate regions. First, the horizontal projection curve is calculated (by calculating the average pixel gray-scale value of each row). The upper and lower boundary points where the gray-scale value drops sharply in the projection curve are found. These boundary points correspond to the upper and lower edges of the eye region. The horizontal projection threshold is set to 0.3 times the average gray-scale value of the eye candidate region; rows with gray-scale values ​​below this threshold are considered valid eye rows. Then, the vertical projection curve is calculated (by calculating the pixel gray-scale value of each column). The average grayscale value of the face is used as the vertical midline of the face bounding box to divide the candidate eye area into left and right parts, corresponding to the left and right eye regions respectively. The vertical projection threshold is set to 0.2 times the average grayscale value of the candidate eye area. Columns with grayscale values ​​lower than this threshold are considered valid eye columns. The initial bounding boxes of the left and right eyes are determined by combining the valid rows and columns of the horizontal and vertical projections. Eye bounding box expansion and cropping: The initial left and right eye bounding boxes are expanded by 5 pixels in all directions. The expansion is based on the occlusion range of different sized eyeglass frames to ensure that the expanded bounding boxes completely cover the pupil, eye fissure, and reflection area of ​​the eyeglass lenses. According to the coordinates of the expanded bounding boxes, the left and right eye area images are cropped from the grayscale image of the preprocessed video frame. The cropped eye area images have 256 grayscale levels, and the resolution can be adaptively adjusted according to the face size. At the same time, the extracted eye area images are associated and stored with the corresponding tracking tags, camera tags, and frame numbers.

[0030] It should be noted that this embodiment uses the SSD face detection model based on the MobileNetV3 lightweight architecture. This model architecture is a well-known technology in the field of computer vision. This embodiment adapts and optimizes the model for classroom scenarios. The specific steps are as follows:

[0031] Model training: The training dataset uses a publicly available face dataset (containing multi-scale, multi-pose, and multi-occlusion face samples), supplemented with 10,000 frames of labeled student faces from real classroom scenarios (labels include face bounding box coordinates and occlusion status); the model input size is fixed at 320×320 pixels; the Adam optimizer is used during training, with an initial learning rate of 0.001, and the learning rate decay strategy is to decrease to 50% of the original value every 5 training epochs; the batch size is set to 32; the loss function is the SSD multi-task loss function, which includes classification loss and bounding box regression loss; the classification loss weight coefficient is set to 1.0, and the regression loss weight coefficient is set to 1.5; the total number of training iterations is 100 epochs; training stops when the validation set accuracy reaches 95%.

[0032] In the model inference stage: The preprocessed classroom video frame images are first subjected to multi-scale scaling at scales of 0.8x, 1.0x, 1.2x, and 1.5x to adapt to the detection requirements of large faces in the front row and small faces in the back row in a classroom setting. The multi-scale images are then input into the trained model, which outputs candidate bounding boxes for each scale along with their corresponding confidence scores. Non-maximum suppression (NMS) is applied to all candidate bounding boxes, with an NMS threshold of 0.45 and a confidence score threshold of 0.6. Only bounding boxes with a confidence score ≥ 0.6 and deduplicated using NMS are retained. The bounding box coordinates are based on the top-left corner of the video frame image, and the output format includes the top-left x-coordinate, top-left y-coordinate, bottom-right x-coordinate, and bottom-right y-coordinate, ensuring that the detection results cover more than 95% of the effective student face regions in the classroom with a false detection rate of less than 5%.

[0033] It should be noted that this embodiment uses the SORT (Simple Online and Real-time Tracking) multi-target tracking framework, which combines Kalman filter prediction with Hungarian algorithm matching. This framework is a well-known technology in the field of multi-target tracking. The implementation details for the classroom scenario are as follows:

[0034] Kalman filter parameter settings: The face tracking state vector is defined as an eight-dimensional vector [x,y,w,h,vx,vy,vw,vh], where x and y are the center pixel coordinates of the face bounding box, w and h are the width and height of the bounding box, and vx, vy, vw, and vh are the motion velocities in the x, y, w, and h directions, respectively. The Kalman filter transition matrix is ​​constructed using a uniform motion model, with the following elements: all diagonal elements are 1, the lower triangular elements corresponding to the velocity terms are 1 (e.g., x row, vx column, y row, vy column), and the remaining elements are 0. The process noise covariance matrix is ​​a diagonal matrix with diagonal elements set to [0.01,0.01,0.01,0.01,0.1,0.1,0.1,0.1]. The observation noise covariance matrix is ​​also a diagonal matrix with diagonal elements set to [0.1,0.1,0.1,0.1]. This parameter setting is suitable for tracking the small head movements of students in the classroom.

[0035] The Hungarian algorithm matching rules are as follows: A cost matrix is ​​constructed using the intersection-union ratio (IOU) of the bounding boxes of two adjacent frames as the similarity index. The element value of the cost matrix is ​​1-IOU, and the IOU is calculated as the ratio of the intersection area to the union area of ​​the two bounding boxes. The IOU matching threshold is set to 0.5. When the IOU of an element in the cost matrix is ​​≥0.5, it is determined to be the same student's face, and a unique tracking identifier (ID) is assigned to it. This ID remains unchanged throughout the entire classroom video acquisition period. If the face bounding box detected in the current frame has no matching historical tracking ID, a new unique ID is assigned to it. If the face bounding box corresponding to a certain tracking ID is not detected for 5 consecutive frames, it is determined that the student has left the detection area or is completely occluded, and the tracking of that ID is terminated.

[0036] Tracking stability optimization: For scenarios with brief occlusion in the classroom (such as students looking down or objects blocking the view), a tracking ID retention mechanism is set up. If the tracking ID does not match the face bounding box for 3 consecutive frames, tracking will not be terminated, but only marked as suspected occlusion. Tracking will resume when a face bounding box that meets the IOU threshold is matched again in subsequent frames, ensuring the continuity of face tracking in classroom scenarios and achieving a tracking success rate of over 90%.

[0037] In this embodiment, classroom teaching organization information is obtained through a wired communication connection established between the edge computing device and the classroom teaching platform and the teacher's terminal. The classroom teaching platform pre-stores the structured lesson plan information for the current class, while the teacher's terminal is used by the teacher to supplement or correct the teaching organization-related content in real time during the classroom teaching process. The edge computing device retrieves the complete lesson plan data for the current class from the classroom teaching platform through an interface protocol. This lesson plan data includes the arrangement of teaching segments in the classroom, the preset start and end times of each teaching segment, the teaching content corresponding to each teaching segment, and the knowledge point information associated with each teaching content. At the same time, the teacher can enter the actual start and end time adjustment information of the teaching segment and the knowledge point association correction information in real time through the touch operation of the teacher's terminal. The edge computing device synchronously receives the correction information and updates the retrieved lesson plan data to ensure that the obtained classroom teaching organization information is completely consistent with the actual classroom teaching process.

[0038] After acquiring classroom teaching organization information, the entire classroom teaching process is broken down into teaching segments according to the criteria for dividing classroom teaching segments. The division criteria are based on the logic and coherence of the teaching content, with the introduction, new knowledge explanation, practice and consolidation, and summary / conclusion segments each divided into independent teaching segments. The new knowledge explanation segment is further subdivided into several sub-segments based on the order in which different knowledge points are explained. Each teaching segment is assigned a unique segment identifier. All teaching segments are arranged sequentially according to their order in the classroom teaching process to form a segment set. After forming the segment set, a mapping relationship is established between each teaching segment and its corresponding knowledge point. Specifically, edge computing devices extract each teaching segment... The teaching content information corresponding to each segment is compared and matched with the knowledge point information in the classroom teaching organization information to determine one or more knowledge points explained in each teaching segment. For a teaching segment explaining a single knowledge point, the segment identifier of the teaching segment is directly associated with the knowledge point identifier of the corresponding knowledge point. For a teaching segment explaining multiple knowledge points, the segment identifier of the teaching segment is associated with the knowledge point identifier of each corresponding knowledge point, and the proportion of explanation time for each knowledge point in the teaching segment is recorded. All association and binding relationships are stored in the local database of the edge computing device to form a mapping relationship table, which facilitates the rapid retrieval of the knowledge points corresponding to each teaching segment in the future.

[0039] After establishing the mapping relationship between teaching segments and knowledge points, the mapping of acquisition time blocks to their respective teaching segments is performed based on the video acquisition time information. The video acquisition time information is obtained through the timestamps embedded synchronously when the camera captures the video stream. Each camera embeds a UTC timestamp accurate to milliseconds in the frame header when capturing video frames. When the edge computing device decodes the video stream frame by frame, it synchronously extracts and records the timestamp information of each video image. Subsequently, acquisition time blocks containing several consecutive frames are defined. Based on the camera's acquisition frame rate of 30 frames per second, each acquisition time block is set to contain 30 consecutive video frames, corresponding to an actual acquisition time of one second. Each acquisition time block is also assigned a unique time block identifier. Simultaneously, based on the timestamps of the first and last video images within each acquisition time block, the actual acquisition time range of each acquisition time block is determined, thus determining the actual acquisition time range of the acquisition time block. After collecting the time range, the edge computing device retrieves the actual start and end time range of each teaching segment stored in the local database. It compares the actual collection time range of each collection time block with the actual start and end time ranges of all teaching segments one by one. If the actual collection time range of a collection time block is entirely within the actual start and end time range of a certain teaching segment, the collection time block is determined to belong to that teaching segment, and the time block identifier of the collection time block is associated with the segment identifier of the teaching segment. If the actual collection time range of a collection time block spans two teaching segments, the teaching segment to which it belongs is the teaching segment corresponding to more than 50% of the frames within the collection time block. If the actual collection time range of a collection time block does not fall within the actual start and end time range of any teaching segment, the collection time block is determined to correspond to a classroom break or invalid collection time, and is not mapped to any teaching segment.

[0040] The association between all collected time blocks and their corresponding teaching segments is stored in a mapping table, which enables precise mapping of each collected time block to its corresponding teaching segment. This allows the eye region image evidence extracted in each time block to be accurately associated with the corresponding teaching segment and knowledge point, providing a precise temporal association basis for evidence retrieval aggregation and scoring based on knowledge graphs.

[0041] In an embodiment of the present invention, the highlight mask and dark nucleus proxy mask of the student's eye region image are extracted, and the index of the symmetrical pupil-masking highlight arc is calculated, including:

[0042] The previously extracted and stored grayscale image of the student's eye region is retrieved. This grayscale image has 256 levels of grayscale, denoted as I. The absolute median difference (MAD) of the eye region grayscale image I is calculated. Specifically, the calculation process involves iterating through all pixels of the eye region grayscale image I, obtaining the grayscale value of each pixel, and calculating the median of all pixel grayscale values, denoted as median(I). Then, the absolute difference between the grayscale value of each pixel and the median(I) is calculated, resulting in a set of absolute differences for all pixels. Finally, the median of this set of absolute differences is calculated. This median is the absolute median difference (MAD) of the eye region grayscale image I. The absolute median difference (MAD) (I) characterizes the grayscale dispersion of the eye region grayscale image I and can effectively suppress the influence of extreme pixels in the highlights or shadows of the eye region image on the grayscale distribution statistics. After calculating the absolute median difference (MAD) (I), the student's eye region image is normalized based on this absolute median difference. The normalization calculation formula is as follows: ,in Let I represent the normalized student eye region image, where I represents the pixel grayscale value of the original student eye region grayscale image, median(I) represents the median of all pixel grayscale values ​​in the original student eye region grayscale image, and MAD(I) represents the absolute median difference of the original student eye region grayscale image. This represents a very small positive number, 0.000001, used to avoid division by zero errors. This value effectively avoids division by zero anomalies that occur when the absolute median difference (MAD) is zero, without significantly affecting the accuracy of the normalization result. This normalization formula maps the gray values ​​of the original eye region grayscale image to a uniform range, suppressing interference caused by fluctuations in the gray values ​​of the eye region image under different lighting conditions, and improving the accuracy of subsequent threshold segmentation.

[0043] After normalization, the normalized eye region image is obtained. Based on this normalized eye region image The Otsu adaptive thresholding segmentation algorithm was used to obtain the specular mask and the dark kernel proxy mask, respectively. When obtaining the specular mask, the normalized eye region image was first analyzed using the Otsu adaptive thresholding segmentation algorithm. The algorithm can automatically find the grayscale histogram that makes the normalized eye region image... The gray value with the largest inter-class variance between the foreground and background regions is used as the specular segmentation threshold, denoted as [value]. Highlight segmentation threshold It can adaptively adapt to the differences in grayscale distribution of images from different eye regions, and then construct a specular mask H(R) using an indicator function. The construction formula is as follows: ,in Indicates a specular mask. This represents an indicator function. The indicator function evaluates to 1 when the condition within the parentheses is true, and evaluates to 0 when the condition within the parentheses is false. This represents the grayscale value of each pixel in the normalized eye region image. This represents the highlight segmentation threshold obtained through the Otsu adaptive thresholding algorithm, which is the gray value of a pixel in the normalized eye region image. Greater than or equal to the highlight segmentation threshold When the value of H(R) for that pixel is 1, the region where that pixel is located is the highlight region in the eye image, corresponding to the bright region formed by reflection from the polarizing screen on the eyeglass lens. When the gray value of a pixel in the normalized eye image is... Less than the highlight segmentation threshold When the value of H(R) for the pixel is 0, the region where the pixel is located is the non-highlight region in the eye image.

[0044] When obtaining the dark nucleus proxy mask, the same Otsu adaptive thresholding segmentation algorithm used to obtain the specular mask is employed. First, the normalized eye region image is processed... The negativeing ​​process is performed to obtain the negativeed normalized eye region image, denoted as . Subsequently, the Otsu adaptive thresholding segmentation algorithm was used to analyze the grayscale histogram of the negativeed normalized eye region image. The algorithm automatically found the grayscale value that maximized the inter-class variance between the foreground and background regions of the negativeed image. This grayscale value was used as the dark kernel segmentation threshold, denoted as... Dark kernel segmentation threshold It can adaptively adapt to the differences in grayscale distribution of the dark nucleus region in different eye region images, and then construct a dark nucleus proxy mask through an indicator function. The construction formula is as follows Where P(R) represents the dark kernel proxy mask, This represents the indicator function, whose value selection rules are consistent with those of the indicator function in specular mask construction. This represents the grayscale value of each pixel in the normalized eye region image. This represents the dark kernel segmentation threshold obtained by processing the negativeed normalized eye region image using the Otsu adaptive thresholding algorithm. The threshold is the gray value of a pixel in the normalized eye region image. Dark kernel segmentation threshold less than or equal to negative When the value of P(R) for that pixel is 1, the region where that pixel is located is the dark nucleus proxy region in the eye image, corresponding to key dark structures in the eye region such as the pupil and the palpebral fissure. When the gray value of a pixel in the normalized eye image... Greater than the negative dark kernel segmentation threshold When the value of P(R) corresponding to the pixel is 0, the region where the pixel is located is the non-dark nucleus proxy region in the eye image. The highlight mask H(R) and dark nucleus proxy mask P(R) obtained through the above process can accurately characterize the highlight region and dark nucleus proxy region in the eye image, providing an accurate image basis for the subsequent calculation of parameters such as pupil ratio and arc band morphology.

[0045] It should be noted that the specular mask is a binary labeling map constructed for the bright pixel regions in the student's eye area image caused by polarization screen reflection or lens specular reflection. It is generated by performing adaptive threshold segmentation on the normalized eye area grayscale image, marking pixels with grayscale values ​​higher than the adaptive threshold as one and marking the remaining pixels as zero, thereby spatially characterizing the coverage range of the lens's specular arc or patch band and using it for subsequent calculations of the area, boundary perimeter, and occlusion relationship with key dark structures. The dark kernel proxy mask is a binary labeling map constructed for the natural dark structures such as the pupil and palpebral fissure in the student's eye area image. It is generated by taking the negative number of the normalized eye area grayscale image and performing adaptive threshold segmentation, marking pixels corresponding to low grayscale values ​​as one and marking the remaining pixels as zero, thereby spatially characterizing the distribution range of the dark kernel region of the eye and using it as a proxy region for the pupil or palpebral fissure for subsequent calculations of the dark kernel area and the intersection area with the specular mask to quantify the degree of pupil occlusion. Among them, polarized screen reflection refers to the fact that the light output from the LCD screen or interactive screen in front of the classroom has obvious polarization properties. When this polarized light shines on the surface of the lens of a student wearing glasses, it will be reflected by the lens surface under the conditions of angle and posture that satisfy the reflection geometry and enter the imaging field of the classroom camera. This will form a bright arc or patch in the eye area image that is significantly brighter than the surrounding skin and eye structure. The intensity of this phenomenon is related to the direction of the screen light emission, the direction of the student's head, the tilt angle of the lens, and the relative position of the camera. Specular reflection of the lens refers to the fact that as a transparent medium with a certain degree of flatness and refractive index difference, the air contact interface of the eyeglass lens will produce an approximate specular reflection effect. When the incident direction of the external light source and the normal of the lens satisfy the angular relationship determined by the law of reflection, the incident light will leave the lens surface in a regular reflection direction and appear as a bright area with clear boundaries and concentrated brightness in the camera image. This may cover key dark structures such as the pupil or palpebral fissure and change the local contrast of the eye area.

[0046] In this embodiment, after obtaining the specular mask and dark kernel proxy mask for the monocular eye region, the monocular pupil occlusion ratio is first calculated. The calculation process begins by performing pixel-level intersection operations on the specular mask and dark kernel proxy mask. For each pixel position in the monocular eye region image, it is determined whether the corresponding pixel value of the specular mask and the dark kernel proxy mask are both 1. If both are 1, the pixel is considered to be a pixel in the intersection region of the two masks. The total number of such intersection region pixels is counted, which is the area of ​​the intersection region between the specular mask and the dark kernel proxy mask. Then, the number of pixels with a pixel value of 1 in the dark kernel proxy mask is counted, which is the area of ​​the dark kernel proxy mask. The formula for calculating the monocular pupil occlusion ratio is: ,in Indicates the pupil coverage ratio of one eye. This represents the area of ​​the intersection region between the specular mask and the dark null proxy mask, and its value is the total number of pixels with a pixel value of 1 in both masks. The value represents the area of ​​the dark kernel proxy mask, which is the total number of pixels with a value of 1 in the dark kernel proxy mask. ε represents a very small positive number of 0.000001 used to avoid division by zero errors. This value will not affect the calculation accuracy of the pupil occlusion ratio and can effectively avoid calculation anomalies in extreme cases. The range of the pupil occlusion ratio for a single eye is 0 to 1. The larger the value, the wider the area of ​​the dark kernel proxy mask covered by the specular mask, that is, the more severe the key structures in the dark part of the eye region are blocked by specular light.

[0047] After calculating the monocular pupil occlusion ratio, the monocular arc band morphological index is calculated. The calculation process begins with connected component analysis of the specular mask. An 8-neighbor connected component labeling algorithm is used to label specular regions with a pixel value of 1 in the specular mask. This algorithm considers eight adjacent pixel positions as connected. It iterates through all pixels in the specular mask, marking the set of mutually connected pixels with a pixel value of 1 as a specular connected region. After labeling, the area of ​​each specular connected region is counted, which is the total number of pixels with a pixel value of 1 in each region. The specular connected region with the largest area is selected as the target specular connected region. The target specular connected region is the bright arc region in the monocular eye region most likely formed by reflection from the polarizing screen. Subsequently, the boundary perimeter and area of ​​this target specular connected region are extracted. The boundary perimeter is calculated using a chain code tracking algorithm, traversing the boundary pixels of the target specular connected region, tracking the boundary pixels sequentially in a clockwise direction, and counting the total number of boundary pixels, which is the boundary perimeter of the target specular connected region. The area of ​​the connected region is the total number of pixels with a value of 1 in the target specular connected region. Simultaneously, pi (π) is introduced, with a value of 3.1416, and the formula for calculating the monocular arc morphological index is as follows: ,in This indicates a morphological index of the unilateral arc band. This represents the perimeter of the boundary of the largest specular connected region, and its value is the total number of pixels on the boundary of this specular connected region. This represents the area of ​​the largest specular connected region, and its value is the total number of pixels with a value of 1 in that specular connected region. Represents pi (π). This index represents a very small positive number used to avoid division by zero errors. It describes the morphological characteristics of the connected regions of the specular highlights. When the specular region is elongated or arc-shaped, the ratio of the square of the boundary perimeter to the area of ​​the connected region is larger, and the index value is higher. When the specular region is blocky or irregularly clustered, the ratio is smaller, and the index value is lower. This allows for the accurate differentiation between symmetrical specular arcs formed by polarizing screen reflections and other unrelated specular regions, providing precise morphological parameter support for the subsequent calculation of the symmetrical pupil-masking specular arc index.

[0048] It should be noted that the monocular pupil occlusion ratio is a proportional indicator used to quantify the degree to which the highlight region in a unilateral eye region obscures key dark structures such as the pupil or palpebral fissure. Its calculation is based on the spatial overlap between the highlight mask and the dark nucleus proxy mask in that unilateral eye region. Specifically, it is obtained by dividing the area of ​​the intersection region of the highlight mask and the dark nucleus proxy mask by the area of ​​the dark nucleus proxy mask, thus reflecting the proportion of the dark nucleus proxy region covered by highlight. A higher monocular pupil occlusion ratio indicates that key dark structures in that unilateral eye are more likely to be obscured by the highlight arc band, leading to a decrease in the visibility of ocular evidence; monocular arc band morphology. The morphological index is a morphological quantification index used to characterize the degree to which the highlight region in a unilateral eye area presents an arc or strip shape. Its calculation preferably uses the largest highlight connected region in the highlight mask as the target region, extracts the boundary perimeter and area of ​​the target region, and uses the ratio of the square of the boundary perimeter to the area of ​​the region as the morphological value, thereby reflecting the thinness and boundary complexity of the highlight region. When the highlight region presents a thin arc, its perimeter is relatively larger than its area, resulting in a higher value for the index. Conversely, when the highlight region presents a blocky white appearance, its perimeter is relatively smaller than its area, resulting in a lower value for the index.

[0049] Calculate the monocular pupil occlusion ratio for both the left and right eyes, as well as the morphological indices of the monocular band for both the left and right eyes. Also calculate the relevant parameters for the monocular pupil occlusion ratios for both eyes. First, calculate the difference between the left and right monocular pupil occlusion ratios, then take the absolute value of this difference. Simultaneously, calculate the sum of the left and right monocular pupil occlusion ratios. The difference ratio is obtained by dividing the absolute value of the difference by the sum. Finally, the left-right symmetry consistency index is obtained by dividing 1 by the difference in the difference ratio. The formula for calculating the left-right symmetry consistency index is as follows: ,in Indicator of left-right symmetry consistency This represents the absolute value of the difference between the monocular pupillary occlusion ratio of the left eye and the monocular pupillary occlusion ratio of the right eye. Indicates the pupil occlusion ratio of the left eye. Indicates the right eye's pupil coverage ratio. This represents a very small positive number used to avoid division by zero errors. The value ranges from 0 to 1. The closer the value is to 1, the more similar the degree of pupil occlusion of the left and right eyes is, the better the consistency of left and right symmetry, and the more it conforms to the characteristics of the symmetrical highlight arc band formed by the reflection of the polarizing screen. The closer the value is to 0, the greater the difference in the degree of pupil occlusion between the left and right eyes, and the less it conforms to the characteristics of this symmetrical artifact.

[0050] The arithmetic mean of the monocular pupil occlusion ratio and the arithmetic mean of the monocular arc band morphological index for both eyes were calculated. The arithmetic mean of the monocular pupil occlusion ratio was calculated by adding the left and right monocular pupil occlusion ratios and then dividing by 2. The arithmetic mean of the monocular arc band morphological index was calculated by adding the left and right monocular arc band morphological indexes and then dividing by 2. Finally, the symmetry pupil occlusion highlight arc band index was obtained by multiplying the left-right symmetry consistency index by the two arithmetic means. This index comprehensively reflects the symmetry of the highlight areas of the left and right eyes and the occlusion... The intensity and morphological characteristics of the pupillary artifacts are considered. When the left and right lenses exhibit symmetrical pupillary obstruction and arc-shaped highlight areas formed by polarization screen reflection, the left-right symmetry consistency index is relatively high. The left-right monocular pupillary obstruction ratio and monocular arc-shaped morphological index are also relatively high. The product of these three factors results in a large index value. When the symmetrical artifact is absent or insignificant, the index value is small. This allows for precise quantification of the intensity of the symmetrical pupillary obstruction highlight arc artifact induced by polarization screen reflection in the eyeglasses, providing a core quantitative basis for the subsequent calculation of evidence reliability weights and eye area visibility evidence values.

[0051] It should be noted that the left-right symmetry consistency index is a symmetry measure used to quantify the degree to which the highlight occlusion effect in the left and right eye areas of a student exhibits symmetrical appearance of the two lenses. Its construction is based on the relative difference between the monocular occlusion ratio of the left and right eyes. Preferably, the absolute value of the difference between the monocular occlusion ratios of the left and right eyes is first calculated, and the sum of these two values ​​is used to form a difference ratio. The smaller the difference ratio, the closer the occlusion ratios of both eyes are. Then, the difference ratio is subtracted from 1 to obtain the left-right symmetry consistency index. The closer the index value is to 1, the more consistent the degree of occlusion between the left and right eyes is, and the more consistent it is with the symmetrical highlight arc phenomenon caused by polarizing screen reflection. The symmetrical highlight arc index is used to comprehensively characterize the polarizing screen. The comprehensive index of the intensity of the symmetrical occlusion highlight arc artifact in the two lenses of the eyeglasses induced by reflection is obtained by multiplying the left-right symmetry consistency index, the arithmetic mean of the monocular occlusion ratio of the left and right eyes, and the arithmetic mean of the monocular arc morphology index of the left and right eyes. The left-right symmetry consistency index is used to emphasize the physical characteristic of the synchronous symmetrical appearance of the two lenses. The arithmetic mean of the monocular occlusion ratio is used to characterize the average degree of occlusion of the highlight on the dark nucleus proxy area. The arithmetic mean of the monocular arc morphology index is used to characterize the degree to which the highlight area is arc-shaped. Thus, the symmetrical occlusion highlight arc index obtains a larger value when both eyes have band-shaped highlights that occlude the key structure of the dark nucleus, and is used as the basis for subsequent calculation of the reliability weight of evidence. The phenomenon of symmetrical highlight arc bands in dual lenses refers to the situation where, when a student wears glasses and faces a polarized display screen or other strong directional light source in front of the classroom, the incident light is regularly reflected at the air contact interface of the left and right lenses and enters the imaging field of the classroom camera. Because the left and right lenses are approximately symmetrical in structure and the student's head orientation and lens posture remain stable for a short period of time, the reflected highlights on the two lenses exhibit a symmetrical distribution in spatial position and shape. Specifically, in the eye area image, arc-shaped or strip-shaped highlight areas with significantly higher brightness than the surrounding areas appear simultaneously on the left and right lenses. Moreover, these highlight areas often extend along the curvature of the lenses to form clearly defined arc bands. They may also cover key dark structures such as the pupil or palpebral fissure, leading to local contrast saturation or reversal in the eye and causing bias in classroom attention assessment based on eye area evidence.

[0052] In embodiments of the present invention, the evidence reliability weight is calculated based on the symmetrical pupillary highlight arc index, and the eye region visibility evidence value is calculated based on the overlap relationship between the highlight mask and the dark nucleus proxy mask, including:

[0053] After calculating the symmetrical pupil-occlusion highlight arc index for a specific time block and camera for the same student, as well as the monocular pupil occlusion ratios for the left and right eyes, the evidence reliability weight and the eye visibility evidence value are calculated. The evidence reliability weight characterizes the credibility of eye behavior evidence and is used to weaken evidence contaminated by artifacts during subsequent scoring aggregation. It is calculated by taking the reciprocal of the sum of 1 and the symmetrical pupil-occlusion highlight arc index. The corresponding formula is: , in the formula Indicates the weight of the reliability of the evidence. This represents the previously calculated symmetrical pupillary highlight arc index, which is used to quantify the intensity of the symmetrical pupillary highlight arc artifact induced by polarizing screen reflection in eyeglasses. The larger the value, the stronger the artifact intensity, the more severe the highlight interference in the eye region image, and the lower the credibility of the corresponding eye region behavioral evidence. At this time, the larger the sum of 1 and the symmetrical pupil-occluded highlight arc index, the smaller its reciprocal, i.e., the evidence reliability weight, means that the weight ratio of the eye region evidence is lower in the subsequent aggregation scoring, thereby achieving automatic weakening of artifact contamination evidence. The smaller the symmetrical pupil-occluded highlight arc index value, the weaker the artifact intensity, the higher the credibility of the eye region evidence. The closer the evidence reliability weight value is to 1, the higher its proportion in the scoring aggregation, ensuring that the uncontaminated valid evidence can play a full role.

[0054] The ocular visibility evidence value is used to characterize the visibility of key structures in the ocular region. It primarily reflects the degree to which key areas such as the pupil and palpebral fissure are obscured by highlights. It is calculated by taking the difference between 1 and the arithmetic mean of the occlusion ratios of the left and right eyes. The corresponding formula is: , in the formula Indicates the visual evidence value of the eye area. Indicates the pupil occlusion ratio of the left eye. This indicates the monocular occlusion ratio of the right eye. Both are previously calculated monocular occlusion ratios, used to characterize the degree to which key structures in the dark area of ​​the monocular eye are obscured by highlights. This represents the arithmetic mean of the pupil occlusion ratios of the left and right eyes. The larger the average value, the more severe the overall obscuration of key structures in the eye region by highlights, and the lower the visibility of key structures in the eye region. In this case, the difference between 1 and this average value, i.e., the eye region visibility evidence value, is smaller. Conversely, the smaller the average value, the higher the visibility of key structures in the eye region, and the closer the eye region visibility evidence value is to 1. The eye region visibility evidence value ranges from 0 to 1. The closer the value is to 1, the better the visibility of key structures in the eye region, and the more effective the corresponding eye region behavioral evidence. The closer the value is to 0, the more severe the obscuration of key structures in the eye region by highlights, and the lower the effectiveness of the corresponding eye region behavioral evidence. This calculation method can accurately quantify the effectiveness of eye region evidence, providing a precise quantitative basis for evidence quality quantification for subsequent knowledge graph-based retrieval aggregation scoring.

[0055] In an embodiment of the present invention, establishing a classroom knowledge graph includes:

[0056] The various nodes in the classroom knowledge graph are instantiated. Student nodes are identified by the "Student" node type, and a unique student node instance is created for each student in the classroom. For example, a "Student_U01" node is created for student number 01, and a "Student_U02" node is created for student number 02. Each student node is configured with basic attributes, including student number, grade, class, and seat number, with attribute values ​​corresponding one-to-one with the actual student information in the classroom. Camera nodes are identified by the "Camera" node type, and a unique camera node instance is created for each camera capturing video streams. For example, a "Student_U01" node is created for student number 01, and a "Student_U02" node is created for student number 02. Each student node is configured with basic attributes, including student number, grade, class, and seat number, with attribute values ​​corresponding one-to-one with the actual student information in the classroom. A Camera_C01 node is created for the camera on the left side of the front of the classroom, and a Camera_C02 node is created for the camera at the back of the classroom. Each camera node is configured with basic attributes, including camera number, installation location, acquisition resolution, acquisition frame rate, and lens angle. The attribute values ​​are consistent with the actual deployment parameters of the camera. Teaching segment nodes are identified by Segment as the node type. A unique teaching segment node instance is created for each completed teaching segment. For example, a Segment_S01 node is created for the teaching segment explaining the concept of a linear equation in one variable in the new knowledge explanation section, and a Segment_S02 node is created for the practice and consolidation section. Each teaching segment node is configured with basic attributes, including segment number, actual start and end times, teaching segment type, and preset duration. These attribute values ​​match the segment information in the classroom teaching organization information. Knowledge point nodes use KnowledgePoint as the node type identifier. A unique knowledge point node instance is created for each knowledge point associated with each teaching segment. For example, a KnowledgePoint_K01 node is created for the concept of a linear equation in one variable, and a KnowledgePoint_K02 node is created for solving a linear equation in one variable. Each knowledge point node is configured with basic attributes, including knowledge point number, knowledge point name, subject, and other attributes. The module hierarchy and attribute values ​​are consistent with the knowledge point system in the curriculum standard. The artifact node uses Artifact as the node type identifier. A unique artifact node instance Artifact_A_POHAI is created for the symmetrical pupil-obscuring high-brightness arc artifact induced by polarized screen reflection in glasses. This artifact node is configured with basic attributes, including artifact type, physical cause description of artifact, and artifact feature description. The artifact type is clearly defined as symmetrical pupil-obscuring high-brightness arc artifact. The physical cause description is formed by the reflection of light output from the polarized LCD screen in front of the classroom through the mirror surface of the student's glasses lens. The artifact feature description is a bright arc band that is symmetrical in position on the left and right lenses and covers the pupil or palpebral fissure area.

[0057] Evidence nodes are identified by the node type "Evidence". A unique evidence node instance is created for each student, each camera, and each time block corresponding to the eye area behavior evidence. For example, an Evidence_E01 node is created for the eye area evidence of student U01, camera C01, and time block t01, and an Evidence_E02 node is created for the eye area evidence of student U01, camera C01, and time block t02. Each evidence node is configured with basic attributes, including evidence number, corresponding time block identifier, and corresponding eye area image storage path. The attribute values ​​are associated with the actual collected eye area image information.

[0058] After instantiating all nodes, create directed edges between various types of nodes. First, create directed edges from evidence nodes to student nodes, with the edge type identified as "ofStudent". These directed edges represent the student subject to which the evidence node belongs. For example, create directed edges from Evidence_E01 to Student_U01 (ofStudent), Evidence_E02 to Student_U01 (ofStudent), and Evidence_E03 to Student_U02 (ofStudent). Second, create directed edges from evidence nodes to camera nodes. The edge type is identified as "fromCamera". This directed edge represents the video capture camera corresponding to the evidence node. For example, a "fromCamera" directed edge is created from Evidence_E01 pointing to Camera_C01, and an "fromCamera" directed edge is created from Evidence_E04 pointing to Camera_C02. Then, a directed edge is created from the evidence node to the teaching segment node. The edge type is identified as "inSegment". This directed edge represents the teaching segment to which the evidence node belongs. For example, an "inSegment" directed edge is created from Evidence_E01 pointing to Segment_S01. A directed edge from Evidence_E02 to Segment_S01 (inSegment) and from Evidence_E05 to Segment_S02 (inSegment) are created. Next, directed edges are created between evidence nodes and artifact nodes, with the edge type identified as `hasArtifact`. These directed edges characterize the artifact type associated with the evidence node. For example, a `hasArtifact` directed edge is created from Evidence_E01 to Artifact_A_POHAI. All evidence nodes containing quantization information for symmetrical pupil-occluded highlight arc artifacts have directed edges pointing to Artifact_A_POHAI. The `hasArtifact` directed edge of `act_A_POHAI` is created; finally, a directed edge is created pointing from the teaching segment node to the knowledge point node, with the edge type identified as `teaches`. This directed edge is used to represent the knowledge point explained in the teaching segment. For example, a directed edge `teaches` is created from `Segment_S01` pointing to `KnowledgePoint_K01`, a directed edge `teaches` is created from `Segment_S02` pointing to `KnowledgePoint_K01`, and a directed edge `teaches` is created from `Segment_S03` pointing to `KnowledgePoint_K02`.

[0059] After creating the directed edges, the node attributes and edge attributes are stored. First, the evidence reliability weight and eye visibility evidence value are stored as numerical attributes of the corresponding evidence nodes, using key-value pairs. The key for the evidence reliability weight is "weight," and the key for the eye visibility evidence value is "value." Then, the symmetric pupil occlusion highlight arc index is stored as the "hasArtifact" directed edge attribute pointing from the evidence node to the artifact node, using key-value pairs. The key is "POHAI," and the value is the calculated symmetric pupil occlusion highlight arc index. If the artifact node attribute storage method is used, the symmetric pupil occlusion highlight arc index corresponding to each evidence node is stored as a key-value pair in the "Artifact_A_POHAI" node. The key is the corresponding evidence node number, and the value is the POHAI value of that evidence node. In this embodiment, it is preferable to store the symmetric pupil occlusion highlight arc index as the "hasArtifact" directed edge attribute, which can directly establish the association between the evidence node and the artifact intensity, facilitating subsequent graph-based retrieval and aggregation calculations.

[0060] It should be noted that the classroom knowledge graph is a graph-structured data organization built for intelligent evaluation of the teaching process. It is used to uniformly represent and manage the learning behavior evidence and teaching process elements collected in the classroom in the form of entities and relationships. It contains several nodes and several directed edges. These nodes include at least evidence nodes, student nodes, camera nodes, teaching segment nodes, knowledge point nodes, and artifact nodes. Evidence nodes represent eye-area observation evidence of a student formed by a camera at a certain acquisition time block and store numerical attributes such as evidence reliability weight and eye-area visibility evidence value. Artifact nodes represent symmetrical pupil-occlusion highlight arc artifacts and, through their association with evidence nodes, carry the symmetrical pupil-occlusion highlight arc. With an index, teaching segment nodes are used to represent segments in the classroom process such as explanation, demonstration, questioning, practice, and comments, and establish pointing relationships with knowledge point nodes to express the teaching content corresponding to the segment. Student nodes and camera nodes are used to identify the evidence source object and the acquisition device object. Directed edges are used to express the attribution relationship between evidence nodes and student nodes, camera nodes, teaching segment nodes, and artifact nodes, as well as the teaching correspondence relationship between teaching segment nodes and knowledge point nodes. This enables the system to quickly retrieve the set of evidence nodes pointing to the target teaching segment through graph query and perform weighted aggregation calculation based on evidence reliability weights on the graph to obtain the segment attention score. At the same time, it can trace the source of artifact contamination along the graph relationship to support interpretable output results.

[0061] In an embodiment of the present invention, a set of evidence nodes pointing to a target teaching segment is retrieved from a classroom knowledge graph. A weighted average calculation is performed on the eye visibility evidence value based on the evidence reliability weight of each evidence node in the evidence node set to obtain a segment attention score. Tracing link information is then generated, including:

[0062] The target teaching segment is identified and its node instances are defined. Any instantiated teaching segment in the classroom knowledge graph is selected as the target teaching segment. For example, the teaching segment node Segment_S01 explaining the concept of a linear equation in one variable is selected as the target teaching segment. Then, the set of evidence nodes pointing to this target teaching segment is retrieved from the classroom knowledge graph. The retrieval process is achieved by traversing all directed edges of type inSegment in the classroom knowledge graph, because inSegment directed edges are the associated edges from evidence nodes to their respective teaching segment nodes. During the retrieval, each inSegment directed edge is checked to determine whether its endpoint is the target teaching segment node Segment_S01. If the endpoint of a directed edge inSegment is Segment_S01, then the starting point of the directed edge is the evidence node pointing to the target teaching segment. All such evidence nodes are collected to form a set of evidence nodes pointing to the target teaching segment Segment_S01. Each evidence node in this set corresponds to eye behavior evidence of a student, a camera, or a time block within the target teaching segment. For example, if evidence nodes Evidence_E01, Evidence_E02, and Evidence_E03 all point to Segment_S01 through the directed edge inSegment, then the set of evidence nodes contains the above three evidence nodes.

[0063] After retrieving the set of evidence nodes, the weighted cumulative visibility value is calculated. The calculation process involves iterating through each evidence node in the set, retrieving the stored numerical attributes of each node, namely the evidence reliability weight and the eye area visibility evidence value. The evidence reliability weight and eye area visibility evidence value of each node are multiplied together to obtain the weighted visibility value for each node. Then, the weighted visibility values ​​of all evidence nodes in the set are summed, and the sum is the weighted cumulative visibility value. The calculation logic is that the weighted visibility value of each evidence node is a quantitative representation of the validity of that node's evidence; the larger the weight and the higher the evidence value, the greater the contribution of that node to the segment score. The summation comprehensively reflects the overall visibility level of all valid eye area evidence within the target teaching segment. Subsequently, the cumulative reliability weight value is calculated. The calculation process also involves iterating through each evidence node in the set, retrieving the evidence reliability weight of each node, and multiplying the evidence reliability weight of each node. The reliability weights of the evidence points are accumulated, and the result is the cumulative reliability weight value. This cumulative value is used to normalize the cumulative weighted visibility value to avoid incomparability of scores due to differences in the number of evidence points. Finally, the ratio of the cumulative weighted visibility value to the cumulative reliability weight value is used as the segment attention score of the target teaching segment. The segment attention score Q ranges from 0 to 1. The closer the value is to 1, the higher the overall attention level of students in the target teaching segment, the lower the degree of artifact contamination of eye area evidence, and the better the effectiveness of the teaching segment. The closer the value is to 0, the lower the overall attention level of students in the target teaching segment, the higher the degree of artifact contamination of eye area evidence, or the poorer the students' true attention state. This score, through weighted aggregation, automatically weakens the influence of evidence points severely contaminated by artifacts, strengthens the contribution of effective evidence, and ensures that the score result can accurately reflect the true teaching effect and students' attention state of the target teaching segment.

[0064] Using the set of evidence nodes corresponding to the target teaching segment as the core, which consists of all evidence nodes pointing to the target teaching segment obtained previously, the algorithm iterates through each evidence node in the set and retrieves the symmetric occlusion highlight arc index stored in the directed edge attribute of each evidence node pointing to the artifact node. The symmetric occlusion highlight arc index value is the quantified value of the artifact intensity of the corresponding evidence node. The larger the value, the more severe the artifact contamination of the eye area evidence corresponding to the evidence node. Then, the arithmetic mean of the symmetric occlusion highlight arc index values ​​of all evidence nodes in the evidence node set is calculated. This arithmetic mean is the segment contamination profile of the target teaching segment. The segment contamination profile ranges from 0 to 1. The closer the value is to 1, the more severe the overall eye area evidence in the target teaching segment is contaminated by the symmetric occlusion highlight arc artifact. The lower the segment attention score is more likely to be caused by artifacts rather than a decrease in students' actual attention. The closer the value is to 1, the less severe the artifact contamination of the eye area evidence in the target teaching segment. The segment attention score can truly reflect the students' overall attention status.

[0065] After generating a fragment contamination profile, a traceability link is further generated. The traceability link is used to locate the specific source of fragment contamination, clarify the cause of abnormal fragment attention scores, and achieve interpretability and traceability of the evaluation results. First, the contribution of each evidence node in the evidence node set is calculated. The contribution is used to characterize the degree of influence of a single evidence node on the attention score of the target teaching fragment. The calculation method is to multiply the evidence reliability weight of each evidence node by the eye area visibility evidence value to obtain the contribution of the evidence node. The smaller the contribution value, the smaller the positive contribution of the evidence node to the fragment attention score, and the more serious the artifact contamination of the corresponding eye area evidence or the lower the evidence validity.

[0066] All evidence nodes in the evidence node set are sorted in ascending order of contribution. A number of the top-ranked evidence nodes are selected as key targets for further investigation, up to 20% of the total number of evidence nodes. If fewer than one node is selected, only one node is chosen. These key targets are the evidence nodes that have the least impact on the segment attention score and are most likely to be most severely affected by artifacts. Next, for each key target, the associated nodes and attribute information are retrieved from the classroom knowledge graph. The student node corresponding to the evidence node is retrieved via the `ofStudent` directed edge to obtain basic attributes such as student ID and seat number, identifying the students most susceptible to artifacts. The camera node corresponding to the evidence node is retrieved via the `fromCamera` directed edge to obtain basic attributes such as camera ID and installation location, identifying potential camera positions with artifact interference. Finally, the symmetrical pupil occlusion highlight arc index value corresponding to the evidence node is retrieved via the `hasArtifact` directed edge to determine the artifact contamination of that evidence node. The retrieved information is then organized into a traceability link based on a fixed logic. This logic arranges the key traceability objects in ascending order of contribution, with each key traceability object corresponding to a sub-link. Each sub-link includes an evidence node number, corresponding student information, corresponding camera information, and a corresponding symmetrical pupil-occluded highlight arc index value. The sum of all sub-links constitutes the traceability link for the target teaching segment. This link clearly determines whether the low score of the target teaching segment is due to artifact interference from a specific camera position, eyeglass reflection artifacts from a specific student group, or a decrease in students' actual attention. If the symmetrical pupil-occluded highlight arc index values ​​of the key traceability objects in the traceability link are all large, it indicates that the low score is mainly caused by artifact contamination, excluding the influence of teacher teaching quality and students' actual cognitive state. If the symmetrical pupil-occluded highlight arc index values ​​of the key traceability objects are all small, it indicates that the low score is mainly caused by a decrease in students' actual attention. This provides precise positioning for teaching optimization, enabling traceable and interpretable evaluation of the teaching process.

[0067] like Figure 2 The diagram shown is a functional block diagram of an intelligent evaluation system for the teaching process based on learning behavior analysis, provided by an embodiment of the present invention.

[0068] In this embodiment, the functions of each module / unit are as follows:

[0069] The data acquisition module is used to acquire video streams from cameras in the classroom and perform face detection and tracking, extract images of students' eye areas, and form acquisition time blocks from several consecutive frames to establish a mapping between acquisition time blocks and teaching segments.

[0070] The artifact index module is used to extract the highlight mask and dark nucleus proxy mask of the student's eye area image and calculate the symmetrical pupil occlusion highlight arc index.

[0071] The evidence parameter module is used to calculate the evidence reliability weight based on the symmetrical pupil-masking highlight arc index and to calculate the eye region visibility evidence value based on the overlap relationship between the highlight mask and the dark nucleus proxy mask.

[0072] The knowledge graph module is used to build a classroom knowledge graph, which includes evidence nodes, artifact nodes, student nodes and teaching segment nodes. The evidence reliability weight and the eye area visibility evidence value are written into the attribute fields of the evidence nodes, and the symmetrical pupil occlusion highlight arc index is associated with the artifact nodes.

[0073] The aggregation scoring module is used to retrieve a set of evidence nodes pointing to the target teaching segment from the classroom knowledge graph. Based on the evidence reliability weight of each evidence node in the evidence node set, a weighted average calculation is performed on the visual visibility evidence value to obtain the segment attention score and generate traceability link information.

[0074] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A learning behavior analysis-based intelligent evaluation method for a teaching process, characterized by, include: S1. Collect video streams from classroom cameras and perform face detection and tracking, extract images of students' eye areas, and construct a collection time block from several consecutive frames to establish a mapping between the collection time block and the teaching segment. S2. Extract the highlight mask and dark nucleus proxy mask from the student's eye region image, and calculate the symmetrical pupil-masking highlight arc index; S3. Calculate the evidence reliability weight based on the symmetrical pupil-masking highlight arc index, and calculate the eye region visibility evidence value based on the overlap relationship between the highlight mask and the dark nucleus proxy mask. S4. Establish a classroom knowledge graph, which includes evidence nodes, artifact nodes, student nodes and teaching segment nodes. Write the evidence reliability weight and the eye zone visibility evidence value into the attribute field of the evidence node, and associate the symmetrical pupil occlusion highlight arc index with the artifact node. S5. Retrieve the set of evidence nodes pointing to the target teaching segment from the classroom knowledge graph. Perform a weighted average calculation on the visual visibility evidence value based on the evidence reliability weight of each evidence node in the evidence node set to obtain the segment attention score and generate traceability link information. 2.The intelligent evaluation method of teaching process based on learning behavior analysis according to claim 1, characterized in that, Establish a mapping between time blocks of data collection and teaching segments, specifically including: Obtain classroom teaching organization information, form a collection of teaching segments, and establish a mapping relationship between each teaching segment and the corresponding knowledge point; Based on the video capture time information, capture time blocks containing several consecutive frames are mapped to their respective teaching segments. 3.The intelligent evaluation method of teaching process based on learning behavior analysis of claim 1, wherein, Extracting the highlight and shadow regions from the student's eye area image, specifically including: Calculate the absolute median difference of the student's eye region image, and normalize the student's eye region image based on the absolute median difference; Based on the normalized image, highlight masks and dark kernel proxy masks are obtained through adaptive threshold segmentation.

4. The intelligent assessment method for the teaching process based on learning behavior analysis according to claim 3, characterized in that, Calculating the symmetrical occlusion specular arc index specifically includes: The ratio of the area of ​​the intersection region between the specular mask and the dark kernel proxy mask to the area of ​​the dark kernel proxy mask is calculated to obtain the monocular pupil occlusion ratio. In the specular mask, the largest specular connected region is selected, and the boundary perimeter and area of ​​the specular connected region are extracted. The ratio of the square of the boundary perimeter to the area of ​​the connected region is calculated to obtain the morphological index of the monocular arc band. Calculate the absolute value of the difference between the monocular pupil occlusion ratio of the left eye and the monocular pupil occlusion ratio of the right eye, and the sum of the monocular pupil occlusion ratios of the left eye and the right eye. Obtain the difference ratio based on the ratio of the absolute value of the difference to the sum. The left-right symmetry consistency index is obtained based on the difference between 1 and the difference ratio. The symmetrical pupillary highlight arc index is obtained by multiplying the left-right symmetry consistency index, the arithmetic mean of the left and right monocular pupillary occlusion ratio, and the arithmetic mean of the left and right monocular arc morphology index.

5. The intelligent assessment method for the teaching process based on learning behavior analysis according to claim 4, characterized in that, Evidence reliability weights are calculated based on the symmetrical pupillary highlight arc index, and ocular visibility evidence values ​​are calculated based on the overlap between the highlight mask and the dark nucleus proxy mask, specifically including: The reciprocal of the sum of 1 and the index of the symmetrical pupil-blocking highlight arc is used as the weight of the evidence reliability. The difference between 1 and the arithmetic mean of the pupil occlusion ratios of the left and right eyes is used as the visual evidence value for the eye region.

6. The intelligent assessment method for the teaching process based on learning behavior analysis according to claim 1, characterized in that, Establishing a classroom knowledge graph includes: Instantiate evidence nodes, student nodes, camera nodes, teaching segment nodes, knowledge point nodes, and artifact nodes; Create directed edges from the evidence node to the student node, camera node, teaching segment node, and artifact node, respectively; Create directed edges from teaching segment nodes to knowledge point nodes; Store the evidence reliability weight and the eye visibility evidence value as numerical attributes of the evidence node; Store the index of the symmetrical pupil-masking specular arc as an edge attribute of the evidence node pointing to the artifact node or an attribute of the artifact node.

7. The intelligent assessment method for the teaching process based on learning behavior analysis according to claim 1, characterized in that, The visual area evidence value is calculated by weighting the evidence reliability weights of each evidence node in the evidence node set. Specifically, this includes: The weighted cumulative visibility value is obtained by summing the products of the evidence reliability weights of all evidence nodes in the evidence node set and the eye visibility evidence values. The reliability weights of all evidence nodes in the evidence node set are summed to obtain the cumulative reliability weight value. The ratio of the weighted cumulative visibility value to the cumulative cumulative reliability weight value is used as the segment attention score.

8. The intelligent assessment method for the teaching process based on learning behavior analysis according to claim 1, characterized in that, Generate traceability information, including: Calculate the artifact contribution value of each evidence node in the evidence node set, wherein the artifact contribution value is the product of the evidence reliability weight and the eye area visibility evidence value; The artifact contribution values ​​are sorted to generate traceability link information, which includes the camera node identifier associated with the evidence node, the student node identifier, and the corresponding symmetrical pupil occlusion specular arc index.

9. A teaching process intelligent evaluation system based on learning behavior analysis, applied in the teaching process intelligent evaluation method based on learning behavior analysis as described in any one of claims 1-7, characterized in that, The system includes: The data acquisition module is used to acquire video streams from cameras in the classroom and perform face detection and tracking, extract images of students' eye areas, and form acquisition time blocks from several consecutive frames to establish a mapping between acquisition time blocks and teaching segments. The artifact index module is used to extract the highlight mask and dark nucleus proxy mask of the student's eye area image and calculate the symmetrical pupil occlusion highlight arc index. The evidence parameter module is used to calculate the evidence reliability weight based on the symmetrical pupil-masking highlight arc index and to calculate the eye region visibility evidence value based on the overlap relationship between the highlight mask and the dark nucleus proxy mask. The knowledge graph module is used to build a classroom knowledge graph, which includes evidence nodes, artifact nodes, student nodes and teaching segment nodes. The evidence reliability weight and the eye area visibility evidence value are written into the attribute fields of the evidence nodes, and the symmetrical pupil occlusion highlight arc index is associated with the artifact nodes. The aggregation scoring module is used to retrieve a set of evidence nodes pointing to the target teaching segment from the classroom knowledge graph. Based on the evidence reliability weight of each evidence node in the evidence node set, a weighted average calculation is performed on the visual visibility evidence value to obtain the segment attention score and generate traceability link information.