Student concentration real-time monitoring method and system based on classroom scene

By acquiring binocular image sequences through high-frequency frame extraction, analyzing pupil data, and decoupling saccade events using a preset saccade judgment threshold, a multi-dimensional eye movement feature model is constructed. This solves the problem of misjudgment in student attention assessment in classroom scenarios, and enables accurate attention monitoring and personalized intervention suggestions.

CN122090386BActive Publication Date: 2026-07-31ZHEJIANG YUNXIAOJIA NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG YUNXIAOJIA NETWORK TECHNOLOGY CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify students' attention levels in classroom settings, especially in distinguishing between rapid saccades and microsaccades. They are also susceptible to environmental interference, leading to misjudgments in attention assessments and an inability to generate personalized intervention recommendations.

Method used

By acquiring binocular image sequences through high-frequency frame extraction, analyzing pupil data, decoupling saccade events using a preset saccade judgment threshold, statistically analyzing frequency and directional entropy, and combining temporal interval variation coefficients, a multi-dimensional eye movement feature model is constructed to generate attention state labels and focus levels.

Benefits of technology

It achieves precise decoupling between rapid and micro-saccades, eliminates environmental interference, accurately assesses focus, distinguishes between latent distractions, generates personalized intervention suggestions, and improves the accuracy of attention monitoring and the targeting of personalized interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090386B_ABST
    Figure CN122090386B_ABST
Patent Text Reader

Abstract

This invention relates to the field of visual monitoring, and more particularly to a method and system for real-time monitoring of student attention in a classroom setting. The method includes: performing high-frequency frame extraction and binocular region cropping on a student's facial video stream to obtain a binocular sub-image sequence; analyzing the visual-motor features of the binocular image sequence to obtain binocular pupil data; decoupling the binocular pupil data based on a preset saccade judgment threshold to obtain rapid saccade events and micro-saccade events, and statistically analyzing the frequencies of the two types of events to obtain rapid saccade frequency and micro-saccade frequency; evaluating the directional divergence of micro-saccade events to obtain micro-saccade directional entropy, and analyzing the temporal interval dispersion of rapid saccade events to obtain temporal interval variation coefficient; mapping the micro-saccade directional entropy and temporal interval variation coefficient in a state space to obtain an attention state label; comprehensively evaluating the attention state label, rapid saccade frequency, and micro-saccade frequency to obtain an attention level, and generating latent distraction intervention suggestions based on the attention level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a method and system for real-time monitoring of student attention in a classroom setting. Background Technology

[0002] Currently, video analytics-based student classroom attention assessment technologies have been developed to some extent. Common methods include: acquiring student facial videos through image acquisition devices, using facial recognition algorithms to locate the eye area, and then tracking the pupil center or fixation point. These technologies typically focus on whether students are looking directly at the blackboard or the teacher. By statistically analyzing the proportion of time students look up and down, or by analyzing basic eye movement parameters such as blink frequency and fixation duration, their attention state can be inferred. Existing technologies set a fixed fixation area, such as the blackboard area. When the time the student's pupil center coordinates fall within this area is less than a preset threshold, it is determined that the student is not paying attention.

[0003] The decoupling of saccade events relies solely on a single velocity threshold for global classification, failing to model the differences between rapid saccades and their temporal intervals, and between micro-saccades and directional divergence. For example, in a classroom environment, head movements and lighting fluctuations can easily trigger spurious velocity peaks, leading to the confusion between rapid and micro-saccades. Furthermore, it cannot quantify the spatial divergence of micro-saccades using directional entropy or characterize the temporal variability of rapid saccades using the coefficient of variation. The single feature dimension makes it difficult to depict the fine eye movement patterns of latent inattention. Moreover, attention assessments often use single temporal statistics such as saccade frequency to directly map levels, without jointly mapping the directional entropy of micro-saccades and the coefficient of variation of the temporal intervals of rapid saccades into a state space. For instance, some students have high-frequency micro-saccades but a focused direction, while some students have disordered rapid saccade intervals but normal frequency. Single indicators are prone to misjudgment, making it impossible to accurately distinguish between focus and latent inattention, and thus difficult to support the generation of personalized intervention suggestions. Summary of the Invention

[0004] This invention provides a method and system for real-time monitoring of student concentration in classroom scenarios. It addresses the technical problems of existing technologies, such as coarse classification of saccade events, single feature dimensions, susceptibility to environmental interference, misjudgment of concentration, and inability to accurately identify latent distraction and generate personalized interventions. The goal is to accurately decouple saccade events, construct a multi-dimensional eye movement feature model, eliminate environmental interference, accurately assess concentration, distinguish latent distraction, and generate personalized intervention suggestions.

[0005] In the first aspect, a method for real-time monitoring of student attention includes: S1. Perform high-frequency frame extraction and eye region cropping on the student's facial video stream to obtain the student's eye sub-image sequence. S2. Analyze the visual-motor features of the binocular image sequence to obtain the student's binocular pupil data; S3. Based on the preset saccade judgment threshold, the saccade events of the pupil data of both eyes are decoupled to obtain the student's rapid saccade events and micro saccade events. The frequency of rapid saccade events and micro saccade events is statistically analyzed to obtain the student's rapid saccade frequency and micro saccade frequency. S4. Evaluate the directional divergence of micro-saccade events to obtain the micro-saccade directional entropy of students, and perform temporal interval dispersion analysis on students' rapid saccade events to obtain the temporal interval variation coefficient of students. S5. Map the micro-saccade direction entropy and the temporal interval variation coefficient in the state space to obtain the student's attention state label. S6. The attention state labels, rapid saccade frequency, and microsaccade frequency are comprehensively evaluated to obtain the student's attention level, and intervention suggestions for the student's latent inattention are generated based on the attention level.

[0006] Preferably, the step of performing high-frequency frame extraction and binocular region cropping on the student's facial video stream to obtain the student's binocular sub-image sequence includes: High-frequency frame sampling is performed on the student's facial video stream to obtain discrete frame images of the student, and the facial regions of the discrete frame images are detected to obtain the student's facial localization box. The pupil regions of both eyes are located using the facial positioning bounding box to obtain the coordinates of the student's left and right eye regions of interest. Based on the coordinates of the region of interest for the left and right eyes, the eye regions of the discrete frames are cropped to obtain the student's right eye sub-image and left eye image. The left and right eye images are associated and encapsulated to obtain the student's double eye image sequence.

[0007] Preferably, the step of analyzing the visual-motion features of the binocular image sequence to obtain the student's binocular pupil data includes: Pupil center localization was performed on the left and right eye images in the binocular image sequence to obtain the coordinates of the student's left and right pupil centers. Based on the coordinates of the left and right pupil centers, pupil displacement is calculated for the left and right pupil centers of adjacent frames in the binocular image sequence, resulting in the student's left eye vertical displacement sequence, right eye vertical displacement sequence, left eye horizontal displacement sequence, and right eye horizontal displacement sequence. Based on the left eye horizontal displacement sequence and the left eye vertical displacement sequence, the velocity component measurement of the left eye displacement between adjacent frames in the binocular sub-image sequence is performed to obtain the student's left eye horizontal velocity sequence and left eye vertical velocity sequence. Based on the right eye horizontal displacement sequence and the right eye vertical displacement sequence, the displacement velocity of the right eye between adjacent frames in the binocular sub-image sequence is analyzed to obtain the student's right eye horizontal velocity sequence and the student's right eye vertical velocity sequence. The velocity components of the left eye horizontal velocity sequence, left eye vertical velocity sequence, right eye horizontal velocity sequence, and right eye vertical velocity sequence are aggregated to obtain the pupil data of the student's two eyes.

[0008] Preferably, the step of decoupling saccade events from the pupil data of both eyes based on a preset saccade determination threshold to obtain the student's rapid saccade events and micro-saccade events includes: Instantaneous velocity calculations were performed on the pupil data of both eyes to obtain the instantaneous velocity amplitude sequences of the student's right eye and left eye. The instantaneous velocity amplitude sequences of the left eye and the right eye are compared with the upper limit threshold of the velocity threshold in the preset saccade judgment threshold to obtain the student's left eye velocity overshoot time set and right eye velocity overshoot time set. By performing binocular intersection discrimination on the left eye velocity overshoot time set and the right eye velocity overshoot time set, the candidate time sequence of the student's eye saccade is obtained; Based on the candidate moments in the candidate moment sequence of eye sacs, eye sacs segments are extracted from the pupil data of both eyes to obtain the student's eye sacs event interval set; Based on the amplitude threshold in the preset saccade judgment threshold, the saccade event interval set is judged to determine the event category, and the student's rapid saccade events and micro saccade events are obtained.

[0009] Preferably, the step of performing frequency statistics on rapid saccade events and microsaccade events respectively to obtain the student's rapid saccade frequency and microsaccade frequency includes: The rapid saccade event and the micro saccade event are divided into fixed windows to obtain continuous equal-length window segments for students. Count the rapid saccade events and micro saccade events within a continuous equal-length window segment to obtain the original count value of the continuous equal-length window segment; The original count values ​​were processed by moving average to obtain the smoothed rapid saccade frequency and smoothed micro saccade frequency of students; By aligning the smoothed rapid saccade frequency with the smoothed micro saccade frequency in time, the student's rapid saccade frequency and the student's micro saccade frequency are obtained.

[0010] Preferably, the step of evaluating the directional divergence of microsaccade events to obtain the student's microsaccade direction entropy includes: The direction angle of the microsaccade event is extracted to obtain the microsaccade angle of the student. Discretize the entire range of microsaccade angles to obtain the microsaccade direction angle range for students; Based on the microsaccade direction angle interval, the frequency of microsaccade angles falling within the interval is statistically analyzed to obtain the frequency of the student's microsaccade direction angle value within the interval of the microsaccade direction angle. Based on the frequency of occurrence of intervals, the uncertainty of the spatial divergence of microsaccade direction is quantitatively evaluated to obtain the microsaccade direction entropy of students.

[0011] Preferably, the step of performing temporal interval dispersion analysis on the student's rapid eye saccade events to obtain the student's temporal interval variation coefficient includes: By correlating adjacent moments of rapid eye saccades, the student's adjacent event moment pairs can be obtained. Extract the time difference between adjacent event time pairs to obtain the time interval value of adjacent event time pairs; The mean value of the students' rapid saccade intervals is obtained by performing mean measurement on the time interval value sequence. Deviation amplitude analysis was performed on the time interval values ​​to obtain the standard deviation of the students' rapid saccade intervals; Based on the standard deviation and mean of rapid saccade intervals, the relative fluctuation of time interval values ​​is mapped by the rate of variation to obtain the coefficient of variation of students' time intervals.

[0012] Preferably, the step of mapping the micro-saccade direction entropy with the temporal interval variation coefficient to obtain the student's attention state label includes: The entropy of micro-eye sacral direction is interval-calibrated to obtain the entropy value interval label for students; Based on the preset coefficient of variation level threshold, the coefficient of variation of time intervals is classified into levels to obtain the coefficient of variation level of students. By combining the entropy value interval label with the coefficient of variation level, a combined label of the student's status is obtained. The student's state combination identifier is mapped to obtain the student's attention state label.

[0013] Preferably, the step of comprehensively evaluating attention state tags, rapid saccade frequency, and microsaccade frequency to obtain the student's attention level, and generating intervention suggestions for the student's latent inattention based on the attention level, includes: Multi-source feature aggregation was performed on attention state labels, rapid saccade frequency, and microsaccade frequency to obtain students' attention characteristics. The students' attention levels are obtained by classifying the attention characteristics into different levels. Based on the pre-set intervention suggestions, the students' attention levels are matched with relevant measures to obtain candidate intervention measures for the students; By assembling candidate intervention measures for students into intervention texts, intervention suggestions for students' latent inattentiveness were obtained.

[0014] Secondly, a real-time student attention monitoring system includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a real-time student attention monitoring method is implemented.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By calculating the instantaneous velocity of pupil data from both eyes and combining preset velocity and amplitude thresholds, precise decoupling of saccade events is achieved. Simultaneously, direction angle extraction, interval frequency statistics, and directional entropy quantification are performed on micro-saccades, and temporal interval calculation and coefficient of variation are calculated for rapid saccades. A multi-dimensional eye movement feature model is built, which can effectively eliminate the interference of velocity pseudo-peaks caused by head micro-movements and light fluctuations in the classroom environment, avoid the confusion between rapid and micro-saccades, make up for the shortcomings of a single feature dimension, and can finely depict the exclusive eye movement pattern of latent inattention.

[0016] 2. By jointly mapping the micro-saccade direction entropy and the coefficient of variation of the temporal interval of rapid saccades into the state space, attention state labels are obtained. Then, by combining the rapid saccade frequency and micro-saccade frequency, multi-source feature aggregation and attention level labeling are completed. This breaks the limitation of relying solely on a single temporal statistical quantity to map the level, solves the problem of attention misjudgment that is easily caused by a single indicator, and achieves accurate differentiation between attention and latent inattention. This supports the generation of personalized latent inattention intervention suggestions tailored to individual students.

[0017] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description

[0018] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic flowchart of a method for real-time monitoring of student focus in a classroom setting according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a real-time student attention monitoring system based on a classroom setting, according to an embodiment of the present invention. Detailed Implementation

[0019] The following reference Figures 1 to 2This invention describes a method and system for real-time monitoring of student attention in a classroom setting, according to embodiments of the present invention. In this description, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature, that is, include one or more of that feature. In the description of this invention, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. When a feature "includes or contains" one or more of the features it encompasses, unless otherwise specifically described, this indicates that other features are not excluded and may be further included.

[0020] In the description of this embodiment, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0021] This embodiment provides a method for real-time monitoring of student concentration in classroom scenarios. Under the premise that the complex classroom environment (head micro-movement, light fluctuation) is prone to interference and eye movement features are single and prone to misjudgment, this method can achieve precise decoupling of rapid eye saccades and micro-saccades, quantification of multi-dimensional eye movement features, fine mapping of attention state, accurate assessment of concentration level, intelligent identification of latent inattention and generation of personalized intervention suggestions.

[0022] Please see Figure 1 , Figure 1 This is a schematic flowchart of a method for real-time monitoring of student attention in a classroom setting according to an embodiment of the present invention. The method generally includes: S1. Perform high-frequency frame extraction and eye region cropping on the student's facial video stream to obtain the student's eye sub-image sequence; S2. Analyze the visual motion features of the binocular sub-image sequence to obtain the pupil data of the student's eyes; S3. Based on a preset saccade determination threshold, the saccade events of the pupil data of both eyes are decoupled to obtain the student's rapid saccade events and micro saccade events. The frequency of the rapid saccade events and the micro saccade events is statistically analyzed to obtain the student's rapid saccade frequency and micro saccade frequency. S4. Evaluate the directional divergence of the micro-saccade events to obtain the micro-saccade directional entropy of the student, and perform temporal interval dispersion analysis on the student's rapid saccade events to obtain the student's temporal interval variation coefficient. S5. Map the micro-saccade direction entropy with the temporal interval variation coefficient in a state space to obtain the student's attention state label. S6. The attention state label, the rapid saccade frequency, and the microsaccade frequency are comprehensively evaluated to obtain the student's attention level, and intervention suggestions for the student's latent inattention are generated based on the attention level.

[0023] In step S1 above, a continuous video stream of students' faces can be continuously acquired by image acquisition devices deployed in the classroom; static images are extracted one by one from the continuous video stream at fixed very short time intervals to form discrete frame images; all pixels of the discrete frame images are traversed one by one, and the brightness and color values ​​of each pixel are compared with the pre-set standard brightness range and standard color range of human facial skin, and the pixels belonging to the student's face are selected and summarized. The system scans facial pixels in the horizontal and vertical directions of the image, recording the leftmost, rightmost, topmost, and bottommost pixel positions. These four positions form the smallest rectangular area as the student's facial positioning box. Within the facial positioning box, the system calculates the total brightness difference between each pixel and its adjacent pixels in the left, right, and above. The set of pixels whose total brightness difference exceeds a preset stable limit is determined as the eye contour position. Within the eye contour position, the system selects the continuous circular distribution of pixels with the smallest brightness value as the pupil region. The system determines the extreme pixel positions in the horizontal and vertical directions of the pupil region, forming the coordinates of the left and right eye regions of interest. Based on the coordinates of the left and right eye regions of interest, the system extracts the image content of the corresponding pixel range from discrete frames, forming the left eye sub-image and the right eye sub-image. The left eye sub-image and the right eye sub-image of the same discrete frame are bound into a single-frame binocular image unit according to the acquisition time. All single-frame binocular image units are then arranged in chronological order of video acquisition time to form the student's binocular image sequence.

[0024] In step S2 above, the pixel brightness values ​​of each frame of the left and right eye images in the binocular image sequence can be read. Taking each pixel as the center, select the top, bottom, left, right and four diagonally adjacent pixels, and filter the pixels whose brightness is lower than the average brightness of the surrounding pixels to form the pupil pixel region. Find the midpoint between the leftmost and rightmost points in the horizontal direction and the midpoint between the top and bottom points in the vertical direction of the pupil pixel region. The intersection of the two midpoints forms the pupil center position. Record the horizontal and vertical information of this position to obtain the coordinates of the left eye pupil center and the right eye pupil center. The horizontal displacement of the pupil in adjacent frames is obtained by subtracting the horizontal displacement of the pupil center in the previous frame from the horizontal displacement of the pupil center in the next frame; the vertical displacement of the pupil in adjacent frames is obtained by subtracting the vertical displacement of the pupil center in the previous frame from the vertical displacement of the pupil center in the next frame; the horizontal and vertical displacements of all adjacent frames are arranged in chronological order to form the left eye horizontal displacement sequence, left eye vertical displacement sequence, right eye horizontal displacement sequence, and right eye vertical displacement sequence; the horizontal displacement of the pupil in adjacent frames is divided by the acquisition time interval between two adjacent frames of the binocular image sequence to obtain the horizontal motion velocity of the pupil; the vertical displacement of the pupil in adjacent frames is divided by the acquisition time interval between two adjacent frames of the binocular image sequence to obtain the vertical motion velocity of the pupil; the horizontal and vertical motion velocities of all adjacent frames are arranged in chronological order to form the left eye horizontal velocity sequence, left eye vertical velocity sequence, right eye horizontal velocity sequence, and right eye vertical velocity sequence; the four velocity values ​​at the same acquisition time node are matched and integrated according to the time dimension to form continuous temporal velocity data, thus obtaining the student's binocular pupil data.

[0025] In step S3 above, the horizontal and vertical pupil velocity at the same time point can be squared and added together, and then the square root of the sum can be taken to obtain the instantaneous pupil velocity amplitude at that time point. The instantaneous pupil velocity amplitudes of all time points are arranged in chronological order to form the instantaneous velocity amplitude sequence of the left eye and the instantaneous velocity amplitude sequence of the right eye. Each value in the instantaneous velocity amplitude sequences of the left and right eyes is compared with the velocity threshold value pre-calibrated according to the physiological movement characteristics of the pupil in the student's focused state, and the time points where the instantaneous velocity amplitude is greater than the velocity threshold are selected to form the set of velocity overshoot times for the left eye and the set of velocity overshoot times for the right eye. The time nodes of the left and right eye velocity overshoot time sets are matched one by one to filter out valid moments with a time difference less than the interval between adjacent frame acquisitions, and arranged in chronological order to form a candidate saccade time sequence. Using each candidate saccade time as the center, a fixed duration matching the frame rate of the classroom video acquisition is extended forward and backward, and pupil data from both eyes within this duration are extracted to form a saccade event interval set. The maximum instantaneous velocity amplitude of each data segment in the saccade event interval set is compared with a pre-defined amplitude threshold value based on the physiological characteristics of rapid and micro-saccades. Saccade behaviors with a maximum instantaneous velocity amplitude greater than the amplitude threshold are judged as rapid saccade events. Saccades with amplitude values ​​less than or equal to an amplitude threshold are classified as micro-saccade events. A fixed window duration is set according to the minimum effective time unit for classroom attention monitoring, dividing the timelines of rapid saccade events and micro-saccade events into consecutive, equally long window segments. Rapid saccade events and micro-saccade events within each window segment are counted separately to obtain the raw count values ​​for each segment. A fixed number of consecutive raw count values ​​are added together and divided by the total number of count values, and then the smoothed rapid saccade frequency and smoothed micro-saccade frequency are calculated sequentially. Based on the original timestamps, the smoothed rapid saccade frequency and smoothed micro-saccade frequency are matched according to time correspondence to obtain the student's rapid saccade frequency and micro-saccade frequency. The reason for using the maximum instantaneous velocity amplitude to distinguish between rapid saccades and microsaccades is that the classroom scene uses high-frequency frame sampling to collect data. The instantaneous velocity amplitude is the instantaneous dynamic feature of pupil movement. Compared with the change of eye angle, it is easier to calculate in real time and has stronger anti-interference ability. It can effectively avoid the judgment bias caused by micro-movement of the head in the classroom, light fluctuations and eye image cropping errors. Moreover, from the perspective of physiological movement characteristics, rapid saccades are active gaze shifts with significantly higher instantaneous pupil velocity amplitude, while microsaccades are unconscious and subtle micro-movements with low and stable velocity amplitude. The difference in velocity amplitude between the two is intuitive and easy to quantify. In contrast, the angle requires complete trajectory calculation, which is not real-time and stable enough to meet the needs of real-time monitoring of classroom attention. In step S4 above, the starting and ending coordinates of the pupil movement can be determined from the pupil center coordinate change trajectory corresponding to the microsaccade event and connected to form a motion line segment; taking the horizontal positive direction of the image pixel coordinate system as the reference, the angle between the motion line segment and the horizontal positive direction is measured to obtain the microsaccade angle; the entire range of the microsaccade angle is divided into segments with fixed equal spans to form continuous and non-overlapping microsaccade direction angle intervals; all microsaccade angle values ​​are matched to the corresponding direction angle intervals and the frequency of falling into each interval is counted; The total frequency is obtained by summing the frequencies of each interval. The frequency of each interval is divided by the total frequency to obtain the frequency proportion. The spatial divergence of the micro-saccade direction is quantified based on the uniformity of the frequency proportion to obtain the micro-saccade direction entropy of the student. The occurrence time of all rapid saccade events is extracted, and two consecutive adjacent occurrence times are combined into adjacent event time pairs. The time interval of the adjacent event time pair is obtained by subtracting the time value of the previous occurrence time from the time value of the later occurrence time. The mean of rapid saccade interval is obtained by summing all the time interval values ​​and dividing by the total number of time interval values. The difference between each time interval value and the mean of rapid saccade interval is calculated, the difference results are squared and summed to obtain the sum of squares. The mean square value is obtained by dividing the sum of squares by the total number of time interval values. The standard deviation of rapid saccade interval is obtained by taking the square root of the mean square value. The coefficient of variation of the student's time interval is obtained by dividing the standard deviation of rapid saccade interval by the mean of rapid saccade interval.

[0026] In step S5 above, a large number of micro-saccade direction entropy values ​​of students in the state of classroom focus and latent distraction can be collected. After being arranged according to the value, the values ​​are evenly divided into continuous non-overlapping value intervals. The current micro-saccade direction entropy value of the students is matched to the corresponding interval to obtain the entropy value interval label. The system collects the coefficient of variation (COP) values ​​of students under highly focused conditions, arranges them according to fluctuation amplitude, determines the level boundary values ​​to form the COP level threshold, and compares the student's current COP with the threshold to determine the corresponding level, thus obtaining the COP level. The entropy interval label and the COP level are concatenated and integrated in a fixed order to form a state combination label that simultaneously carries two types of eye movement features. Based on classroom test data, the system establishes a correspondence between the state combination label and the attention state, and matches the student's current state combination label to the unique corresponding attention state to obtain the student's attention state label.

[0027] In step S6 above, the attention state labels, rapid saccade frequency, and microsaccade frequency at the same time point can be combined into focus information for a single time point based on the original timestamp, and all focus information can be connected in the order of collection time to form focus features. Eye-tracking data of students' actual learning status in the classroom is collected, and three states—high concentration, general concentration, and latent inattention—are linked to the corresponding numerical ranges of eye-tracking data. Various data points in the concentration characteristics are matched to their corresponding state ranges. The student's concentration state is determined by the majority of the data points belonging to the range and labeled with a corresponding level, thus obtaining the student's concentration level. For different concentration levels, classroom-appropriate intervention methods are developed to form a set of pre-set intervention suggestions. Student concentration levels are matched with these pre-set intervention suggestions to obtain suitable candidate intervention measures. These candidate intervention measures are then translated into standardized text for real-time classroom implementation, arranged logically by highlighting the state, providing guiding actions, and clarifying learning objectives, forming intervention suggestions for students' latent inattention.

[0028] The beneficial effects are as follows: high-quality binocular images are obtained through high-frequency frame extraction and precise eye region cropping; pupil displacement and velocity are accurately analyzed to form complete time-series data; rapid eye movements and micro-eye movements are reliably decoupled and their frequencies are smoothed using a dual threshold of velocity and amplitude, effectively resisting environmental interference such as head movements and light fluctuations in the classroom; eye movement features are finely quantified by micro-eye movement direction entropy and rapid eye movement temporal interval variation coefficient; combined with dual-feature joint state mapping, the limitations of single indicators are broken, avoiding misjudgment of attention; finally, attention level is comprehensively calibrated based on multi-source features, and personalized classroom intervention suggestions are matched to achieve accurate monitoring of student attention and intelligent identification and targeted guidance of latent inattention.

[0029] In step S1 above, the method of performing high-frequency frame extraction and binocular region cropping on the student's facial video stream to obtain the student's binocular sub-image sequence includes the following steps: Step S101: High-frequency frame sampling is performed on the student's facial video stream to obtain discrete frame images of the student, and the facial regions of the discrete frame images are detected to obtain the student's facial localization box. Step S102: Locate the pupil regions of both eyes in the facial positioning box to obtain the coordinates of the student's left eye region of interest and right eye region of interest. Step S103: Based on the coordinates of the left eye region of interest and the right eye region of interest, the eye region of the discrete frame is cropped to obtain the student's right eye sub-image and left eye sub-image. Step S104: Associate and encapsulate the left eye image and the right eye image to obtain the student's double eye image sequence; In step S101 above, a continuous video stream of students' faces can be continuously acquired by image acquisition devices deployed in the classroom scene. Static images are extracted one by one from the continuous video stream at fixed, very short time intervals. Each extracted independent static image is a discrete frame image of the student. All pixels in the discrete frame image are traversed one by one, and the brightness and color values ​​of each pixel are compared with the pre-set standard brightness range and standard color range of human facial skin. If the brightness value of the pixel falls within the standard brightness range of facial skin and the color value falls within the standard color range of facial skin, then the pixel is determined to belong to the student's face pixel.

[0030] Pixels that meet the criteria are grouped together to form a set of all pixels belonging to the student's face. The positions of pixels in the face pixel set are scanned sequentially in the horizontal direction of the image, and the positions of the leftmost and rightmost pixels are recorded. The positions of pixels in the face pixel set are scanned sequentially in the vertical direction of the image, and the positions of the topmost and bottommost pixels are recorded. The four positions of the leftmost, rightmost, topmost, and bottommost pixels in the horizontal direction are used as the four sides of a rectangle. The area enclosed by these four positions is the smallest rectangular area that can completely contain the face pixel set. This smallest rectangular area is the student's face positioning box.

[0031] In step S102 above, the pixel range covered by the student's face positioning box can be used as the only processing area. Within this area, the brightness value of each pixel is compared one by one with the brightness values ​​of the four pixels adjacent to it (up, down, left, and right). Taking one pixel as an example, let's assume that the brightness of this pixel is... The total brightness difference is And the brightness of the pixel adjacent to this pixel is The brightness of the adjacent pixel below is The brightness of the adjacent pixel on the left is The brightness of the adjacent pixel to the right is L. night ,but: ; The total brightness value of each pixel within the facial localization bounding box area is calculated sequentially using the formula described above. The total brightness difference of each pixel is then sorted from largest to smallest. The set of pixels with the highest total brightness difference exceeding a preset stability limit is identified as the eye contour location, where the brightness is significantly lower than the surrounding area. Pixels with a total brightness difference not exceeding the preset stability limit are determined not to belong to the eye contour location. Within the pixel range corresponding to the eye contour location, brightness values ​​are compared point by point to identify the set of consecutive circularly distributed pixels with the lowest brightness value within that range; this set represents the circular pupil region. Within the circular pupil region, the leftmost, rightmost, topmost, and bottommost pixels in the horizontal direction, vertical direction, and horizontal direction are determined. The four pixel positions corresponding to the left pupil region are combined to form the coordinates of the left eye region of interest; the four pixel positions corresponding to the right pupil region are combined to form the coordinates of the right eye region of interest.

[0032] In step S103 above, the complete pixel matrix of the discrete frame is used as the processing basis. Based on the horizontal and vertical pixel ranges determined by the coordinates of the left eye region of interest, the pixels within the range are extracted and combined into independent image content, which is the student's left eye sub-image. Based on the horizontal and vertical pixel ranges determined by the coordinates of the right eye region of interest, the pixels within the range are extracted and combined into independent image content, which is the student's right eye sub-image.

[0033] In step S104 above, the left-eye and right-eye images taken from the same discrete frame are bound according to the video acquisition time node; the bound left and right eye images are combined to form a single-frame binocular image unit. According to the time sequence of video stream acquisition, the single-frame binocular image units corresponding to the discrete frame are arranged sequentially; the complete set of all single-frame binocular image units after arrangement is the student's binocular image sequence.

[0034] The beneficial effects are as follows: fixed-interval frame extraction completely preserves facial motion information; dual-dimensional pixel filtering improves facial recognition accuracy; the minimum rectangular positioning box accurately limits the processing range, reduces invalid interference, and improves image processing efficiency; limiting the processing area reduces the computational range and speeds up processing. Brightness difference sorting accurately locks the eye contour, improving eye recognition accuracy; low-brightness circular pixel filtering stably locates the pupil; standardized coordinate extraction provides a precise basis for eye cropping, accurately cropping and removing irrelevant information, reducing data volume. Eye separation processing improves the integrity and purity of eye images; time binding ensures image temporal consistency; ordered arrangement constructs a continuous binocular motion dataset, providing standardized raw data support for visual motion feature analysis.

[0035] In step S2 above, the method for analyzing the visual-motion features of the binocular image sequence to obtain the student's binocular pupil data includes the following steps; Step S201: Perform pupil center localization on the left and right eye images in the binocular image sequence to obtain the coordinates of the student's left and right pupil centers. Step S202: Based on the coordinates of the left pupil center and the right pupil center, the pupil displacement is calculated for the left pupil center and the right pupil center of adjacent frames in the binocular image sequence, respectively, to obtain the student's left eye vertical displacement sequence, right eye vertical displacement sequence, left eye horizontal displacement sequence, and right eye horizontal displacement sequence. Step S203: According to the left eye horizontal displacement sequence and the left eye vertical displacement sequence, the velocity component measurement is performed on the left eye displacement between adjacent frames in the binocular sub-image sequence to obtain the student's left eye horizontal velocity sequence and left eye vertical velocity sequence. Step S204: Based on the right eye horizontal displacement sequence and the right eye vertical displacement sequence, analyze the displacement velocity of the right eye between adjacent frames in the binocular sub-image sequence to obtain the student's right eye horizontal velocity sequence and the student's right eye vertical velocity sequence. Step S205: The horizontal velocity sequence of the left eye, the vertical velocity sequence of the left eye, the horizontal velocity sequence of the right eye, and the vertical velocity sequence of the right eye are converged to obtain the pupil data of the student's eyes.

[0036] In step S201 above, each frame of the left and right eye images in the binocular image sequence can be examined one by one. The binocular image sequence is a continuous set of images formed by extracting the binocular regions from the student's facial video stream after high-frequency frame extraction and arranging them in chronological order. High-frequency frame extraction is the operation of continuously extracting independent frames from the student's facial video stream at fixed and extremely short time intervals. The frame density of this operation is much higher than the frame rate of conventional video playback, which can completely capture the subtle movement trajectory of the pupil in a short period of time. The left eye image is an independent image that retains only the student's left pupil and its surrounding eye area; the right eye image is an independent image that retains only the student's right pupil and its surrounding eye area. The brightness value of each pixel in the left eye image is read. Taking the currently detected pixel as the center, the pixels directly adjacent to that pixel in the four diagonal directions are selected as the surrounding pixels. All pixels with brightness values ​​lower than the overall average brightness of the surrounding pixels are filtered out. These pixels together form the pixel area of ​​the left eye pupil.

[0037] Find the midpoint between the leftmost and rightmost points in the horizontal direction of the left pupil pixel region, and then find the midpoint between the topmost and bottommost points in the vertical direction. The intersection of these two midpoints is the center of the left pupil. Record the horizontal and vertical positions of this center position in the image as the coordinates of the left pupil center. Similarly, read the brightness value of each pixel in the right eye sub-image. Using the currently detected pixel as the center, select the pixels directly adjacent to it in the four diagonal directions as the surrounding pixels. Filter out pixels with brightness lower than the overall average brightness of the surrounding pixels to form the right pupil pixel region. Determine the midpoint of the horizontal and vertical intersection of the right pupil pixel region, and record the horizontal and vertical positions of this point as the coordinates of the right pupil center. The pupil center coordinates are used to accurately mark the horizontal and vertical position information of the pupil in the core position of the eye sub-image.

[0038] In step S202 above, the coordinates of the left pupil center corresponding to two temporally adjacent frames in the binocular image sequence can be retrieved, and then, according to the formula, Calculate the horizontal displacement of the left eye in two adjacent frames, where the horizontal position of the center of the left pupil in the previous frame is... In the next frame, the horizontal position of the center of the left pupil is... The horizontal displacement of the left eye in two adjacent frames is The horizontal displacement of the left eye calculated from adjacent frames is arranged sequentially according to the image acquisition time to form a left eye horizontal displacement sequence. The vertical displacement of the left eye pupil between the two frames is obtained by subtracting the vertical position of the left eye pupil center in the previous frame from the vertical position of the left eye pupil center in the next frame. The vertical displacement of the left eye calculated from adjacent frames is arranged sequentially according to the image acquisition time to form a left eye vertical displacement sequence. The right eye pupil center coordinates of adjacent frames are processed using the same position difference calculation method as the left eye pupil center coordinates. The position difference calculation method refers to the calculation method of subtracting the corresponding position value in the same direction of the previous frame from the corresponding position value in the next frame to obtain the displacement value of the adjacent frame. The right eye pupil center coordinates of adjacent frames are processed to obtain the right eye horizontal displacement sequence and the right eye vertical displacement sequence, respectively. The displacement sequence is a set of continuously changing pupil position values ​​arranged in chronological order.

[0039] In step S203 above, the formula can be used. Calculate the horizontal motion velocity of the left eye in a single frame, where the horizontal displacement of the left eye in adjacent frames is... The acquisition time interval between two adjacent frames of the binocular image sequence is, The horizontal motion speed of the left eye in a single frame is The calculated horizontal motion velocity of the left eye is arranged in the order of image acquisition time to form a horizontal velocity sequence of the left eye. The vertical displacement of the left eye in each group of adjacent frames in the vertical displacement sequence of the left eye is divided by the acquisition time interval between two adjacent frames of the binocular image sequence to obtain the vertical motion velocity of the left eye corresponding to that frame. The calculated vertical motion velocity of the left eye is arranged in the order of image acquisition time to form a vertical velocity sequence of the left eye. The velocity component measurement is the process of dividing the displacement value by the acquisition time interval to obtain the pupil motion velocity value per unit time.

[0040] In step S204 above, the horizontal displacement of the right eye in each group of adjacent frames in the right eye horizontal displacement sequence is divided by the acquisition time interval between two adjacent frames in the binocular image sequence to obtain the horizontal motion velocity of the right eye in that frame. The calculated horizontal motion velocities are arranged in the order of image acquisition time to form a horizontal motion velocity sequence of the right eye. Similarly, the vertical displacement of the right eye in each group of adjacent frames in the right eye vertical displacement sequence is divided by the acquisition time interval between two adjacent frames in the binocular image sequence to obtain the vertical motion velocity of the right eye in that frame. The calculated vertical motion velocities are arranged in the order of image acquisition time to form a vertical motion velocity sequence of the right eye.

[0041] In step S205 above, the velocity values ​​belonging to the same image acquisition time node in the left eye horizontal velocity sequence, left eye vertical velocity sequence, right eye horizontal velocity sequence, and right eye vertical velocity sequence can be matched to each other. The unified matching method is as follows: using the acquisition time of the same frame image as the sole criterion, the velocity values ​​corresponding to the same acquisition time in the four sequences are grouped into the same data set; the matched velocity values ​​are then integrated into a continuous overall data set according to the chronological order of acquisition time. Velocity component convergence is the process of uniformly matching and integrating the independent velocity sequences of the four directions of the eyes into a complete time-series data set according to the time dimension; the integrated overall time-series velocity data is the student's pupil data. The pupil data is a time-series information set that completely records the real-time movement velocity of the student's pupils in the horizontal and vertical directions, providing direct raw motion data support for subsequent saccade event decoupling.

[0042] The beneficial effects are as follows: Precisely locating the center coordinates of both pupils provides accurate positional data for displacement calculation; accurately calculating the displacement of pupils in adjacent frames and forming continuous temporal displacement data provides reliable support for velocity calculation; accurately measuring the motion velocity of both eyes in various directions and generating temporal velocity data, comprehensively depicting the real-time dynamic motion characteristics of both pupils. Finally, the temporal aggregation of multi-dimensional velocity data is completed, constructing a standardized and integrated dataset of both pupil motion velocity, comprehensively ensuring the accuracy, completeness, consistency, and standardization of each stage of pupil position detection, displacement calculation, velocity measurement, and data integration, providing stable and reliable raw motion data support for subsequent saccade event decoupling.

[0043] In step S3 above, based on the preset saccade judgment threshold, the saccade events of the pupil data of both eyes are decoupled to obtain the student's rapid saccade events and micro saccade events. The frequency of the rapid saccade events and micro saccade events is statistically analyzed to obtain the student's rapid saccade frequency and micro saccade frequency. Step S301: Perform instantaneous velocity calculation on the pupil data of both eyes to obtain the instantaneous velocity amplitude sequence of the student's right eye and the instantaneous velocity amplitude sequence of the left eye; Step S302: The instantaneous velocity amplitude sequence of the left eye and the instantaneous velocity amplitude sequence of the right eye are compared with the upper limit threshold of the velocity threshold in the preset saccade judgment threshold to obtain the student's left eye velocity overshoot time set and right eye velocity overshoot time set. Step S303: Perform binocular intersection discrimination on the left eye velocity overshoot time set and the right eye velocity overshoot time set to obtain the student's eye saccade candidate time sequence; Step S304: According to the candidate times in the candidate time sequence of saccades, saccade segments are extracted from the pupil data of both eyes to obtain the student's saccade event interval set; Step S305: Based on the amplitude threshold in the preset saccade judgment threshold, the saccade event interval set is judged for event category to obtain the student's rapid saccade events and micro saccade events; Step S306: Divide the rapid saccade event and the micro saccade event into fixed windows to obtain continuous equal-length window segments for the student. Step S307: Count the rapid saccade events and micro saccade events within the continuous equal-length window segment to obtain the original count value of the continuous equal-length window segment; Step S308: Perform a moving average on the original count values ​​to obtain the smoothed rapid saccade frequency and smoothed microsaccade frequency of the students. Step S309: Time-align the smoothed fast saccade frequency and the smoothed micro saccade frequency to obtain the student's fast saccade frequency and the student's micro saccade frequency.

[0044] In step S301 above, the pupil data is a time-series set that completely records the real-time movement speed of the student's pupils in the horizontal and vertical directions. When processing the left eye movement speed at the same time point, the instantaneous velocity amplitude of the left eye is calculated according to the following formula: ; Based on the chronological order of video capture, the instantaneous velocity amplitude of the left eye at each time point is calculated and arranged sequentially to form a sequence of instantaneous velocity amplitudes of the left eye. When processing the right eye motion velocity at the same time point, the instantaneous velocity amplitude of the right eye is calculated according to the following formula: ; Based on the chronological order of video capture, the instantaneous velocity amplitude of the right eye at each time point is calculated and arranged sequentially to form a sequence of instantaneous velocity amplitudes of the right eye.

[0045] in For a moment Left eye horizontal movement speed, For a moment Left eye vertical movement speed, For a moment Right eye horizontal movement speed, For a moment Right eye vertical movement speed, For a moment Instantaneous velocity amplitude of the left eye, For a moment Instantaneous velocity amplitude of the right eye.

[0046] In step S302 above, the velocity threshold is a fixed value pre-calibrated based on the physiological pupillary movement characteristics of students in a focused state during class. When calibrating the velocity threshold, pupillary movement data of multiple students during focused class are collected, and the instantaneous pupillary velocity amplitude when the student is focused and no saccades occur is extracted. After summarizing the amplitude values, the upper limit of the interval with the highest frequency of occurrence is selected as the velocity threshold. The amplitude value of each time point in the left eye instantaneous velocity amplitude sequence is extracted one by one, and each value is compared with the velocity threshold. When the left eye instantaneous velocity amplitude is greater than the velocity threshold, this time point is marked as a left eye velocity overshoot moment. All marked left eye velocity overshoot moments are summarized in chronological order to obtain a set of left eye velocity overshoot moments. The amplitude value of each time point in the right eye instantaneous velocity amplitude sequence is extracted one by one, and each value is compared with the velocity threshold. When the right eye instantaneous velocity amplitude is greater than the velocity threshold, this time point is marked as a right eye velocity overshoot moment. All marked right eye velocity overshoot moments are summarized in chronological order to obtain the right eye velocity overshoot moment set.

[0047] In step S303 above, the left-eye velocity overshoot time set is the set of time points where the instantaneous velocity amplitude of the left eye exceeds the velocity threshold; the right-eye velocity overshoot time set is the set of time points where the instantaneous velocity amplitude of the right eye exceeds the velocity threshold. Each time point in the left-eye velocity overshoot time set is taken out one by one and matched with the time points in the right-eye velocity overshoot time set. If the time difference between two time points is less than the interval between two adjacent frames during video capture, these two time points are determined to be the same valid time point. The determined valid times are arranged in chronological order; the arranged valid times are combined to form a candidate saccade time sequence.

[0048] In step S304 above, the candidate saccade time sequence is a set of valid moments of binocular synchronization velocity overshoot arranged chronologically; the binocular pupil data is time-series data of continuously recorded pupil movement velocity. Taking each candidate moment in the candidate saccade time sequence as the center, a fixed duration matching the frame rate of the classroom video is extended in both directions before and after this center moment; the binocular pupil data within this duration range is completely extracted. Each extracted data segment corresponds to an independent saccade behavior; the extracted data segments corresponding to all independent saccade behaviors are summarized in chronological order to form a saccade event interval set.

[0049] In step S305 above, the amplitude threshold is a fixed value pre-calibrated based on the physiological movement characteristics of rapid and micro-saccades in the classroom. When calibrating the amplitude threshold, real saccade data from multiple students in the classroom are collected to distinguish between rapid saccades (intentional eye movements) and micro-saccades (unintentional micro-movements). The instantaneous velocity amplitudes corresponding to the two types of saccades are extracted, and the boundary value between the rapid and micro-saccade amplitude intervals is used as the amplitude threshold. The maximum instantaneous velocity amplitude of each data segment within the saccade event interval set is extracted one by one, and this maximum amplitude is compared with the amplitude threshold. When the maximum instantaneous velocity amplitude of this data segment is greater than the amplitude threshold, the saccade behavior corresponding to this data segment is determined as a rapid saccade event; when the maximum instantaneous velocity amplitude of this data segment is less than or equal to the amplitude threshold, the saccade behavior corresponding to this data segment is determined as a micro-saccade event.

[0050] In step S306 above, the duration of the fixed window is pre-set according to the minimum effective time unit for classroom attention monitoring. Based on the complete timeline of the rapid eye movement (REM) event, the timeline is sequentially divided into multiple time segments according to the pre-set fixed window duration. Each segment is of identical length, and the segments are consecutive without gaps. The segmented time segments are arranged chronologically to form continuous, equal-length window segments corresponding to the REM event. Similarly, based on the complete timeline of the micro-saccade (MS) event, the timeline is sequentially divided into multiple time segments according to the same fixed window duration as the REM event. Each segment is of identical length, and the segments are consecutive without gaps. All segmented time segments are arranged chronologically to form continuous, equal-length window segments corresponding to the MS event.

[0051] In step S307 above, the continuous equal-length window segment consists of multiple time intervals of consistent duration, each independent time interval being a dedicated statistical unit. The occurrence time of a rapid eye-saccharidation event is automatically matched with the time intervals of each statistical unit within the corresponding continuous equal-length window segment to determine the unique statistical unit to which each rapid eye-saccharidation event belongs. After all rapid eye-saccharidation events are assigned to intervals, rapid eye-saccharidation events belonging to the same statistical unit are automatically accumulated; the total accumulated count is directly used as the original rapid eye-saccharidation count value for the current statistical unit. The original rapid eye-saccharidation count values ​​corresponding to the statistical units are arranged sequentially according to the chronological order of the continuous equal-length window segments to form the original count values ​​for rapid eye-saccharidation events. The occurrence time of a micro-eye-saccharidation event is automatically matched with the time intervals of each statistical unit within the corresponding continuous equal-length window segment to determine the unique statistical unit to which each micro-eye-saccharidation event belongs. After all micro-saccade events are assigned to intervals, micro-saccade events belonging to the same statistical unit are automatically accumulated and counted; the total number obtained is directly used as the raw count value of the current statistical unit. According to the time sequence of consecutive equal-length window segments, the raw count values ​​of micro-saccades corresponding to the statistical units are arranged sequentially to form the raw count value of micro-saccade events.

[0052] In step S308 above, the moving average process selects a fixed number of consecutive raw count values ​​as a group of calculation units. First, consecutive rapid saccade count values ​​are selected, and all raw count values ​​within the group are summed to obtain a total value. This sum is divided by the total number of raw count values ​​involved in the calculation within the group, and the result is the smoothed rapid saccade frequency corresponding to the current calculation unit. Then, new consecutive raw count values ​​are selected sequentially according to time, and the above addition and division calculation steps are repeated. All calculated results are arranged in chronological order to form the smoothed rapid saccade frequency. The same number of consecutive micro-saccade raw count values ​​as the rapid saccade calculation are selected as a group of calculation units, and all raw count values ​​within the group are summed to obtain a total value. This sum is divided by the total number of raw count values ​​involved in the calculation within the group, and the result is the smoothed micro-saccade frequency corresponding to the current calculation unit. New consecutive raw count values ​​are selected sequentially according to time, and the above addition and division calculation steps are repeated. All calculated results are arranged in chronological order to form the smoothed micro-saccade frequency.

[0053] In step S309 above, the smoothed fast saccade frequency is a set of smoothed fast saccade values ​​arranged chronologically after being processed by moving average; the smoothed micro-saccade frequency is a set of smoothed micro-saccade values ​​arranged chronologically after being processed by moving average. Using the original timestamp of the video capture as the sole matching criterion, the values ​​with the same original timestamp in the smoothed fast saccade frequency and smoothed micro-saccade frequency are matched one-to-one. The matched smoothed fast saccade frequency value is the student's fast saccade frequency at that timestamp; the fast saccade frequencies corresponding to the timestamps are combined chronologically to form the complete fast saccade frequency. Similarly, the matched smoothed micro-saccade frequency value is the student's micro-saccade frequency at that timestamp; the micro-saccade frequencies corresponding to all timestamps are combined chronologically to form the complete micro-saccade frequency.

[0054] The beneficial effects are as follows: the instantaneous pupillary movement rate is accurately quantified by synthesizing velocity components while preserving complete velocity characteristics; thresholds are calibrated based on classroom physiological characteristics to accurately screen and match saccade moments, eliminating interference and ensuring recognition accuracy; complete extraction of saccade full-cycle data ensures the integrity of the analysis; rapid saccades and micro-saccades are accurately decoupled to ensure accurate classification; unified statistical dimensions enable frequency quantification and standardized calculation; data is smoothed by moving average to eliminate fluctuations and improve numerical reliability; and finally, time-series alignment is completed to generate a standardized saccade frequency, providing accurate, reliable, and standardized eye movement frequency characteristics to support attention assessment.

[0055] In step S4 above, the directional divergence of micro-saccade events is evaluated to obtain the micro-saccade directional entropy of students, and the temporal interval dispersion analysis of students' rapid saccade events is performed to obtain the temporal interval variation coefficient of students. Step S401: Extract the direction angle of the micro-saccade event to obtain the student's micro-saccade angle; Step S402: Discretize the entire range of microsaccade angles to obtain the microsaccade direction angle range of the student. Step S403: Based on the microsaccade direction angle interval, perform interval fall frequency statistics on the microsaccade angle to obtain the frequency of the student's microsaccade direction angle value appearing in the interval within the microsaccade direction angle. Step S404: Based on the frequency of occurrence in the interval, perform uncertainty quantification assessment on the spatial divergence of the micro-saccade direction to obtain the micro-saccade direction entropy of the student. Step S405: Correlate adjacent time points of the rapid eye saccade events to obtain the student's adjacent event time pairs; Step S406: Extract the time difference between adjacent event time pairs to obtain the time interval value of adjacent event time pairs; Step S407: Perform mean measurement on the time interval value sequence to obtain the mean of the students' rapid saccade intervals; Step S408: Perform deviation amplitude analysis on the time interval value to obtain the standard deviation of the student's rapid saccade interval; Step S409: Based on the standard deviation and mean of rapid saccade intervals, the relative fluctuation of the time interval values ​​is mapped by the rate of variation to obtain the student's time interval variation coefficient.

[0056] In step S401 above, the micro-saccade event originates from the student's unconscious pupil micro-movement behavior record obtained after high-frequency frame extraction, binocular region cropping, pupil data analysis, and saccade event decoupling in a classroom setting. From the trajectory of the pupil center coordinate changes corresponding to the micro-saccade event, the starting and ending coordinate points of the pupil movement are identified; these points are then connected by a straight line to form a pupil movement line segment. Using the horizontal positive direction of the binocular sub-image pixel coordinate system as a fixed measurement reference, the angle between the pupil movement line segment and the horizontal positive direction of the pixel coordinate system is measured; this angle is the student's micro-saccade angle.

[0057] In step S402 above, the entire range of possible microsaccade angles is divided into continuous segments with fixed and equal angle spans; each segment forms an independent angle range that constitutes a direction angle interval. All direction angle intervals completely cover all microsaccade angles without overlap or omission; all direction angle intervals together constitute the student's microsaccade direction angle interval.

[0058] In step S403 above, all micro-saccade angle values ​​obtained after extracting the direction angles of all micro-saccade events are extracted. Each micro-saccade angle value is compared with the micro-saccade direction angle interval one by one to determine the unique direction angle interval to which the micro-saccade angle value belongs. When each micro-saccade angle value falls into the corresponding direction angle interval, the count value of that direction angle interval is increased by one unit. After completing the interval assignment comparison and counting of all micro-saccade angle values, the final count value corresponding to each direction angle interval is the frequency of the student's micro-saccade direction angle value appearing in the interval within the micro-saccade direction angle.

[0059] In step S404 above, the frequencies of all intervals within the microsaccade direction angle interval are summed to obtain the total frequency value. The frequency of each interval is divided by the total frequency value to obtain the frequency proportion of that interval. The frequency proportion values ​​of all intervals are compared for uniformity; the closer the frequency proportion values ​​are, the higher the spatial divergence of the microsaccade direction. According to the degree of spatial divergence, it is directly mapped to the corresponding quantized value; this quantized value is the student's microsaccade direction entropy.

[0060] In step S405 above, the rapid saccade events originate from records of students' active gaze shifting behavior obtained after high-frequency frame extraction, binocular region cropping, pupil data analysis, and saccade event decoupling in a classroom setting. The effective occurrence times corresponding to all rapid saccade events are extracted; according to the chronological order of the effective occurrence times, every two consecutive adjacent effective occurrence times are paired. Each pair of paired effective occurrence times represents the student's adjacent event time pair.

[0061] In step S406 above, any pair of adjacent event time points is selected, and the unified timeline values ​​of the previous and next valid occurrence times within that pair are obtained. The timeline difference between the previous and next valid occurrence times is the time interval value of that pair of adjacent event time points. The time interval values ​​are calculated sequentially for all pairs of adjacent event time points in the above manner. The time interval values ​​are then arranged in chronological order to form a complete set.

[0062] In step S407 above, all time interval values ​​in the time interval value set are summed sequentially to obtain the total time interval value; the total time interval value is divided by the total number of time interval values ​​in the time interval value set to obtain the average rapid eye saccade interval of the students.

[0063] In step S408 above, the difference between each time interval value in the time interval value set and the mean of rapid saccade intervals is calculated to obtain a single deviation value; each single deviation value is multiplied by itself to obtain a deviation square value. All deviation square values ​​are summed sequentially to obtain a total deviation square; the total deviation square is divided by the total number of time interval values ​​in the time interval value set to obtain the average deviation square value. The square root of the average deviation square value is then taken to obtain the final value; this final value is the standard deviation of the student's rapid saccade interval.

[0064] In step S409 above, the standard deviation of the rapid saccade interval is divided by the mean of the rapid saccade intervals to obtain the calculation result; this calculation result directly corresponds to the relative fluctuation of the rapid saccade interval value. The higher the relative fluctuation, the larger the value of the calculation result; this calculation result is the student's coefficient of variation of the temporal interval.

[0065] The beneficial effects are as follows: it can accurately lock the pupil movement coordinates and standardize the angle measurement benchmark, stably obtaining micro-saccade angle values; by uniformly dividing angle intervals, accurately matching angles and intervals, and statistically analyzing frequencies, it solidifies the data foundation for micro-saccade directional divergence assessment; it quantifies the spatial divergence degree by frequency proportion and maps it to entropy values, improving the accuracy of micro-saccade directional feature quantification. Simultaneously, it standardizes the extraction and pairing of rapid saccade moments, accurately calculates and collects temporal interval data; it accurately calculates the interval mean and standard deviation, and obtains the coefficient of variation through the ratio, intuitively representing the degree of temporal interval fluctuation. Overall, it achieves accurate quantification of micro-saccade directional divergence and rapid saccade temporal dispersion, providing reliable multi-dimensional eye movement feature support for attention state assessment.

[0066] In step S5 above, the micro-eye saccade direction entropy and the temporal interval variation coefficient are mapped in the state space to obtain the student's attention state label. Step S501: The entropy of the micro-eye saccade direction is calibrated in intervals to obtain the entropy value interval labels for students; Step S502: Based on the preset coefficient of variation level threshold, classify the coefficient of variation of the time interval into levels to obtain the coefficient of variation level of the students. Step S503: Combine the entropy value interval label with the coefficient of variation level to obtain the student's state combination label. Step S504: Perform state mapping on the student's state combination identifier to obtain the student's attention state label.

[0067] In step S501 above, the microsaccade direction entropy is a value obtained after evaluating the divergence of student microsaccade events, used to reflect the spatial uniformity of microsaccade direction divergence. First, a large number of microsaccade direction entropy values ​​are collected in a real classroom setting when students are attentive and when they exhibit latent inattention. All collected values ​​are arranged in ascending order. Then, based on the overall distribution of the arranged values, the entire value range is evenly divided into multiple continuous and non-overlapping fixed value intervals, with each interval covering a consistent value range. The student's current microsaccade direction entropy value is compared one by one with the pre-divided fixed value intervals to find the unique fixed value interval that completely contains the value. This unique fixed value interval is used as the student's entropy interval label. The entropy interval label is a feature identifier used to clearly mark the value range to which the student's microsaccade direction divergence belongs.

[0068] In step S502 above, the coefficient of variation (CVA) is calculated by first determining the standard deviation of the students' rapid saccade intervals, and then dividing that standard deviation by the mean of the rapid saccade intervals. This CVA value reflects the relative fluctuation range of the rapid saccade intervals. CVA values ​​are collected in advance in a real classroom while multiple students maintain high concentration. These values ​​are arranged from low to high fluctuation range. Several fixed level boundary values ​​are determined based on key nodes in the value distribution; these boundary values ​​together form a preset CVA level threshold. The students' current CVA is compared sequentially with the preset CVA level thresholds. Based on the comparison results, the level interval to which the coefficient belongs is determined, and the fixed level corresponding to this level interval is the student's CVA level. The CVA level is a feature identifier used to clearly indicate the level to which the fluctuation of the students' rapid saccade intervals belongs.

[0069] In step S503 above, the entropy interval label is used to reflect the divergence of the student's micro-saccade direction; the coefficient of variation level is used to reflect the fluctuation of the student's rapid saccade time interval. Following a pre-set fixed arrangement, the entropy interval label and the coefficient of variation level are sequentially concatenated and integrated, merging the two feature identifiers into an inseparable whole identifier; this whole identifier is the student's state combination identifier. The state combination identifier is a comprehensive feature identifier that simultaneously carries the two core eye movement features of the student: micro-saccades and rapid saccades.

[0070] In step S504 above, the state combination label is a comprehensive label integrating two core features of students' eye movements (both eyes). Based on extensive classroom testing data, a one-to-one correspondence is established between the state combination labels and different student attention states in the classroom setting. The student's current state combination label is matched one by one with all the contents of this correspondence to accurately find the attention state uniquely corresponding to that state combination label; this uniquely corresponding attention state is the student's attention state label. The attention state label is the final result label that directly determines the student's actual current state of attention in the classroom.

[0071] The beneficial effects are as follows: It accurately quantifies the spatial dispersion uniformity of micro-saccades and stably labels their numerical range, providing precise micro-saccade feature basis for attention state determination; it accurately characterizes the relative fluctuation amplitude of rapid saccade time intervals and clearly divides the fluctuation level hierarchy, providing reliable rapid saccade feature basis for attention state determination; it integrates the core eye movement features of micro-saccades and rapid saccades to form a unified comprehensive identifier, ensuring the synergistic association and complete presentation of the two types of eye movement features; and it achieves accurate mapping from eye movement features to attention state, directly determining the actual attention state of students in the classroom, and improving the accuracy and uniqueness of attention state recognition.

[0072] In step S6 above, the attention state label, rapid saccade frequency, and microsaccade frequency are comprehensively evaluated to obtain the student's attention level, and intervention suggestions for the student's latent inattention are generated based on the attention level. Step S601: Multi-source feature aggregation is performed on the attention state label, rapid saccade frequency, and microsaccade frequency to obtain the student's attention characteristics; Step S602: Classify the attention characteristics to obtain the student's attention level. Step S603: Based on the preset intervention suggestions, perform measure correlation matching on the student's attention level to obtain the student's candidate intervention measures; Step S604: The candidate intervention measures for students are assembled into an intervention text to obtain intervention suggestions for students' latent inattentiveness.

[0073] In step S601 above, the original timestamps generated during the facial video stream acquisition process can be retrieved first. Then, based on each independent original timestamp, the attention state label, rapid saccade frequency, and microsaccade frequency corresponding to that timestamp can be found. The attention state label is a unique identifier used to mark the student's current attention state after mapping the microsaccade direction entropy with the temporal interval variation coefficient. The rapid saccade frequency is the stable number of rapid saccade events per unit time obtained after performing fixed window division, event count, moving average smoothing, and temporal alignment. The microsaccade frequency is the stable number of microsaccade events per unit time obtained after performing fixed window division, event count, moving average smoothing, and temporal alignment. The attention state label, rapid saccade frequency, and microsaccade frequency under the same timestamp are combined in a fixed order to form the focus information of a single time node. Then, the focus information formed by combining all time nodes is sequentially connected according to the time sequence of video acquisition to form a continuous and uninterrupted complete information set. This continuous and uninterrupted complete information set is the student's focus characteristic.

[0074] In step S602 above, eye-tracking data of multiple students during actual classroom learning can be collected first, simultaneously recording their actual performance in three states: high concentration, general concentration, and latent inattention. Then, these three states are associated with corresponding attention state labels, rapid eye saccade frequency, and micro-saccade frequency data, respectively, to statistically determine fixed numerical ranges for each of the three data categories under each state. Each data point in the concentration feature is compared sequentially with the numerical ranges corresponding to the three states to determine the state range to which each data point belongs. The state range results for the three data points are then summarized, and the student's final concentration state is determined based on the range to which the majority of data points belong. This state is then mapped to a preset level classification, and the resulting level classification is the student's concentration level. Specifically, when statistically calculating the numerical ranges, all data of the same type for all students in the same state are first summarized. These data are then arranged in ascending order, and after removing outliers at the beginning and end, the minimum and maximum values ​​of the remaining data are taken as the boundaries of the numerical range for that state.

[0075] In step S603 above, intervention methods can be developed based on the actual classroom teaching scenario, targeting different levels of focus, to form a complete set of pre-set intervention suggestions. High focus levels correspond to intervention methods of continuous attention and positive encouragement; medium focus levels correspond to intervention methods of gentle prompts and eye guidance; and low focus levels correspond to intervention methods of proactive questioning and content focus. Then, the student's focus level is matched perfectly with the level categories in the pre-set intervention suggestion set, selecting only the intervention method that exactly matches the student's focus level; this selected intervention method is the candidate intervention measure suitable for that student.

[0076] In step S604 above, the practical content of the candidate intervention measures can first be transformed into standardized text content suitable for real-time classroom use. Then, following the logic of first clarifying the student's current state, then providing specific guidance actions, and finally specifying the learning objectives that need attention, these standardized texts can be organized and arranged in sequence. The resulting coherent and directly executable complete text content is the intervention suggestion for the student's latent inattentiveness.

[0077] The beneficial effects are as follows: it enables precise temporal alignment and aggregation of multi-source eye-tracking features, constructing a complete and continuous attention feature system; it scientifically defines the criteria for judging attention levels to improve the accuracy and objectivity of level labeling; it achieves precise matching between intervention measures and attention levels to ensure the pertinence and applicability of intervention measures; at the same time, it standardizes the text format of intervention recommendations to ensure their feasibility and practicality, comprehensively improving the accuracy of attention assessment, the reliability of level judgment, the pertinence of intervention matching, and the effectiveness of intervention recommendations in implementation.

[0078] The flowchart provided in this embodiment is not intended to indicate that the operations of the method will be performed in any particular order, or that all operations of the method are included in every case. Furthermore, the method may include additional operations. Within the scope of the technical concept provided by the method in this embodiment, additional variations can be made to the above method.

[0079] It should be understood that in some embodiments, the components may be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods may be implemented using software or firmware stored in memory and executed by a suitable instruction execution system.

[0080] This embodiment also provides an intelligent educational companion system based on speech signal feature extraction, such as... Figure 2 As shown, the computer device 30 includes a memory 20, a processor 10, and a computer program 21 stored on the memory 20 and running on the processor 10. When executed by the processor 10, the computer program 21 implements the steps of the intelligent education companion method based on voice signal feature extraction of any of the above embodiments.

[0081] The computer program 21 used to perform the operations of this invention may be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages ​​and procedural programming languages. The computer program 21 may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a Local Area Network (LAN) or Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, Field-Programmable Gate Arrays (FPGAs), or Programmable Logic Arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing status information of the computer-readable program instructions.

[0082] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.

Claims

1. A method for real-time monitoring of student concentration based on a classroom scene, characterized by, include: S1. Perform high-frequency frame extraction and eye region cropping on the student's facial video stream to obtain the student's eye sub-image sequence; S2. Analyze the visual motion features of the binocular sub-image sequence to obtain the pupil data of the student's eyes; S3. Based on a preset saccade determination threshold, the saccade events of the pupil data of both eyes are decoupled to obtain the student's rapid saccade events and micro saccade events. The frequency of the rapid saccade events and the micro saccade events is statistically analyzed to obtain the student's rapid saccade frequency and micro saccade frequency. S4. Evaluate the directional divergence of the micro-saccade events to obtain the micro-saccade directional entropy of the student, and perform temporal interval dispersion analysis on the student's rapid saccade events to obtain the student's temporal interval variation coefficient. S5. Map the micro-saccade direction entropy with the temporal interval variation coefficient in a state space to obtain the student's attention state label. S6. The attention state label, the rapid saccade frequency, and the microsaccade frequency are comprehensively evaluated to obtain the student's attention level, and intervention suggestions for the student's latent inattention are generated based on the attention level.

2. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The process of performing high-frequency frame extraction and binocular region cropping on the student's facial video stream to obtain the student's binocular sub-image sequence includes: The student's facial video stream is sampled at high frequency to obtain discrete frame images of the student, and the facial region of the discrete frame images is detected to obtain the student's facial positioning box. The pupil regions of both eyes are located using the facial positioning frame to obtain the coordinates of the student's left eye region of interest and right eye region of interest. Based on the coordinates of the left and right eye regions of interest, the discrete frame images are cropped to obtain the student's right and left eye sub-images. The left eye image and the right eye image are associated and encapsulated to obtain the student's eye image sequence.

3. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The process of analyzing the visual-motion features of the binocular image sequence to obtain the student's binocular pupil data includes: Pupil center localization is performed on the left and right eye images in the binocular image sequence to obtain the coordinates of the student's left and right pupil centers. Based on the coordinates of the left and right pupil centers, pupil displacement is calculated for the left and right pupil centers of adjacent frames in the binocular image sequence to obtain the student's left eye vertical displacement sequence, right eye vertical displacement sequence, left eye horizontal displacement sequence, and right eye horizontal displacement sequence. According to the left eye horizontal displacement sequence and the left eye vertical displacement sequence, the velocity component measurement is performed on the left eye displacement between adjacent frames in the binocular sub-image sequence to obtain the student's left eye horizontal velocity sequence and left eye vertical velocity sequence. Based on the right eye horizontal displacement sequence and the right eye vertical displacement sequence, the displacement velocity of the right eye between adjacent frames in the binocular sub-image sequence is analyzed to obtain the student's right eye horizontal velocity sequence and the student's right eye vertical velocity sequence. The horizontal velocity sequence of the left eye, the vertical velocity sequence of the left eye, the horizontal velocity sequence of the right eye, and the vertical velocity sequence of the right eye are converged by line velocity component aggregation to obtain the pupil data of the student's eyes.

4. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The step of decoupling saccade events from the pupil data of both eyes based on a preset saccade determination threshold to obtain the student's rapid saccade events and micro-saccade events includes: Instantaneous velocity calculations were performed on the pupil data of both eyes to obtain the instantaneous velocity amplitude sequence of the student's right eye and the instantaneous velocity amplitude sequence of the left eye; The instantaneous velocity amplitude sequence of the left eye and the instantaneous velocity amplitude sequence of the right eye are compared with the upper limit threshold of the velocity threshold in the preset saccade determination threshold to obtain the student's left eye velocity overshoot time set and right eye velocity overshoot time set; By performing binocular intersection discrimination on the left eye velocity overshoot time set and the right eye velocity overshoot time set, the candidate saccade time sequence of the student is obtained; According to the candidate times in the candidate time sequence of the eye saccades, the eye saccade segments are extracted from the pupil data of both eyes to obtain the eye saccade event interval set of the student; Based on the amplitude threshold in the preset saccade determination threshold, the saccade event interval set is classified into event categories to obtain the student's rapid saccade events and micro-saccade events.

5. The method for real-time monitoring of student concentration in a classroom setting as described in claim 4, characterized in that, The frequency statistics of the rapid saccade events and the microsaccade events are performed separately to obtain the student's rapid saccade frequency and microsaccade frequency, including: The rapid saccharidation event and the micro saccharidation event are respectively divided into fixed windows to obtain continuous equal-length window segments for the student; The rapid saccharidation events and micro saccharidation events within the continuous equal-length window segment are counted to obtain the original count value of the continuous equal-length window segment; The original count values ​​are processed by a moving average to obtain the smoothed rapid saccade frequency and smoothed microsaccade frequency of the students; The smoothed rapid saccade frequency and the smoothed micro saccade frequency are time-aligned to obtain the student's rapid saccade frequency and the student's micro saccade frequency.

6. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The process of evaluating the directional divergence of the microsaccade events to obtain the microsaccade direction entropy of the student includes: The direction angle of the microsaccade event is extracted to obtain the microsaccade angle of the student. Discretize the entire range of the microsaccade angle to obtain the microsaccade direction angle range of the student; Based on the microsaccade direction angle range, the frequency of the microsaccade angle falling within the range is statistically analyzed to obtain the frequency of the student's microsaccade direction angle value appearing within the range of the microsaccade direction angle. Based on the frequency of occurrence of the interval, the degree of spatial divergence of the microsaccade direction is quantitatively evaluated to obtain the microsaccade direction entropy of the student.

7. The method for real-time monitoring of student attention in a classroom setting as described in claim 6, characterized in that, The temporal interval dispersion analysis of the student's rapid eye saccades is performed to obtain the student's temporal interval variation coefficient, including: By correlating adjacent moments of the rapid eye saccade events, the student's adjacent event moment pairs can be obtained; Extract the time difference value of the adjacent event time pairs to obtain the time interval value of the adjacent event time pairs; The mean value of the student's rapid saccade interval is obtained by performing mean measurement on the time interval value sequence. The deviation amplitude is extracted from the time interval value to obtain the standard deviation of the student's rapid saccade interval; Based on the standard deviation of the rapid saccade interval and the mean of the rapid saccade interval, the relative fluctuation of the time interval value is mapped by the rate of variation to obtain the coefficient of variation of the time interval of the student.

8. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The step of mapping the micro-saccade direction entropy to the temporal interval variation coefficient in a state space to obtain the student's attention state label includes: The entropy of the micro-eye saccade direction is calibrated in intervals to obtain the entropy value interval label of the student; Based on a preset coefficient of variation level threshold, the coefficient of variation of the time interval is classified into levels to obtain the coefficient of variation level of the student. The entropy value interval label and the coefficient of variation level are combined and aggregated to obtain the student's state combination label; The student's state combination identifier is mapped to obtain the student's attention state label.

9. The method for real-time monitoring of student attention in a classroom setting as described in claim 1, characterized in that, The student's attention level is obtained by comprehensively evaluating the attention state label, rapid saccade frequency, and microsaccade frequency, and intervention suggestions for latent inattentiveness are generated based on the attention level, including: Multi-source feature aggregation is performed on the attention state label, the rapid saccade frequency, and the microsaccade frequency to obtain the student's attention characteristics; The student's attention level is obtained by classifying the attention characteristics into different levels. Based on the preset intervention suggestions, the student's attention level is matched with the corresponding measures to obtain the candidate intervention measures for the student; The candidate intervention measures for the students are assembled into intervention texts to obtain intervention suggestions for the students' latent inattentiveness.

10. A real-time student attention monitoring system, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the steps of the real-time monitoring method for student attention according to any one of claims 1-9.